July's AI Video Generation Battle: Over 20 Billion in Capital Floods In—Who's Made the Finals?

AI video generationworld modelSeedanceKling AIPixVersefundingunicornFlovaAIVidu
1 hours agoSource: blockweeks.com
July's AI Video Generation Battle: Over 20 Billion in Capital Floods In—Who's Made the Finals?

Aishi Technology

July is not over yet, and the AI video generation track is already burning hot.

From July 3 to July 23, in just 20 days, five companies announced large-scale financing in succession.

Kling AI
raised $3 billion, Shengshu Technology raised $500 million, Aishi Technology raised several hundred million dollars in Series C+, Zhixiang Future raised 1.5 billion RMB in Series C, and even FlovaAI, a new company established less than a year ago, raised $80 million in angel round.

The total financing of the first four unicorn-level companies exceeds 26.2 billion RMB. That is to say, the money poured in during July 2026 alone is more than the entire track raised in the past two years (2024-2025) combined.

Why did they all concentrate on financing in July? What exactly happened? This should not be a coincidence.

Kling AI: BAT historic co-investment, ARR quadrupled in one year

On July 3,

Kling AI
announced a $3 billion Series A financing cap, led by CPE Yuanfeng and Guofang Innovation, with 34 institutions participating—in the investor list, Tencent, Alibaba Cloud, and Baidu appeared together for the first time. FountainBridge Capital served as financial advisor.

The $3 billion Series A financing cap corresponds to a post-investment valuation of approximately $18 billion; the currently signed portion is about 19 billion RMB.

Why did capital dare to give this number? It has to be that Kuaishou's Q1 2026 report gave the market confidence.

Kuaishou's Q1 2026 financial report disclosed that

Kling AI
achieved single-quarter revenue of over 650 million RMB, a year-on-year increase of over 300%. The annualized revenue run rate (ARR) soared from $100 million in March 2025 to nearly $500 million in March 2026, quadrupling year-on-year.

This number is unmatched in the AI video track—except for ByteDance's Seedance, no other domestic company has reached this scale.

Kling's revenue is driven by two wheels: B-end enterprise customer API calls + P-end paid member subscriptions. As of June 2026,

Kling AI
has over 100 million global users and nearly 50,000 enterprise customers. For example, Kling deeply participated in the virtual scene and special effects production of the popular Chinese historical drama "The Peaceful Year," and supported the generation of hundreds of high-quality shots in the Hollywood series "The David Dynasty."

In terms of products, in February this year, Kling launched the 3.0 series model, achieving up to 15-second video generation, storyboard control, subject consistency, and simultaneous audio and video output.

Two details are worth noting. First, the valuation is lower than the rumored $20 billion target in May, with more pragmatic pricing, and the deal actually closed faster; second, a performance clause was embedded in the announcement—if Beijing Kling fails to complete an IPO before October 2031, investors have the right to demand repurchase at principal plus 8% annual simple interest.

The combination of spin-off, financing, equity incentives, and repurchase clauses points to a clear direction: Kling is paving the way for an independent listing.

Shengshu Technology: $500 million financing on the eve of IPO, from video model to "world model"

On July 6, Shengshu Technology announced a $500 million Series B+ financing.

But if you look at the timeline, the company has taken three rounds in 5 months: over 600 million RMB in Series A+ in February (led by Zhongguancun Science City and Xinglian Capital), 2 billion RMB in Series B in April (led by Alibaba Cloud), plus this $500 million round in July, totaling over 5 billion RMB.

Three rounds in 5 months, this pace indicates one thing: capital is scrambling to get on board. Why the urgency?

At the end of March this year, Shengshu Technology completed a joint-stock reform—a standard move before listing. Market sources say it may start the Hong Kong IPO process as early as the first half of the year. Therefore, capital is rushing to board on the eve of the IPO.

Shengshu Technology's core product is

Vidu
, launched globally in July 2024, and has iterated to the

Vidu
S1 version in three years—a real-time interactive model supporting voice-controlled frames and unlimited duration generation (1 minute to 2 hours).

In 2025, Shengshu Technology achieved over 10x growth in users and revenue, with over 400 million cumulative generated videos; especially, its B-end customer lineup is quite impressive: in the advertising and e-commerce field, it has JD.com, Alibaba 1688, Amazon, Meituan, Focus Media, BlueFocus, L'Oréal, Anta; in the film and animation field, it covers Tencent Animation, Yuewen Group, CCTV Animation, iQiyi, Mango TV; in the gaming field, it serves Lilith Games and 37 Interactive Entertainment.

Shengshu Technology is not satisfied with just video generation; this year it began to extend towards embodied intelligence.

In April 2026, Shengshu Technology officially released the general world action model Motubrain, adopting the World Action Model (WAM) technical route, unifying perception, prediction, and action modeling, enabling robots to have "one brain to foresee" and "one brain to multitask," while the commercial version MotuBrain entered industrial-level verification.

This Tsinghua-affiliated company, established only three years ago, is rapidly evolving from a "video model company" to a "general world model platform." After the $500 million arrives, the funds will be mainly used for the research and development of the "general world model."

This is the story that capital truly values.

Aishi Technology: Three-year unicorn, 150 million global users

On July 14, Aishi Technology's Series C+ round was led by Alibaba. Together with the $300 million Series C round completed in March (led by CDH Investments, with nearly 20 participants including China Ruyi and 37 Interactive Entertainment), the total Series C financing reached 2.98 billion RMB, with a post-investment valuation exceeding $2 billion.

Alibaba had previously led Aishi's $60 million Series B round in September 2025, and this time it continued to increase its stake in Series C+. Ant Group also led the A2 round of over 100 million RMB as early as April 2024.

Alibaba's continuous investment in Aishi shows that it regards Aishi as a strategic layout point in the AI video track.

As of March 2026, Aishi Technology has over 150 million global users, 15 million monthly active users, and an annual recurring revenue (ARR) of $40 million.

Interestingly, co-founder Xie Xuzhang revealed that the training cost of Aishi's same-level model is only about 10% of that of its peers. The key path to this cost advantage lies in screening high-quality data, reducing ineffective training iterations, and utilizing an optimized Diffusion Transformer architecture to improve resource utilization, lower single training costs, and achieve engineering experience reuse.

In January this year, Aishi released

PixVerse
R1, claiming it to be the world's first universal real-time world model supporting 1080P. Unlike traditional "pre-recorded" video generation, R1 achieves "real-time dynamic generation"—users can input new instructions at any time during video playback, and the scene can achieve natural and smooth transitions of lighting and camera within about 0.5 seconds.

爱诗科技

Aishi's approach differs from other players: while others are competing on generation quality, it is competing on interaction capability. Just one week after the release of R1, China Ruyi announced a strategic cooperation with Aishi and invested $14.2 million.

The latest C+ round of financing clearly states its use: for video generation foundation models, real-time world models, and global growth.

ZhiXiang Future: New Unicorn, Alternative Approach of "Model + Agent + Hardware"

ZhiXiang Future is a "new unicorn" on this list—after a 1.5 billion yuan C round, its post-investment valuation exceeds $1 billion. Founder Mei Tao is a foreign academician of the Canadian Academy of Engineering and former vice president of JD.com.

The company has completed three rounds of financing in the past three months, totaling over 2.1 billion yuan.

On July 23, ZhiXiang Future announced a C round of 1.5 billion yuan financing, led by the National Social Security Fund, Sichuan Industrial Revitalization Fund, and ICBC Investment, with participation from Shanghai Film New Vision Fund, Huace Film & TV, and 18 other institutions—a combination of national long-term funds, local state-owned capital, and industrial capital, which is rare in the AI video track.

ZhiXiang launched the world's first open-use video generation DiT architecture model as early as May 2024, and at the just-concluded WAIC 2026, it released the multimodal creation agent vivago R1, focusing on unlimited-length video generation and editing.

Currently, its products cover over 100 countries, serving more than 50 million users and 40,000 enterprise customers. Its full-year revenue in 2025 exceeded 100 million yuan, and Q1 2026 revenue has already surpassed last year's total. Mei Tao stated during the 2026 Chain Expo that due to the high cost of large model pre-training, the company expects to achieve monthly breakeven by 2029.

Mei Tao was also straightforward: "We don't compete with ByteDance on underlying capabilities; we focus on commercial marketing and professional film and television collaboration."

ZhiXiang's path is clear: avoid direct confrontation with big players and deeply cultivate vertical scenarios. It first built a user base with image models, then evolved toward video and world models.

FlovaAI: $80 Million Angel Round, Guo Lie's Third Departure

FlovaAI's founder Guo Lie is an unavoidable name in China's mobile internet history. In 2013, Guo Lie created FaceQ, followed by FaceU, and in 2018, the team was acquired by ByteDance for $300 million. After that, he participated in incubating Qingyan Camera,

Xingtu
, and

CapCut
—that's right,

CapCut
was something he helped create.

In 2025, Guo Lie set off again, founding Shenzhen Yuzhou Technology, and launched Flova.ai in October. Sequoia China, IDG, and Sky9 Capital have invested a total of over $80 million.

Daring to give $80 million in an angel round, investors are betting not on the product but on Guo Lie himself. Every consumer tool product he has made in the past became a phenomenal hit.

爱诗科技

Flova is an AI-native video creation Agent platform. Users express their creative needs through conversation, and the Agent completes the entire process from scriptwriting, character design, storyboarding, to image/video/audio generation, and timeline assembly for editing.

Unlike the mainstream canvas format on the market, Flova designs a "storyboard" workspace. The Agent receives instructions in the dialog box, and the work process is displayed on the left storyboard. Users view results directly in the middle preview area and can double-click to intervene and modify if unsatisfied.

Multiple early users reported that Flova's biggest feature is its "memory"—after creating over a dozen episodes, it can still remember a specific storyboard from the first episode. This long-context memory capability is very prominent among similar products.

At this stage, Flova adopts a zero-margin strategy. Guo Lie said: "During the development period, we will try to achieve zero gross profit." The team is not in a hurry to pursue short-term revenue but focuses on the effective consumption ratio: how much of the token consumption users find valuable, and how much is wasted due to product inadequacies.

Why All Crowded in July? Possible Reasons

Now back to the core question: Why did these five companies coincidentally conduct intensive financing in July? Piecing together all the information, I found four reasons forming a resonance.

First, ByteDance's Seedance "siphon effect" forces everyone to accelerate.

Since the release of ByteDance Seedance 2.0 in February this year, its monthly revenue has exceeded 1 billion yuan, with a penetration rate of about 95% in the short drama industry, and it occupies over 80% of the market share by daily token consumption.

Volcano Engine has raised its 2026 MaaS revenue target from last year's 1.5 billion to 15 billion—a tenfold increase, almost entirely relying on the Seedance model alone. Moreover, Seedance 2.1 is about to be released, with expected performance improvement of another 20%, and the low-end version's price is being pushed down to 0.5 yuan per second.

When ByteDance is using a "model capability gap + price war" dual attack on the market, other players, if they don't quickly raise money, hoard computing power, and accelerate iteration, may not even have the qualification to stay at the table.

Everyone is looking for a fulcrum to counter ByteDance in the long run. But the approaches are completely different: Kling takes a platform-based full-modal route, Shengshu takes a tech geek route, Aishi takes a real-time interaction differentiation route, Zhixiang takes a vertical scenario deep cultivation route, and FlovaAI takes an Agent workflow route.

Capital is not investing in the same story but betting on different technical routes and business models.

Second, commercialization has been proven, and capital sees an exit path.

Last year, everyone was still discussing "Can AI video make money?" This year, the data is out: Kling's ARR is nearly $500 million, Shengshu's 2025 revenue grew over 10 times, Aishi's ARR is $40 million, and Seedance exceeded 1 billion in a single month.

Advertising, e-commerce, short dramas, gaming, film and television—the paid scenarios for AI video are being validated one by one. This state of "visible cash flow" is the direct trigger for concentrated capital entry.

Third, the track has upgraded from "video generation" to "world model," and the narrative has been elevated.

In the first half of this year, the competitive landscape of foundational large models has become more concentrated. Primary market funds are beginning to seek the next direction, and multimodal generation and world models have become the consensus exit.

Note a trend: None of the four companies that received massive funding claim to be solely doing "video generation." Kling talks about an All-in-One multimodal creative workflow, Shengshu talks about "general world models" and embodied intelligence, Aishi talks about "real-time world models" and interactive entertainment, and Zhixiang talks about "native multimodal world models" and causal reasoning.

Video generation is just the entry point; the real endgame is the world model—an AI system that can understand, reason, and construct the physical world. Investors are not looking at this year's video generation market, but at who will emerge with a world model three years from now.

Fourth, the Hong Kong IPO window is open, and capital is scrambling for "tickets."

In January this year, Zhipu and

MiniMax
both saw their stock prices surge after listing on the Hong Kong Stock Exchange. In the first quarter, Hong Kong IPO fundraising exceeded HK$100 billion in 79 days, setting a record for the fastest pace, with over 350 companies in the queue.

Shengshu has completed its share reform, and IPO rumors are rife; Kling's spin-off financing with repurchase clauses is essentially a prelude to listing. The intensive financing in the primary market is largely preparing for tickets to the secondary market.

Finally, I want to say that after this July, the AI video track has entered the finals. The landscape will likely become much clearer: leading companies are holding heavy funds to rush for listings, mid-tier companies are being repriced by the world model narrative, and new players like Flova are trying to bypass the model arms race with Agents, carving out a piece of the pie from the application layer.

But the problem behind the hustle is that money comes fast, but burning it is even faster. Video generation is the most compute-intensive AI category. Over 26 billion sounds impressive, but spread across each company's annual computing bill, it may not be ample.

This article is from WeChat public account "IT Juzi" (ID: itjuzi521), author: Wu Meimei