Original title: (After MiniMax’s stock price rebounds 92%: what’s left after subtracting the model company’s impact?)
Original author: Insight Beating
On August 31, MiniMax surged 16.18%. Counting from the July low, the stock has already rebounded 92%, and its market cap has returned to HK$124.9 billion.
Three months ago, the market had just held a memorial for this company. The stock price fell by more than 80% from its peak, and a market cap of HK$300 billion was subsequently wiped out.
On August 26, MiniMax released its interim report. Revenue increased 283% year over year, and B2B revenue grew 703%. By August, the company’s ARR had already exceeded $800 million, and the token consumption in July was 20 times that of January.
A few days later, what pushed MiniMax back into the discussion center was a model it hadn’t made itself.
H3 Max.
infinite channel
On August 3, MiniMax open-sourced H3, a full-modal model with 33B parameters that can handle text, images, videos, and speech at the same time. fal took its open-sourced weights, continued with post-training, optimized together with its own inference infrastructure, and finally produced H3 Max.

According to fal, generating the same 5-second video, H3 Max’s throughput can reach about 35 times MiniMax’s official endpoint.
On August 29, Rehan Sheikh, an engineer at fal, a U.S. generative AI infrastructure company, connected H3 Max to Twitch and began streaming. The content wasn’t a pre-prepared video. While the footage played, the model generated afterward. A 5-second segment, 768p, with synchronized audio; the generation time was about 3 seconds.

Faster than playback.
That’s why viewers can decide the direction of the plot in the live comment barrage; within a few seconds, the new storyline appears on screen. As long as the machines keep running, this channel can theoretically stream forever. A weekend drew around 5 million onlookers, and even people who didn’t like AI-generated content gave it a name: “infinite gruel machine.”
At the bottom, it comes from the same model. What truly widens the gap is what happens after the model is delivered: post-training, inference optimization, hardware scheduling, and the deployment approach.
A media outlet, when reporting it, wrote the headline directly as:
(Post-Train Beats Parent)
Post-training beats the parent father
At first glance, this looks like a very neat engineering optimization—yet it collides with a bigger issue behind MiniMax’s rebound this round.
If models can be open-sourced, taken away, and even further trained—and others can make them faster and better—then what exactly should the market price for a model company?
To answer this question, let’s first look at the company fal.

This company doesn’t train base models. Its business is to take in models trained by other companies, run them on its own infrastructure, and then sell them to developers. The faster new models come out, the faster its shelf gets updated.
After bringing models in, fal will redo inference optimizations, hardware scheduling, and deployment—pushing speed and cost to a position more suitable for commercial calls. More than 2.5 million developers use it to call generative models; the company’s ARR is already about $400 million, with a team of only 70 people.
Over the past year, fal’s revenue has grown by about 60 times, while its valuation has risen from $4.5 billion to around $8 billion, according to reports.
Sequoia Capital conducted a deep interview with fal’s founding team. During the interview, Sequoia partner Pat Grady mentioned that in fal’s second and third quarters of 2025, they observed that the top-five video model leaderboard’s half-life is only 30 days.
The term “half-life” comes from physics; originally it referred to the time it takes for half the atomic nuclei of a radioactive substance to decay. Used here, it means the time it takes for about half of the Top 5 video model leaderboard to be replaced by new models.
A model can reach the top of the charts and then be replaced within a month. But no matter how the model changes, developers always need a platform to run them.
fal got stuck in this position.
Developers pay for each generated image and video segment based on usage volume. After plugging into a platform like fal, whenever which model is currently best, you can switch and use it instantly—without having to re-integrate APIs or rebuild infrastructure. Models can change month to month; the platform’s position is much harder to replace.
fal CEO Burkay Gur once said, “When content generation becomes infinite, finite things become more valuable.”

This company itself was born as a business after models rapidly depreciate. H3 is made by MiniMax, but the users, revenue, and call data for H3 Max will first stay with fal. The more successful H3 Max is, the sharper this problem becomes.
And fal is still a private company—its shares can’t be bought in the public market.
From 1,100 Hong Kong dollars to 200 Hong Kong dollars
When MiniMax went public, the market believed a different story.
This company builds text, speech, and video models. For consumers, it has Talkie and Hailuo AI; for businesses, it has APIs. Compared with startups that do only a single model, it looks more like a complete multimodal commercial closed loop.
On January 9 this year, MiniMax listed with an issue price of HK$165. Public subscriptions were oversubscribed by 1,848 times; the stock rose 109% on the first day. On March 10, during trading, the stock price broke HK$1,100, and its market cap once exceeded Baidu.
That was when the market believed this story the most.
The turning point came in June. On June 1, MiniMax released its flagship model M3 with 428B parameters. Compared with GLM-5.2 and DeepSeek V4 in the same period, its model size was smaller. The official emphasis was on training and engineering efficiency: one system could train four models autonomously for 12 hours without human intervention, and accelerate CUDA kernels by 9.4x.

Then the problem came. The community found that some of the results officially published came from self-testing; there was a clear gap between third-party actual tests and the公开 numbers. The synchronized adjustment to billing rules also caused some older users’ quota consumption speed to increase substantially.
On June 5, the company publicly apologized, saying, “M3 needs more compute power; this is because our work hasn’t been up to standard.” Ten days later, after deploying only two weeks of the flagship model, it permanently cut the price by 50%. Immediately afterward, on June 16, Zhipu released GLM-5.2; the first place on the open-source leaderboard changed hands right away, and M3’s lead lasted only two weeks. On July 9, the first day after the ban was lifted, MiniMax fell 20%. A few days later, the stock price had already dropped more than 80% from its high point.
The very next day after the crash, Yan Junjie announced that he would no longer take a salary before achieving AGI, and would use his personal shares to reward the team and support open-sourcing. That same day, the company completed a share placement worth 16 billion Hong Kong dollars, raising about 7x subscriptions.
By August 26, MiniMax had produced a mid-year report pretty enough. Revenue grew 283% year over year, B-end revenue grew 703%, and R&D spending continued to rise rapidly. The next day, the stock price rose only 4%.
The market is clearly still waiting for something else.
Core asset
The biggest problem for model companies is that product updates are too fast.
In early June, M3 just took the top spot; half a month later, GLM-5.2 overtook it. Then DeepSeek V4, GLM-5.3, and Kimi K3 were released in sequence. Now, whoever is first today might change hands as soon as the next model comes out—the lead window is only a few weeks.
Leading only counts in weeks. Relying on a single model is obviously not enough to support a company’s long-term valuation.
So over the past year, the market has been searching for the model company’s true core assets.
The most direct answer is the next generation of models—but the next generation will also become outdated quickly. As long as the industry keeps iterating at high speed, “building new models” only ensures the company isn’t eliminated; it can’t explain why it’s worth holding long-term.
A more convincing answer is the ability to manufacture models.
This year, many model releases have already proven that with the same base, relying only on post-training and the environment can still produce huge leaps in capability.
智谱 GLM-5.3 and GLM-5.2 use the same 753B-parameter base model; what really changed is the larger-scale training environment and reinforcement learning post-training. Terminal-Bench 3.0 jumped from 4.6 to 28.3; the base wasn’t retrained, yet the effect improved dramatically.
Later, Tang Jie called it a “control variable experiment.”
During the training and evaluation process, Kimi K3 built more than 50 million sandbox environments; a large amount of machine time went into waiting for model inference. Freezing, restoring, and forking were compressed to the tens-of-milliseconds scale, with the goal of letting the model repeatedly try, fail, and restart in massive real environments.
Cline even let K3 modify the scaffolding that trains itself. After 17 hours without human intervention, performance improved from 77.5 to 88.8.
The part of model companies that’s truly hard to replicate starts shifting from a set of weights to a whole system of production weights, post-training weights, and evaluation weights.
Anthropic has spent more than $1 billion in a year on reinforcement learning environments. For the same reason, OpenAI bought tens of thousands of Macs for computer-use agents.
The barrier and investment to build this system have been rising rapidly, and it’s not controlled only by model companies.
Cursor’s Composer 2 is a very typical example. Its underlying layer is Kimi K2.5, open-sourced from the Moonshot/“dark side of the moon,” but the valuation the market ultimately gave to Cursor is clearly unrelated to Kimi K2.5.
Cursor has accumulated a large amount of real programming behavior, developer workflows, and reinforcement learning systems trained around those data. The base model comes from others; what’s truly critical is what remains after the model is used.
fal is the same. It takes the H3 weights, builds H3 Max. When users generate videos on fal, the call revenue first goes into fal, and the usage data also stays with fal first.
Upstream model companies provide the most expensive raw materials, while downstream may end up keeping the truly compoundable things. Technical complexity and value-capture ability don’t naturally correspond.
Thirty years ago, something similar happened in the personal computer industry. Chips were the most complex part inside the machine, but ordinary consumers recognized Dell and Compaq. Intel later spent more than a decade to reinsert itself into the consumer’s field of view through “Intel Inside.”
Subtract model from the model company
So judging a model company’s value might involve doing a subtraction. Subtract the model weights, replicable inference optimizations, and the post-training capabilities that can be taken away by third parties; what remains is what the company can truly hold long-term.
MiniMax just happened to run an experiment like this proactively. On August 3, it open-sourced H3. Within a short time, hundreds of derivative models appeared, downloads reached tens of millions, and chip platforms such as Huawei Ascend, Mobiledge/Ampere? (沐曦), and AMD quickly completed adaptation. A large number of developers and partners got onboard.
Even though the weights are open-sourced, enterprise contracts and API revenue have not leaked away—just in July before H3 open-sourcing, MiniMax’s token consumption was already 20 times that of January, and these are still recorded on its own books.
More importantly, H3 didn’t open everything. The two hosted modules H3-Context-IR and Regenerate-2 K are still held by MiniMax. If you want to output higher-spec 2K videos, you still need to call the official services. This is a very typical open-source strategy: diffuse underlying capabilities so the ecosystem can grow, while keeping the portion that can truly generate commercial feedback at the service end.

That’s why H3 Max’s significance for MiniMax isn’t just “someone optimized my model faster.” For the first time, it demonstrated the subtraction operation to the market.
The interim report provides part of the answer: B-end revenue grew 703%, and ARR surpassed $800 million. At least up to now, open-sourcing hasn’t caused all value to leak away. But H3 Max also demonstrates the other half of the risk: customer relationships, call revenue, and usage data may settle on downstream platforms.
Long versus short contest
In late August, MiniMax’s stock price rebounded from the lows to nearly double, but the short-selling ratio also rose to about 20%, setting a record at a high level.
What the long-sellers see is what remains after subtraction. For example: enterprise customers, API revenue, real usage, data reflow, and the system capability to continue training the next-generation models. In public markets, there aren’t many assets like this. Model companies you can buy are limited, so MiniMax comes with a noticeable scarcity premium.
H3 Max’s explosive popularity gave this money a very intuitive reference point. If a company like fal is worth $8 billion, and one of fal’s hottest new products is built on H3—then how much is the company behind H3 worth?
That’s the long-sellers’ view.
The short-sellers’ logic is that if weights can be taken for free, post-training can be done by fal, and inference speed can be increased by 35 times by third parties—then why should all those capabilities still be fully reflected in MiniMax’s valuation?
This also explains the abnormal stock and financial-report trend for MiniMax over the past few months: the reports only reflect how much this company makes now. What the market truly cares about is what it can retain after open-sourcing. For the first time, H3 Max put this answer on the table.
Reflow
After putting H3 out, MiniMax actually faced a very old business problem. A company can voluntarily give up a layer of scarcity—provided it knows where the money in the next layer is.
Red Hat made a similar choice.
Linux source code has never belonged to Red Hat. Anyone can download it for free, modify it, and re-release it. What Red Hat sold at the end wasn’t Linux itself, but stable versions that enterprises are willing to pay for—plus long-term maintenance, technical support, and a certification system.
The more widely open-source software spreads, the bigger the market for this kind of service becomes. In 2019, IBM spent $34 billion to acquire Red Hat—not, of course, a Linux source code that anyone can download.
Whatever is released must be able to bring value back.
This is also the real thing MiniMax needs to prove next after open-sourcing H3.
Now the market has already seen the diffusion. Models get downloaded, adapted, and even undergo second-round training—and third-party products like H3 Max have started to emerge.
Next, we need to see whether this diffusion can ultimately be reflected in MiniMax’s own revenue and assets.
It could be more enterprise customers; it could be continued growth in API consumption; it could also be the data and feedback left behind during developers’ usage, and then fed back into post-training to become part of the next-generation model.
If these things can form a loop, the more H3 gets taken away, the more entry points MiniMax gets.
If no loop is formed, the situation is completely different.
Model influence may grow bigger and the ecosystem may become more prosperous, but the most valuable customer relationships, usage data, and transaction entry points ultimately settle in other hands. By then, MiniMax will provide the underlying capabilities but capture only a small portion of the value produced by this diffusion.
So what H3 open-sourcing truly needs to be tested on isn’t download numbers, or how many companies announced compatibility. Those metrics only prove that it can travel far enough.
What needs to be watched over the next few quarters is how much money can still come back after it spreads. Red Hat has enterprise service contracts; Google has search and advertising. MiniMax also needs to find its own pipeline.
----END-
Original link
