Video title: Silicon Valley Coordinates x Fireworks Co-founder Benny Chen: open-source models, token growth, inference optimization, and model customization
Video author: Silicon Valley Vector
Edited by: Peggy, BlockBeats


Editor’s note: Against the backdrop of open-source models accelerating toward the closed-source frontier and ongoing declines in inference prices, industry discussions are shifting from “who has the strongest model capabilities” to “who can deploy models into production at lower cost.” But when model capabilities converge and rising token consumption becomes a shared concern, a more fundamental question starts to emerge: What are enterprises truly willing to pay for—cheaper model calls, or proprietary intelligence that can reliably complete specific tasks?


Recently, (Silicon Valley Coordinate) host Cao Qingyun sat down with Chen Yufei, Fireworks AI’s co-founder, for a conversation. Fireworks sits between models and enterprise applications, mainly providing customers with open-source model inference, performance optimization, and customization services. Compared with merely discussing whether open source can catch up to closed source, Chen Yufei’s observation is closer to real workloads: where Tokens flow, why enterprises pay, and what is still missing for models to move from concept validation to production.



In this conversation, Chen Yufei breaks down “who will win between open source and closed source” into a set of deeper, structural questions: can Token growth be converted into revenue, can general capabilities replace vertical accumulation, can low-priced models pass enterprise evaluations, and how inference platforms create value between cloud vendors and application companies.


First, the scale of open-source model usage and their commercial value are diverging. In the past, catching up on capabilities and the calling price were the main metrics for judging open-source competitiveness. Today, the Fireworks platform processes roughly 400 to 500 trillion Tokens per day, and the actual usage of open-source models has expanded rapidly. But free traffic, promotional subsidies, and differences in model pricing can cause Token statistics to overestimate part of the demand. Customers may call low-priced models heavily, yet still allocate the highest budget to the closed-source models with the best results. This means that the next phase of open source is no longer just about expanding traffic, but proving that it can reach—and even surpass—frontier models on high-value tasks, and converting cost advantages into willingness to pay.


Second, general-purpose models and vertical models are beginning to evolve in different directions. In the past, when frontier models upgraded, they could directly eliminate a batch of fine-tuned models. Now, vertical applications such as law, healthcare, and programming are accumulating more granular evaluations, data, and workflows, and their optimization targets are gradually separating from those of frontier labs. General-purpose models need to raise the ceiling of capabilities; vertical models need to deliver reliable results within limited scenarios. The former can solve a broader range of problems; the latter understands better how users define “correct.” This means vertical companies’ barriers are not just having a customized model, but the ability to continuously translate industry needs into evaluation systems and repeatedly migrate as base models update.


Third, the bottleneck for enterprise AI deployment is shifting from model supply to evaluation capabilities. In the past, enterprise proof-of-concepts often relied on trial experiences and subjective judgment. Now that AI is entering production settings like call centers, legal search, and medical assistance, “looks like it works well” can no longer support procurement decisions. Enterprises must know in which tasks a model is effective, when it fails, and how much cost and quality change when switching from closed source to open source. Evaluation is therefore no longer a supporting tool—it is infrastructure connecting procurement, training, and production deployment. Whoever can define tasks, build test distributions, and continuously update standards truly controls model selection.


Fourth, the value of inference platforms is shifting from “selling cheap compute” to organizing models, hardware, and workflows. In the past, inference optimization was mostly understood as reducing the cost of each single Token. Now, caching, task splitting, model routing, and context management can directly change task completion rates. Different models don’t need to compete for the same position; they can each serve as an executor or an advisor. Fireworks’ business logic is also built on this: rather than building heavy, asset-intensive hardware, it links revenue to customers’ real-world usage by training, customization, and ongoing inference. But the main competitor on this path is not a single new cloud provider; it is large cloud vendors that can simultaneously control compute, software, and customer entry points.


Fifth, the rise of open-source models may not weaken infrastructure demand. Instead, it may compress model-layer premiums and push value further toward inference and compute. Tech giants continue to increase capital expenditures—not just to capture near-term returns on investment, but to measure the long-term risk of missing the AI cycle. However, for new cloud companies that rely on external financing, debt costs, project delays, and payback periods still impose more direct constraints. Even if models become cheaper, it doesn’t mean building and running AI systems becomes lighter at the same time.


If you compress this conversation into one judgment, it would be this: catching up in open-source models is just the starting point. The next phase of AI commercialization will revolve around evaluation, customization, inference efficiency, and customer workflows.


The following is the original text content (for easier reading and understanding, the original content has been reorganized somewhat):


TL;DR


· Open-source models are rapidly catching up to closed-source models, but Token growth does not equal revenue growth. Enterprise budgets still prioritize the models with the best results.

· Distillation can only help open-source models close the gap in the short term; long-term competitiveness still depends on independent training, evaluation data, talent, and compute.

· Upgrading general-purpose models no longer necessarily eliminates vertical models. Barriers in scenarios like law and healthcare are shifting from model capability to evaluation, data, and workflows.

· The key for enterprise AI to move from PoC to production is not adding more model choices, but building an evaluation system that can measure quality, cost, and failure boundaries.

· Inference optimization has expanded from reducing the cost of each single Token to caching, task splitting, and model routing. System design itself can enable combinations of multiple models to outperform a single frontier model.

· Customized models are better suited for vertical SaaS platforms that serve many similar customers, while a single enterprise often lacks a sufficiently broad data distribution and continuous evaluation capability.

· Fireworks’ competitiveness is not about owning GPUs; it’s about connecting training, customization, and ongoing inference. But the real long-term competitor is still the cloud giants that can cover the entire technology stack.

· Widespread open-source models may not weaken hardware demand. More likely, they will compress the model-layer premium and redistribute industry value toward compute, inference infrastructure, and enterprise workflows.


Interview highlights


Open source will keep catching up, but distillation is not a long-term answer


The Fireworks team largely comes from Meta’s PyTorch ecosystem. Based on experience with software development such as operating systems and databases in the past, the team has long believed that valuable software will eventually emerge in an open-source form that is competitive. Large models may develop along similar paths, but this round of catching up is happening faster than Chen Yufei originally expected.


On one hand, the annual recurring revenue of closed-source model companies like OpenAI and Anthropic continues to grow rapidly. On the other hand, open-source models in China and the US are also quickly closing the capability gap and pushing inference prices very low. Open source and closed source are not a zero-sum relationship where one grows while the other must decline.


Regarding the distillation of some open-source models through outputs from closed-source models, Chen Yufei believes this approach can help a model reach a higher level in the short term, but it is hard to become a reliable long-term path. Closed-source vendors can gradually patch APIs and security mechanisms, reducing the likelihood that inference processes and related data can be extracted.


But progress in open-source models doesn’t rely entirely on distillation. As long as the evaluation system keeps improving, data costs keep falling, and there is enough talent, GPUs, and engineering capacity, open-source teams can still independently train competitive models. Companies like Meta, which have compute and talent resources, have no reason to stay far behind closed-source frontier models for the long term.


This does not mean closed-source models will lose the market. Chen Yufei compares the two to Apple and Android: an open ecosystem may accommodate larger usage scale, while closed-source products that keep leading can retain high-value customers who are less sensitive to price thanks to stable user experience, brand recognition, and enterprise trust.


Enterprise procurement tends to be conservative. “No one gets fired for buying IBM” describes the psychology behind such decisions. As long as customers still believe that procuring top closed-source models is more reliable, companies like OpenAI and Anthropic can maintain certain advantages. Compared with short-term technical gaps, it may be harder to change users’ mindsets and market entry capabilities.


Chen Yufei expects that closed-source model companies’ ARR will continue to grow, but the figure may not fully reflect the economic benefits they ultimately capture. For example, some revenue may need to be shared with cloud vendors, so ARR growth doesn’t necessarily translate into the same proportion of accounting revenue and profit. Whether closed-source model vendors’ negotiating power is declining can only be judged after the related companies disclose more complete financial information.


In the long run, open source may achieve larger calling scale, while closed source continues to capture higher-margin demand. What determines the commercial boundary between the two is not only model capability, but also brand, channels, customer trust, and how revenue is allocated.


Tokens are flowing to open source, but revenue is still following results


According to Chen Yufei, the Fireworks platform currently processes roughly 400 to 500 trillion Tokens per day. Based on the public figures it cites, this scale already exceeds the enterprise API traffic for Gemini and OpenAI disclosed externally.


The main demand on the platform comes from programming, healthcare, collaborative office work, and deep research. Among these, more and more tasks that were previously not considered programming are being redefined as “programming problems.”


Office software like PowerPoint and Excel is a typical example. When models can operate these tools through code or structured tools, existing code generation, tool calling, and reinforcement learning methods can be transferred to office scenarios. Many vertical SaaS companies are also packaging industry tools as environments that models can call, and then using reinforcement learning to make models familiar with specific workflows.


Chen Yufei roughly categorizes current demand into two types: programming and deep research. Different vertical applications combine the two to create specialized agents for scenarios such as law, healthcare, and office work.


He is particularly optimistic about collaborative office products aimed at non-technical users. Previously, most industry resources went into “science-student needs” such as mathematics and programming, but demands like generating slide decks, processing documents, and producing videos for a broader range of knowledge workers may grow faster.


The development of computer-operating agents has progressed more slowly than previously expected. More than a year ago, such products were once thought capable of bypassing complex APIs and directly using computers like humans. In practice, to reduce costs, enterprises have shifted more workflows to pure text. Text models are cheaper and easier to embed into existing workflows. As virtual machine costs fall and execution environments improve, computer-operating agents may still gain more room, but the specific optimization path is still being explored.


However, Token growth does not directly translate into commercial value. Public model routing leaderboards are often affected by free tiers and promotional activities, and the same number of Tokens can correspond to completely different prices. Traffic generated by low-priced models and traffic generated by expensive frontier models do not carry the same meaning in terms of revenue.


Chen Yufei believes revenue reflects customers’ real choices better than the number of Tokens. When enterprises are not sensitive to price, budgets still flow to the models that deliver the best results. Therefore, when Fireworks helps customers customize models, its primary goal is usually not merely to achieve “good value for money,” but to reach—or even exceed—top closed-source models on specific tasks.


The programming market also shows a split between usage and revenue. Token consumption will likely keep growing, but if commonly used models keep getting cheaper, lower prices will stimulate more calls without necessarily expanding market revenue proportionally. Just like widespread electricity use doesn’t mean all profits go to power generation companies, growth in AI call volume can’t directly tell us where the industry’s value ultimately lands.


In the coming period, the usage of open-source models may continue to rise, but whether related spending grows in sync depends on whether enterprises are willing to invest resources into customization. If more and more companies can make customized open-source models challenge closed-source models on specific tasks, then traffic and revenue will further shift toward open source. If customization costs, failure rates, and organizational capabilities continue to be obstacles, closed-source models will still capture most of the high-value demand.


General-purpose models raise the capability ceiling; vertical models accumulate task moats


Every time a new frontier model is released, the market asks again: will models that enterprises have fine-tuned with their own data and funds quickly lose value?


Chen Yufei said that in 2025, it is indeed true that many customized models are being replaced by the next generation of base models, but by 2026, this situation is already clearly decreasing. In his view, this is a positive sign that the commercial value of vertical models is stabilizing.


The reason is that general-purpose models and vertical applications are optimizing in different directions. Frontier labs need to prove that models can solve more difficult problems in mathematics, science, and drug R&D. Vertical companies in areas like law and healthcare, meanwhile, focus on specific user needs, task completion rates, citation accuracy, and whether workflows are reliable.


“Doing a good job on an agent for a vertical domain is like solving the Riemann Hypothesis—these might be completely different kinds of work.”


Take the legal AI company Harvey as an example. Its advantage comes not only from which base model it uses, but also from how it decomposes legal tasks, its evaluation system, data processing, and product workflow. Healthcare is similar: different products need to establish dedicated standards around how doctors work, the sources of information, and accuracy requirements.


These evaluations and data won’t automatically become invalid when new models are released. After upgrading base models, vertical companies can transfer their existing training environments to the new models without having to understand the entire industry again. As evaluations become more granular and data cleaning pipelines become more mature, vertical models form an ability that continuously accumulates around specific tasks.


Theoretically, frontier model companies could also concentrate resources to enter the legal or healthcare markets and beat existing products in a single domain. But this approach may not align with their commercial goals. General-purpose model companies need to serve broad customers; if they pour a large amount of resources into one vertical domain, it may be hard to support the market scale and valuation they pursue.


Chen Yufei compares general-purpose models to chain restaurants that have to cater to a wide range of tastes, while vertical models are more like restaurants focused on a small number of dishes. The latter doesn’t need to solve every problem—if it is significantly better than general-purpose solutions on target tasks, it can build an independent paid space.


He also calls this kind of imbalanced but highly specialized capability “Jagged Intelligence.” A legal model may not be able to solve the Riemann Hypothesis, but it can outperform larger general-purpose models in certain kinds of legal tasks. For enterprise customers, whether work can be completed reliably is often more important than whether the model has broader capabilities.


Not every enterprise is a fit for customized models. Chen Yufei believes the best customers are vertical SaaS companies that serve many similar organizations, not a single hospital, law firm, or end enterprise.


Vertical SaaS platforms can reach a large number of customers, understand common needs across different organizations, and also build evaluation systems that cover more scenarios. A single company usually lacks sufficiently broad data distributions and the ability to run continuous evaluations. Rather than independently training models, it’s often better to codify internal processes into skills, tools, or agents, and then call external base models to complete tasks.


In customized reinforcement learning projects, GPUs are the main cost, and data and training environments are just as critical. Teams need to check whether the model is showing reward hacking—getting high scores by exploiting evaluation loopholes without actually completing the task. At the same time, they must ensure the training environment is stable and can scale up training.


The initial environment doesn’t necessarily need to reach tens of thousands of instances. Chen Yufei says that just a thousand training environments may be enough to start reinforcement learning, because during training the model repeatedly explores and generates many trajectories. If each training step runs 128 times per environment, then a thousand initial data points could generate roughly 128,000 training trajectories.


Customized models still need to be “retrained” after they go live. When new open-source base models are released, the platform can migrate the existing training and evaluation systems to the new models. As long as the evaluation standards and task environments are already fixed, migration itself may not take long. What truly generates ongoing work is customers adding new scenarios, new workflows, and more difficult tasks.


Therefore, a vertical company’s core asset is not the weights of any single generation of model, but the data, evaluations, and workflows that can be migrated repeatedly as base models update.


Enterprise AI is stuck at PoC; what it lacks isn’t models but evaluation


When enterprises choose open-source or closed-source models, the first thing they consider is not only performance and price, but trust.


Because enterprise contracts typically last one to two years, buyers need to assess whether the vendor and its ecosystem can support business needs in the long run. Open-source models may be cheaper and more flexible, but closed-source vendors still have an edge in brand communication, success stories, and market recognition.


Chen Yufei believes that the most obvious shortcoming of open-source models and their service providers right now is not technical, but market entry capability. Closed-source model companies can continuously strengthen user recognition through demonstrations of new capabilities and research results, while the open-source ecosystem lacks a unified system for market promotion and customer communication.


Deployment methods are also changing. Early on, enterprises focused more on on-prem deployment, but now more and more workloads are shifting to the cloud. As long as vendors can build trust in areas like security and virtual private cloud setups, multi-tenant cloud services are often more economical than enterprises dedicating their own GPUs.


A single enterprise may not be able to keep a batch of GPUs at high utilization. A platform can aggregate the needs of different customers to improve hardware utilization. As inference cost becomes a larger share of enterprise operating expenditures, the cost advantages of shared infrastructure become even more pronounced.


One of the biggest obstacles for enterprises to move from concept validation to production is the lack of rigorous evaluation. Even if many companies are already using AI in critical areas like call centers, the way they judge model performance is still close to judging by feel during trials—not by establishing test standards that can be run repeatedly.


Without evaluation, enterprises can neither reliably compare open-source and closed-source models nor determine whether model updates truly improve the business. Chen Yufei therefore believes that talent able to design AI evaluations is still scarce. A reliable evaluation system may help enterprises complete model replacements and save long-term spending far greater than what it would cost to build in the first place.


This is no different in essence from unit testing in traditional software companies. In the past, SaaS companies needed to run a suite of tests before delivering software. Now, vertical AI companies need to build an evaluation set, ensuring that agents meet certain standards before going into production. The industry’s focus is shifting from “writing tests” in part to “writing evaluations,” but the core remains the same: turning product quality into measurable, accumulable organizational capability.


The development of third-party data companies can also be used as an indicator for observing enterprise AI adoption. If their customers keep concentrating on a small number of frontier labs, it suggests there is still limited room for ordinary enterprises to build independent AI capabilities. If revenue sources become increasingly diversified, it may indicate that more enterprises are starting to buy data, build evaluation pipelines, and train their own models.


What enterprises truly lack is not more models, but a set of standards that can define correctness, identify failures, and support procurement decisions.


Inference optimization shifting from reducing Token costs to system orchestration


Fireworks’ first core business is inference optimization. Chen Yufei believes there is still plenty of room here, because when a model is first released, its runtime efficiency often differs significantly from what it looks like after long-term optimization.


The platform’s problem is how to shorten this optimization cycle, so that when a new model is released it is as close as possible to the best state. Besides the model runtime itself, caching, scheduling, and supporting infrastructure all affect the final cost. Many optimizations come from large amounts of detailed engineering work, which is why inference optimization remains a labor-intensive business.


Compared with first-party APIs provided by model developers, Fireworks serves many base models as well as customized models, which allows it to observe more diverse traffic patterns and workloads. Model vendors first need to reliably serve their own models to a broad user base. Third-party platforms can then optimize more finely based on the specific characteristics of each customer’s requests.


Model routing is not only for cost reduction. In a joint experiment between Fireworks and Harvey, the system assigns one model as the task executor and another as the advisor, splitting long-context tasks into several shorter sub-tasks. Because the effectiveness of large models may decline as context length increases, a reasonable division of labor can make the overall system perform better than using the strongest model alone.


Chen Yufei likens this process to CPU performance optimization. Developers need to understand the characteristics of the processor, then break the program into workloads that fit that processor. Similarly, in an AI system, base models can be seen as a type of text processor: teams first need to understand the capability boundaries of different models, then design task frameworks, and finally place the right model in the right spot.


Hardware optimization can’t be judged only by standalone metrics like GPU, HBM, CPU, and network. When models are designed at the beginning of training, they are usually already tailored to the compute-to-memory ratio of the existing hardware. Migrating from NVIDIA GPUs to AMD GPUs with similar features is relatively easy, while moving to ASICs with a different architecture may require a lot of work.


Whether a new hardware platform succeeds doesn’t only depend on whether its parameters are close to NVIDIA’s; it also depends on scale of supply, software ecosystems, and whether economic incentives motivate developers to optimize models for it. Even if the hardware configuration is similar to NVIDIA’s, if developers can’t earn enough return from using it, it is hard to form an ecosystem.


The final optimization of an inference platform is not just about the model’s runtime speed, but how well the model matches the task and the hardware.


Fireworks doesn’t own hardware, but it connects training and inference


In the face of AI cloud vendors expanding into the software layer, Fireworks has no plans to build its own hardware or purchase GPUs with heavy capital assets for the time being. Chen Yufei describes the company’s model as closer to being a “sub-landlord”: obtaining compute from infrastructure providers, while focusing on compute software, inference services, and model customization.


The difference between Fireworks and general AI cloud platforms is that most of the traffic on its platform does not come from unmodified base models, but from customized models. Chen Yufei believes that competition in base model inference will become increasingly intense, and profits are more likely to come from specialized models that deliver better results for customers.


Fireworks connects training and inference into one business. The company helps customers train models and expects inference traffic after launch to keep growing. Compared with providers that only charge for training fees or consulting fees, this revenue structure aligns the platform’s incentives more closely with customers’ long-term outcomes: customers generate more value through the model, and Fireworks earns more revenue from continuous inference.


Some reinforcement learning service providers only offer training or consulting; their revenue may not be directly correlated with the usage outcomes after a model goes live. Fireworks’ business logic is different: it improves model performance through training, and then earns returns from subsequent inference traffic. It does not sell one-off training projects; instead, it sells the ability for customer models to continuously generate usage.


The competitors that can truly cover the entire chain are not single new cloud companies, but large cloud service providers like AWS, Microsoft Azure, and Google Cloud. They have compute, and they can also build reinforcement learning services, inference platforms, and development tools.


On one hand, Fireworks competes with cloud vendors; on the other hand, it also relies on and collaborates with their infrastructure. The company’s cooperation with Azure is currently the closest: customers can use Azure credits to purchase Fireworks services, and it also works with AWS and Google Cloud.


Chen Yufei believes that Fireworks’ current inference optimization capabilities are still better than those of comparable services from cloud vendors, but single-point performance is not the most solid barrier. More importantly, it completes the entire machine learning ops cycle—from training, deployment, and evaluation to inference—providing a coherent experience for developers. Compared with new cloud companies, Fireworks covers a longer chain. Compared with cloud giants, it needs to maintain an advantage in specialization and execution speed.


Even if, in the future, AI can automatically complete core functions, inference optimization, and post-training optimization, understanding customer needs may still be difficult to fully replace. A system might solve complex math problems, but it may not be able to spontaneously understand an enterprise’s processes, constraints, and evaluation criteria. Fireworks’ long-term value still ultimately depends on whether it can translate customer needs into models, evaluations, and deployable products.


Open source compresses model-layer premiums; compute remains the ultimate constraint


In the next one to two years, Fireworks’ biggest challenge is still compute. Whether the company can obtain enough computing resources at a reasonable price will directly limit revenue growth. Chen Yufei summarizes Fireworks’ current state as being “constrained by compute,” and expects that at least over the next year the market overall may still maintain this status.


This also affects his view of tech giants’ AI capital expenditures. The market usually judges whether companies like Google and Meta are overinvesting based on ROI, but for these companies, another risk may be bigger: if AI ultimately delivers extremely high returns and they miss this competitive cycle because they reduced investment, they may permanently lose their seat at the table.


Therefore, even if there is over-investment in the short term, the cost might only be pressure on financial statements—and the need to digest depreciation over the next several years. For tech giants with long-term cash flows and financing capability, this is usually not a survival risk; missing a key technology cycle, however, might be.


The situation for new cloud companies is different. They need to evaluate returns on investment more rigorously, including financing costs and the payback period. Chen Yufei suggests watching relevant companies’ bond ratings, their abilities to extend projects, and their capabilities for refinancing. If in the next 3 to 6 months there are participants that cannot keep borrowing due to credit issues, debt restructuring and whether the market still has funding willing to take over will become a key checkpoint for testing AI infrastructure financing capabilities.


Regarding the market narrative that “open-source models catching up will reduce hardware demand,” Chen Yufei is skeptical. In his view, the first thing open-source capability improvements weaken is the premium at the model layer. If overall AI spending doesn’t change, value may instead shift toward companies that provide compute and infrastructure.


However, whether AI capital expenditures will ultimately yield enough returns remains an unanswered question. Compute projects often take years to recoup investment, and the supply-and-demand landscape after a year is hard to predict. The clearest signal at this stage is that whether it’s Fireworks or many of its customers, they are still looking for more compute.


Competition between open source and closed source won’t be determined solely by model leaderboards. As the capability gap continues to narrow, industry value will depend more on who can define tasks, validate results, and deploy models into production with sustainable costs.


[Video link]