Whether the artificial intelligence industry is building a foundation for a construction bubble whose costs are difficult to recoup remains one of the most fiercely debated questions on Wall Street. However, in the latest episode of the Next Big Thing podcast, Dylan Patel, founder of the semiconductor and AI infrastructure research firm SemiAnalysis, offered a completely different observation: Anthropic not only continues to grow its revenue at a rapid pace, but may have officially entered a profitable stage as early as the second quarter of 2026.

Patel said in an interview that Anthropic achieved profitability and positive free cash flow in the second quarter. Based on the information he has, Anthropic was profitable in April and May, when the accounting had been closed, and its free cash flow was positive. At the time the interview was recorded, although June’s accounts had not been fully settled, the operating trends were still heading in the same direction.

He further claims that Anthropic’s annualized recurring revenue—its ARR—has surpassed $50 billion, and its gross margin is above 70%. In May 2026, Anthropic said its annualized revenue has crossed $47 billion. If the growth trend continues, the $50 billion threshold Patel refers to is not unreasonable.

Claude Code becomes the revenue engine for Anthropic

The core driver behind Anthropic’s rapid revenue growth comes from enterprise customers and the programming/developer tool Claude Code.

In February 2026, Anthropic said Claude Code’s annualized revenue has exceeded $2.5 billion, up more than double from the beginning of the year; adoption speed among enterprise customers continues to accelerate.

Patel believes Anthropic can lead OpenAI in the enterprise market. The key is not necessarily that Claude is smarter on every benchmark test, but that Claude has higher “token efficiency” in real-world workflows.

He points out that sometimes OpenAI’s models can complete tasks that Anthropic’s models can’t, especially in cutting-edge science, math, or hard programming problems. But some work may require triple the time and generate four times the tokens. For enterprise tasks that require humans to continuously check, modify, and provide feedback, both waiting time and token consumption directly affect costs and the user experience.

By contrast, Claude Code often completes work with fewer tokens and in less time, letting users move on to the next round of revisions faster. SemiAnalysis therefore still uses Anthropic models as its main tool for now; OpenAI Codex is more often used for long tasks that can run overnight without frequent human intervention.

When companies choose models, they can’t just compare the list price per million tokens. The true costs are: how many tokens the model needs to complete the entire task, how much time it takes, and how many times human intervention is required.

SemiAnalysis has 90 employees and spends $11 million per year using AI

Patel also uses SemiAnalysis as an example to show that companies’ AI spending is growing at an astonishing rate.

In November 2025, before Claude Code became widely popular, SemiAnalysis’s recurring AI spending per year was often under $100,000. At the time, the company mainly bought a $200-per-month ChatGPT plan for each employee; if employees had needs, it also provided Claude or xAI services.

But as Claude Code quickly became widespread after December 2025, SemiAnalysis’s annualized AI spending rose to $4 million by the end of January 2026, and is now about $11 million. If the highest-spending week is multiplied by 52, it corresponds to $14 million per year.

For a company of only about 90 people, this means each employee could consume more than $120,000 worth of AI resources per year on average.

Patel admits that this is an extremely large expense, but the company hasn’t reduced its AI budget as a result. The reason is that AI enables the team to develop more products, increase sales revenue, and improve employee productivity. For SemiAnalysis, AI spending has already reached about one-third of personnel costs; he estimates that by year-end it could even approach half of personnel costs.

AI budgets burn through early, and companies start cutting other software

This kind of change is also starting to appear at other companies.

Patel said that many companies initially budgeted for an entire year of AI, only to use it all up in the first or second quarter. As a result, businesses face two options: limit employees’ use of AI, or cut other spending so more resources flow to AI.

He has observed that some companies are starting to cancel traditional SaaS products because employees can use AI to build internal tools that replace existing software. Other companies choose to reduce hiring—and even cut labor costs—rather than limit AI usage.

Patel warns that companies that choose to suppress AI usage might be able to control costs in the short term, but they could gradually fall behind in productivity and product development speed.

In his observations, AI costs are not simply going down or going up—they follow a loop of “improved model efficiency, followed by an expansion in usage scope.” When new models are released, the number of tokens needed to complete the same tasks drops, so company spending may temporarily decline. But employees will quickly use the newly gained capability to take on more work, and total spending rises again.

The cheapest model isn’t necessarily the model with the lowest cost

Patel breaks enterprise AI workloads into two categories.

The first category is automation work already integrated into fixed workflows. For example, once a company receives customer documents, it hands them to a model to check specific fields or perform classification. As long as model quality meets requirements, the company can gradually switch to cheaper models while maintaining accuracy, thereby lowering inference costs.

Patel estimates that, at the same level of capability, the cost-effectiveness of AI models could improve by about 60x per year. He cites DeepSeek V3 and GPT-4 as examples, saying they are about two years apart, but the inference cost gap for similar-capability reasoning can at one point be hundreds of times.

The second category is assistant-style AI work, such as programming, research, data analysis, and complex decision-making. The best way to optimize costs for these tasks is often not to switch to weaker, cheaper models, but to use the newest, smartest models with the highest token efficiency.

An older model may need 100,000 tokens and multiple rounds of back-and-forth to complete a task, while a newer model may need only 25,000 tokens and a single instruction. Even if the newer model’s per-token price is higher, the overall cost and employees’ time cost may still be lower.

Therefore, Patel believes that for knowledge workers, the “smartest model” is often also the one with the lowest practical cost.

Reasoning models drive up memory demand; the shortage could last for several years

In addition to model companies’ revenue, Patel also has a strongly bullish view of the AI hardware supply chain—especially the memory market.

In the past, the memory industry was usually seen as going through a business cycle every 18 to 24 months: insufficient supply pushes prices up, manufacturers expand production, and then excess capacity leads to price declines. But Patel believes that generative AI has introduced a structural change to this memory upcycle.

The key turning point comes from reasoning models and AI agents.

The input and output of traditional chatbots are usually short, but reasoning models need to preserve a large amount of intermediate reasoning steps. AI agents may read documents for long periods, call tools, run programs, and maintain context. This material is stored in KV Cache, causing the memory capacity required for inference to increase rapidly.

Patel says that when context length increases from 1,000 tokens to 100,000 tokens, the model weights themselves may not change, but the amount of data that the KV Cache needs to read and store grows dramatically. Therefore, even if compute doesn’t rise proportionally, memory demand can still jump sharply.

SemiAnalysis estimates that over the next three years, global memory production capacity will only increase by about 20% to 30% per year, but demand driven by AI could be close to doubling. If supply can’t catch up quickly, memory prices may keep rising until smart phones, laptops, and other consumer electronics are forced to reduce usage, freeing more capacity for AI data centers.

Patel even predicts that in the future iPhones and MacBooks may rise in price by several hundred dollars due to higher memory costs, not just a symbolic increase of $100.

However, he does not deny that the memory market will eventually see a downturn cycle. When supply catches up with demand, the gross margins that memory makers are expanding rapidly could still drop sharply. The difference is that this time the supply-demand imbalance may not be repaired within a few quarters, but could persist as a structural shortage lasting for years.

AI agents turn CPUs from a side role into a new bottleneck

Another change in AI investment is that CPUs are back in focus.

In the first stage of generative AI, models mainly rely on GPUs for model training and short-context inference, so the importance of CPUs is relatively low. But when model training shifts from pre-training to reinforcement learning, and inference shifts from chat to AI agents, CPU demand also starts to surge.

Reinforcement learning requires the model to verify answers across a variety of environments—for example, running program unit tests, compiling code, operating simulated websites, or checking engineering systems. These tasks rely heavily on CPUs rather than GPUs.

AI agents also keep searching the web, querying databases, calling Python, compiling code, and deploying services during real work. GPUs handle model inference, but the process of models interacting with the external world often requires CPU processing.

Patel points out that OpenAI and Anthropic have begun leasing large amounts of existing CPU capacity from Amazon, Google, and Microsoft. Products such as Intel, AMD, Arm, Amazon Graviton, and NVIDIA Vera have therefore benefited.

Still, he also reminds that CPU demand today includes part of a “channel inventory replenishment effect.” Over the past three years, data centers deployed large numbers of AI accelerators without allocating enough CPUs. With the rise of AI agents, cloud providers are now making up for past shortfalls. Once this backlog is cleared, CPU growth rates could return to a more stable level.

Even if the ratio of CPU to GPU quantities is similar, their market value can differ enormously. A high-end AI accelerator may be worth around $50,000, while a CPU is about $5,000. Therefore, most of the capital expenditures for AI data centers will still go to GPUs, AI ASICs, and memory—not CPUs.

Large-scale commercial use of CPO may not arrive until 2029, and the lifetime of copper cables may be longer than expected

For data center networking, Patel also believes demand will continue growing. As model sizes and AI clusters expand, the share of network equipment in an AI system’s total cost could rise from under 10% to over 10%. Entering the era of co-packaged optical components—that is, the CPO era—this could even reach 20% to 30%.

But the market may be overestimating the rollout speed of CPO.

Patel believes that CPO is unlikely to be fully deployed across the board in 2027, and that truly large-scale rollout may have to wait until the end of 2028 to 2029. The main reason is not that the technical concept is infeasible, but that manufacturing yields, packaging costs, capacity, and chip design are still not mature.

For the next several generations of GPUs, NVIDIA will still use copper interconnects as much as possible. Vendors will switch to more expensive optical solutions only when the telecom signal can’t support the required distance and bandwidth.

This also implies that the lifetimes of traditional optical transceiver modules, active cables, backplane connectors, and high-speed copper cables may be longer than the market originally expected. Patel says SemiAnalysis previously also expected CPO to become mainstream earlier, but after observing progress in downstream chips, it has pushed the timing back.

Data centers add 50% more each year; power becomes the biggest constraint on AI expansion

AI infrastructure ultimately still has to come back to the problem of electricity.

Patel estimates that in 2026 the world will add about 20GW of data center capacity, 30GW in 2027, and 50GW in 2028. With such rapid expansion, electricity generation, transmission, power conversion, construction work, and political approvals all become bottlenecks at the same time.

The hardest part to solve is power transmission. Large transmission lines involve local governments, regulators, utilities, and residents’ interests, and the approval and construction timeline could take as long as several years.

That’s why an increasing number of data centers are adopting “generate after the fact,” building power plants directly on-site within the parks. Patel predicts that in a few years, as much as half of the power demand for newly added data centers may come from on-site generation.

In addition to large-scale combined-cycle natural gas power plants, operators are also starting to use reciprocating engines, industrial gas turbines, diesel engines, and even convert train, ship, or truck engines into power-generation equipment.

These options might not be the most environmentally friendly, the cheapest, or the easiest to maintain, but when the power grid can’t supply electricity on time, companies may be willing to use any viable method to get GPUs online first.

Patel describes the AI industry as forming an energy-solution spectrum stretching from “converting a truck engine into an on-site generator” all the way to “sending data centers into space.”

This article: Anthropic turns from loss to profit in Q2, with ARR surpassing $50 billion! SemiAnalysis founder: the AI bottleneck is electricity—first appeared in Chain News ABMedia.