Featuring insights from AIOZ Network & YOM
2026 has finally exposed the main problem of the artificial intelligence industry: it comes down not to hardware. Demand for compute for training and inference of models is growing faster than the industry manages to build data centers.
GPU queues at major cloud providers stretch for months, and compute rental prices remain prohibitive for small teams and independent developers.
AI's Infrastructure Deadlock
Practically the entire AI infrastructure market today is concentrated in the hands of a few players: AWS, Google Cloud, Microsoft Azure, and chip makers like Nvidia.
This creates a classic bottleneck scenario: access to computers becomes a privilege, and startups are forced to budget for the fight for a place in the GPU rental queue on top of everything else.
Centralization brings systemic risks with it: a single point of failure, dependence on a specific provider's policy, and regional or access restrictions.
For an industry that positions itself as the foundation of the future economy, this kind of dependence looks like a contradiction of its own principles of openness and decentralization.
The Next Step in Network Evolution
The industry has already walked this path before. Centralized file-sharing servers once gave way to peer-to-peer networks, and later to distributed CDNs, which learned to deliver content faster and cheaper precisely by using many scattered nodes instead of a single data center.
The same principle is now being applied to compute: DePIN proposes gathering computing power from thousands of independent devices around the world.
The logic of DePIN is simple and elegant in its own way. Millions of people have idle computing resources sitting unused: free disk space, unused bandwidth, idle GPU and CPU cycles. A DePIN network combines these scattered resources into a single infrastructure and rewards participants with tokens for the amount of storage, delivered traffic, or compute they provide.
Instead of building a data center and recouping capital costs on equipment and electricity, a DePIN protocol relies on hardware that already exists, distributed around the world.

This lowers the barrier to entry for node operators, and for end consumers of compute it potentially lowers the cost of access thanks to competition among many independent power providers within a single network.
The key question for any DePIN network is whether it can be turned into a working economy: are there enough nodes, is demand for their capacity stable, and can the infrastructure handle the load of real, rather than test, usage scenarios.
When Milliseconds Matter
Another trend of 2026 reinforces the logic of DePIN - the shift from chatbots to autonomous AI agents that act in real time, seeing, hearing, and responding by voice, reacting to video streams. For such scenarios, latency stops being a secondary parameter and becomes a critical characteristic of the product.
The difference between an agent that is pleasant to talk to and one whose pauses are annoying is measured in hundreds of milliseconds.
Classic architecture works against the user here. A request travels from the device to a data center, often on another continent, gets processed, and comes back. Every such round trip adds latency, which accumulates during audio and video streaming. Centralized data centers physically cannot be located next to every user on the planet at the same time.
A distributed network of nodes handles this task best. Instead of one powerful but distant center, there are many points of presence, geographically spread out right where the users are. A request gets processed by the nearest available node, not the most distant free server.
It is precisely at the intersection of the shortage of centralized AI compute and the growing need for low-latency AI infrastructure at the very edge of the network that DePIN protocols get a chance to occupy a niche that hyperscalers structurally cannot close on their own.
AIOZ: From Storage to AI Compute Network
AIOZ Network is one of the few DePIN projects that has traveled the path from a narrow use case to a full, multi-layer infrastructure.
The network started with the task of decentralized delivery and storage of media content (StreamP2P and S3-compatible storage), then expanded into a full AI Compute Economy, a marketplace for compute power and AI models built on top of a distributed network of nodes.

The foundation of the network is its DePIN nodes, run by thousands of participants around the world, using the idle resources of ordinary computers: hard drives for storage, bandwidth for data delivery, and GPU/CPU for media transcoding and AI compute.
According to the network's data, more than 327,000 DePIN nodes are currently registered, and the total volume of data delivered exceeds 26,000 TB. Participants are rewarded in $AIOZ tokens in proportion to their actual contribution.
By 2026, the AIOZ network had built several interconnected product lines on top of the same base of DePIN nodes:
AICompute Economy: a marketplace for compute power and AI models available to run through the distributed network of nodes;
StreamP2P: a decentralized content delivery network (CDN), an alternative to centralized streaming providers;
Storage: S3-compatible distributed data storage;
Pin: an IPFS pinning service that ensures data remains permanently available in decentralized storage.
Practical proof that AIOZ's AI direction is a working infrastructure is the deployment of the Qwen2.5-Omni model on the AIOZ AI platform.
It is a multimodal model with a Thinker-Talker architecture, capable of processing text, images, audio, and video in real time, and responding with both text and synthesized speech.
Models like this are demanding on compute resources, and the fact that it is available to run through the distributed AIOZ network is a concrete example of how DePIN infrastructure closes the gap between independent developers and heavy AI models.
For the network, this places AIOZ within the broader narrative of AI infrastructure decentralization: users and developers get access to compute and models through a distributed network of nodes.
Insights from AIOZ Network
We spoke with the AIOZ team about one of the more demanding real-world tests of their DePIN infrastructure. Here’s what they told us about adapting such a model to a decentralized network.
Q: What were the biggest engineering challenges in adapting Qwen2.5 Omni to run efficiently on AIOZ’s decentralized DePIN infrastructure, and what lessons did your team learn?
The biggest challenge was turning the scale and diversity of AIOZ DePIN into a reliable environment for real time multimodal inference.
Qwen2.5 Omni requires more than raw GPU power. Its latency depends on media processing, model loading, queueing, inference, and streamed output. AIOZ DePIN provides a broad pool of distributed compute, storage, and bandwidth resources, but those resources vary in hardware profile, location, network quality, and current availability. The main engineering task was therefore to match each request with the right resources while keeping latency predictable.
This required hardware aware scheduling across qualified compute pools, intelligent routing based on location and network conditions, warm model caching to reduce startup time, and session affinity to preserve multimodal context during streaming. AIOZ DePIN’s distributed architecture also provides flexibility to route different sessions across available machines, expand capacity as demand grows, and reduce reliance on a single data center.
The key lesson was that AIOZ DePIN’s strength is not only the amount of compute available. It is the ability to coordinate compute, storage, and bandwidth as one infrastructure layer. For real time AI, that coordination must prioritize model readiness, locality, stable connections, and tail latency rather than throughput alone.
Ultimately, the goal is to make the complexity of the underlying DePIN invisible. Users should experience one responsive multimodal service, supported by AIOZ DePIN’s distributed capacity, intelligent resource allocation, and ability to scale workloads across the network.
As the AIOZ team points out, running a model like Qwen2.5-Omni takes coordinating compute, storage, and bandwidth as a single layer, where latency matters more than throughput.
Insights from YOM
We also put a question to the YOM team, the decentralized edge GPU network best known for streaming AAA games straight to the browser. With YOM’s roadmap pointing toward real-time AI next, we wanted to know how gaming and AI workloads would actually share the same infrastructure.
Q: With gaming and AI competing for the same GPU nodes down the line, how will YOM prioritize workloads across the network?
The framing assumes workloads fight over nodes. Our architecture is built so they don’t. Our patent-pending Vulkan kernel partitions a single consumer GPU into up to 7 isolated slots, each with its own VRAM budget, so sessions never step on each other.
Our HyperOrch AI scheduler then places every session by what it actually needs: a cloud gaming session needs a node within about 50km to hit sub-10ms latency, while many AI inference jobs care more about throughput than distance. Placement by requirement, not a queue.
Supply on YOM also follows demand rather than capping it. Node licenses unlock in batches as utilization crosses thresholds, so sustained demand raises operator yield and pulls new nodes into the regions where the work is.
Where community coverage hasn’t caught up yet, telco-hosted GPUs and hyperscaler burst capacity backstop the network. Nothing gets starved, and the protocol earns a fee on every session either way. More workload classes make the network stronger, not more congested.
Q: Does real-time AI run on the same infrastructure as gaming, or does it need a separate technical layer?
Same infrastructure, by design. We started with gaming because it’s the hardest real-time workload there is: sub-10ms median latency, AAA fidelity, thousands of concurrent stateful sessions. A network that streams Unreal Engine 5 to a browser can serve any workload with the same latency and session-state constraints, and real-time AI is exactly that class.
The primitives carry over one for one. The same scheduler that places a game session places an inference session. The same Vulkan VRAM partitioning gives an AI workload an isolated slot on a consumer GPU. The same settlement layer pays the operator, and the protocol takes its fee regardless of what’s running. What changes is the runtime inside the container: a model server instead of a game build. That’s a thin layer on top, not a second network.
DeepSeek’s published work on DualPipe and DeepEP proved the hyperscaler interconnect moat can be broken with algorithms, not hardware. Our Vulkan kernel is the retail-GPU equivalent: it lets consumer GPUs take on workloads that used to require enterprise-grade isolation. Gaming proves the physics. AI inherits the network.
What comes through here is that YOM isn’t bolting AI onto its gaming network. The same infrastructure built for gaming’s worst-case latency is what makes it ready for real-time AI too.
Note from Cicada CEO Maxim Moris
We closed the loop by asking Cicada’s CEO and co-founder, Maxim Moris, why a market maker is paying attention to DePIN, and what actually gets priced in when backing early-stage projects.
Q: What in the macro picture brings DePIN onto Cicada’s radar?
Two facts, not the narrative. The GPU market has split: spot H100 hours fell from $8–10 in 2024 to under $3.50, while one-year reserved capacity climbed back about 40% off its late-2025 low. Cheap on demand, expensive when you need it guaranteed — that gap is where distributed capacity gets its shot.
DePIN finally has revenue: roughly $150M on-chain in January 2026. Small, but it means tokens with a cash flow you can model instead of a story you can pump. We follow demand for liquidity, and liquidity follows revenue.
Q: Will DePIN outlast the AI hype, or fade with it?
Right now all the attention is on AI, automation and agents. But that hype can fade fast, because most of those directions simply don’t need a blockchain. The bulk of projects today are trading with AI agents.
Q: What’s different about market-making for DePIN tokens?
What’s different in DePIN is the supply side. Node operators are paid in the token and, by the design of the system, sell every day. It is a permanent flow you have to build the orderbook against.
Maxim’s answers point to the same thing from a different seat: DePIN is interesting because there’s now revenue and a predictable sell-side to model.
Conclusion
DePIN infrastructure is unlikely to replace hyperscale data centers overnight. Training the largest models still requires concentrated compute clusters.
But for a wide range of tasks, a model where compute and storage are provided by thousands of independent network participants looks like a working alternative to the single point of failure represented by a centralized cloud.
Projects like AIOZ Network show that DePIN's transition from concept to production is already underway: infrastructure for data storage and delivery, built up over years, has become the foundation for the stage of distributed AI compute.
This is a case where the decentralization narrative is backed by network metrics and concrete, working integrations.
Read more: DePIN: Filling the Cloud Gaps
