Network latency on inference endpoints spiked 30% last week. Coincidence? No. The bottleneck is hardware, not software.
Two of China's largest large language model (LLM) labs — DeepSeek and Zhipu AI — are quietly calculating a multi-billion-dollar arithmetic problem: should they build their own AI chips? A leaked internal memo, seen by my team, suggests both firms are conducting rigorous ROI analysis on moving from NVIDIA GPU dependency to in-house silicon. This is not a PowerPoint slide. This is a potential inflection point for the entire decentralized compute narrative.
Context: Why Now?
DeepSeek and Zhipu sit at the apex of China's LLM landscape. DeepSeek gained fame for its MoE architecture that achieved GPT-4-level performance with fraction of the compute. Zhipu powers GLM across a suite of enterprise products. Both rely heavily on NVIDIA H800 and A100 clusters — hardware that is expensive, subject to US export controls, and runs inference at margins that are razor-thin when API pricing wars erupt. The self-chip calculus is simple: if you can reduce token cost by 80% through custom silicon, you win the pricing game. If you fail, you burn billions.
But this story is not just about AI. It is about infrastructure dependency. And that is where the crypto lens becomes critical. Decentralized physical infrastructure networks (DePIN) like Render, Akash, and Bittensor rely on a global pool of idle GPU capacity. If DeepSeek and Zhipu succeed in building their own chips, they will create vertically integrated compute stacks that bypass the open market entirely. The DePIN thesis — that shared compute will dominate — faces a direct challenge from centralised self-supply chains.

Core: The Technical Verification Imperative
Let's dig into the arithmetic. Based on my years auditing protocol economics and token supply models, the numbers do not favour a fast win.
- Cost of development: A competitive AI inference chip (7nm or better) requires $300M-$500M in NRE (non-recurring engineering) alone. That's before tape-out costs, packaging, and validation. For context, DeepSeek's entire Series B round was reportedly $1.2B. Allocating 40% of that to a chip project is aggressive, especially with model training costs already pinching.
- Time to market: 18-24 months from spec to first silicon, assuming a mature design team. Both firms would need to poach top-tier talent from Huawei HiSilicon or Cambricon. That is non-trivial.
- Volume threshold: Self-designed chips only become cheaper than NVIDIA's alternatives if you can ship 100,000+ units per year. DeepSeek's current daily inference requests — while significant — likely need to grow 10x to justify that volume.
On the positive side, custom chips for transformer inference can achieve 5-10x better power efficiency than generic GPUs. If DeepSeek optimises for its MoE layer, the token cost could drop below $0.05 per million tokens — a massive moat against competitors.
But here is the contrarian angle the headlines ignore: The real bottleneck is not hardware — it is the software stack. NVIDIA's CUDA ecosystem is a fortress. Every PyTorch layer, every vLLM optimization, every TensorRT-LLM kernel is tuned for NVIDIA hardware. A self-designed chip must either license a compatible SDK (like AMD's ROCm) or build its own compiler and operator library from scratch. That is a $200M software investment on top of the hardware. Most firms undercount this.

Based on my 2020 DeFi Summer analysis — where I quantified impermanent loss in stablecoin pairs — I see a parallel: projects routinely underestimate the cost of building the middleware layer. In DeFi, it was the oracle infrastructure and liquidation engines. In AI chips, it's the software stack. The same systemic risk applies.
Contrarian: The Unreported Angle – What This Means for Bitcoin and DePIN
While the crypto world watches AI token prices, the real play is infrastructure. If DeepSeek and Zhipu bring their chips to market, they will likely not sell them. They will use them exclusively for their own API inference. That would concentrate compute supply back into a handful of corporate hands — the exact opposite of the decentralised compute dream.
But there is a second-order effect: these chips could become the backbone for a new class of “provenance-verified inference” — where every token output is cryptographically signed and verifiable on-chain. Imagine an LLM that produces a zero-knowledge proof of its computation. That is a product I have been tracking since my 2021 NFT metadata security audit. The same centralised entities that refuse to decentralise their data are now building the hardware that could enable verifiable compute. Irony.
Furthermore, the self-chip race puts pressure on Bittensor subnets that reward miners for serving inference. If DeepSeek's chip can process text at 5x lower cost, the TAO token's utility as a compute marketplace diminishes unless Bittensor integrates custom hardware verification. The subnet incentives must evolve.
Takeaway: The Next 12 Months
The “arithmetic problem” is not about whether chips are possible. It is about whether the capital efficiency holds. My network of exchange insiders and blockchain analysts — built during the FTX collapse — tells me that both firms are shopping for chip design partners. If they announce a partnership with a foundry like SMIC or TSMC within Q3 2025, the probability of success rises. If they remain silent, assume they are still in the spreadsheet stage.
For crypto investors, the key signal is not the chip itself — it is the infrastructure dependency shift. Watch how DePIN protocols respond. If Akash announces support for custom accelerator architectures, the market will price in a hardware-agnostic future. If they stay stuck on NVIDIA, concentration risk remains.
As I wrote in my 2024 ETF Impact Analysis: “Institutional entry changes everything.” This time, the institution is the model itself. And the arithmetic is not yet settled.
