YouTube1 month ago

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Lin Qiao, founder of Fireworks AI, argues that the future of artificial intelligence will be dominated by specialized, customized models rather than a single general-purpose model controlling all intelligence. Qiao predicts that token costs will fall 10x over the next three years, driving 100x growth in usage, while maintaining that open-source models are increasingly viable for enterprise applications despite competition from frontier labs.

The Specialization Thesis

Qiao fundamentally rejects the notion of AGI—a single superintelligent model solving all problems optimally. He argues that majority of enterprise data remains private and locked within companies, never used to train public models, making specialized intelligence built on proprietary data superior to general-purpose solutions. He contends that every company is built on unique beliefs and operational principles, meaning true intelligence must reflect each organization's distinct needs, values, and customer understanding. This cannot be replicated by external companies building generic solutions.

Open Source vs. Frontier Models

While acknowledging the value of foundation models from companies like OpenAI and Anthropic as "power lines" distributing intelligence infrastructure, Qiao sees open-source models crossing a critical quality threshold where they solve most enterprise problems at 90% the capability of frontier models while being significantly cheaper. Open models also grant full control over weights, enabling customization—a fundamental advantage over closed APIs. Fireworks currently processes over 40 trillion tokens daily, with the majority coming from customized models rather than off-the-shelf variants. Within a year, this could reach 20-100x growth.

Cost and Usage Dynamics

Qiao attributes high current token costs to supply chain constraints in chip manufacturing rather than inherent model economics. As competition increases and infrastructure scales, costs will compress dramatically. The 10x cost reduction will unlock a 100x usage explosion because organizations will transition from cost-conscious experimentation to routine deployment. However, cost optimization happens at multiple layers: using fewer tokens per task, customizing models to be more precise, and improving infrastructure efficiency through specialized inference deployment.

Enterprise Reality Check

Qiao acknowledges that product-market fit and durable business models have decoupled for AI startups. Many companies have customer demand but cannot scale without bankrupting themselves due to token costs. Large incumbents with existing traffic face the same problem—they cannot afford to roll out AI features to their entire customer base. This has forced enterprises to choose between renting frontier models or owning their intelligence through open models, customization, and deployment control. Both Harvey and Lagora in legal AI, for example, benefit from building proprietary knowledge into specialized systems rather than relying solely on general models.

The Multi-Model Future

Qiao envisions millions of specialized models—one per application or use case—rather than a few dominant AGI systems. This requires intelligent routing mechanisms, specialized tuning, and customized deployment optimization. Companies will increasingly implement hierarchical approaches, routing simple tasks to cheaper open models and complex problems to more capable (but expensive) frontier models. The future isn't a single power line but a distributed ecosystem where companies own customized layers enabling them to optimize cost, quality, and speed for their specific workloads.

Business Model and Stack Position

Fireworks operates in the specialized intelligence platform layer, between chip providers and application companies. The company focuses on helping enterprises customize models, optimize inference, and deploy efficiently. Gross margins currently sit at 30-40% rather than traditional SaaS levels (80%+) because Fireworks prioritizes growth velocity and experimentation during hyperscaling. The company explicitly avoids moving upmarket into applications or downmarket into data centers unless timing and business fundamentals clearly warrant it. Currently, the focus remains on product innovation and market expansion.

Infrastructure and Bottlenecks

While Jensen Huang's "five-layer AI cake" (applications, models, infrastructure, chips, energy) frames the ecosystem, Qiao argues the true constraint is supply chain—specifically manufacturing capacity for chips and the physical infrastructure bottlenecks in data center construction and operation. Building data centers or custom chips makes sense only when workloads stabilize and usage reaches sufficient scale. Currently, AI workloads remain too dynamic and experimental to justify hardware-level specialization. Innovation in system design for very large models (10+ trillion parameters) and heterogeneous chip deployment (combining GPU and SRAM-based accelerators) remains inadequately addressed.

National Sovereignty and Open Source

Concerns about Chinese open-source models dominating the landscape are overstated because the open ecosystem attracts multiple participants rather than relying on a single provider. If China restricts access, the US can develop its own open models. The bigger risk is over-dependence on any single closed model provider. Companies should implement guardrails around all models—open or closed—because providers inevitably embed their own judgments and design principles, which may not match enterprise needs.

Long-Term Vision

Qiao believes every company will eventually own their own intelligence as a necessity, not an option—mirroring how every software company owns its stack. This requires not just choosing between build versus buy, but understanding when each makes sense. As specialized models proliferate and costs fall, enterprises will increasingly optimize for ROI rather than token efficiency, shifting focus to measurable business impact rather than raw usage metrics.

You just got that from one episode.

Point stillscase at the podcasts, videos, and articles you follow and get the same thing: a summary you can search, question, and hand to your agent in any chat. Free to start.

Try it free