Inference cost is the new battleground in AI, shifting focus from training to serving tokens cheaply. Learn how GPUs, batching, and distillation are driving a race to the bottom and what it means for the future of models. Chapters: 00:00 The Real AI Race 00:16 Inference Cost per Token 00:31 The New Moat 00:47 Training vs. Inference Cost Over Time 01:00 Race to the Bottom 01:14 DeepSeek Shocks the Market 01:26 AWS Joins the Fray 01:38 Hidden Economics 01:52 Three Cost Killers 02:07 Latency Reduction 02:22 Throughput Boost 02:37 Size vs. Accuracy 02:52 Hardware Shakeup 03:08 Tokens per Second 03:24 Andrew Feldman Quote 03:39 Inference Chip Market 03:55 Open-Source Cost Advantage 04:10 LeCun on Cost 04:24 Open-Source Dominance 04:40 Karpathy's Moat 04:56 Cost Plunge 05:11 The 100x Gap 05:27 The Hidden Subsidy 05:43 Caching Wins 05:59 Model Compression 06:15 Quantization in Action 06:31 The Student Outperforms 06:47 The Open-Source Edge 07:03 Tokenomics 07:19 Caching Saves 07:35 Subsidized Inference 07:51 The Final Frontier Sources & further reading: ⢠OpenAI API Pricing ā https://openai.com/api/pricing/ ⢠DeepSeek-V2 Technical Report ā https://arxiv.org/abs/2405.04434 ⢠Speculative Decoding Paper ā https://arxiv.org/abs/2211.17192 ⢠Groq LPU Benchmarks ā https://groq.com/ ⢠AWS Bedrock Pricing ā https://aws.amazon.com/bedrock/pricing/ ⢠AI Inference Chip Market Report ā https://www.marketsandmarkets.com/











