DeepSeek released V4 Flash, a coding model priced at roughly $0.28 per unit of output versus $25 for Anthropic's Claude Opus 4.8—a 99% discount while performing comparably on complex coding benchmarks. This triggered a full price war: OpenAI cut GPT-5.6 Luna's costs 80%, Google launched efficiency-focused Gemini variants, and Meta shifted to closed-source aggressive pricing. The article frames this as commoditization of AI capability—when performance gaps narrow, buyers shop on price rather than provider loyalty, potentially creating markets for "intelligent routers" that automatically select models by cost and speed. OpenAI's bet is that volume (not margin) sustains profitability; Anthropic alone maintains premium positioning on safety claims. The practical question for practitioners: which models should you route to for which tasks when the capability floor keeps rising but pricing floor keeps dropping? The source doesn't detail V4 Flash's actual latency, token limits, or failure modes—just Arena.ai leaderboard position and the price differential.
reply