Why it matters
For enterprise teams, this shift highlights the rapid commoditisation of high-performance inference. Accessing flagship-level capabilities at a fraction of the cost allows for more ambitious scaling of agentic workflows and real-time applications without compromising on budget or speed.
Key points
- V4-Flash-0731 surpasses DeepSeek's previous flagship in nine performance benchmarks
- Pricing is 30 percent lower than the latest discounted rates for OpenAI's GPT-5.6 Luna
- The release intensifies the price war among frontier model providers for high-speed inference
- The model targets developers needing high throughput for complex automation tasks



