DeepSeek launches V4-Flash-0731 to challenge OpenAI on performance and cost

DeepSeek has launched V4-Flash-0731, a new model that outperforms its own flagship across nine key benchmarks. This release significantly undercuts the market by offering pricing 30 percent lower than OpenAI's recently discounted GPT-5.6 Luna model.

For enterprise teams, this shift highlights the rapid commoditisation of high-performance inference. Accessing flagship-level capabilities at a fraction of the cost allows for more ambitious scaling of agentic workflows and real-time applications without compromising on budget or speed.

  • V4-Flash-0731 surpasses DeepSeek's previous flagship in nine performance benchmarks
  • Pricing is 30 percent lower than the latest discounted rates for OpenAI's GPT-5.6 Luna
  • The release intensifies the price war among frontier model providers for high-speed inference
  • The model targets developers needing high throughput for complex automation tasks
Generative AI AI Apps & Platforms AI Agents & Automation
All AI news

More AI news

Models

OpenAI disbands preparedness team responsible for assessing catastrophic ai risks

OpenAI has reportedly dissolved its preparedness team, the group tasked with evaluating and mitigating potential catastrophic risks from advanced AI models. This internal restructuring follows several high profile departures from the company safety and alignment divisions.

Models

Deepseek raises V4 API pricing by up to eleven times ahead of reported IPO

DeepSeek has implemented a substantial price increase for its V4 API, with some costs rising by up to eleven times as of 16 August 2026. This shift signals the end of the low cost strategy the company utilised to disrupt the market in 2025. The adjustment coincides with industry reports suggesting the firm is preparing for an initial public offering.

Models

Writer releases enterprise-optimised GLM-5.2 variant with token-saving technology

Writer has launched a new variant of the open-source GLM-5.2 model, specifically post-trained to meet enterprise reliability standards. This release introduces a specialised harness designed to significantly reduce token consumption, directly addressing the high operational costs associated with large-scale model deployment.

Ready to build something
extraordinary?

15 minutes. No pitch deck. Just a conversation about what AI can do for your team.

Talk directly with our AI specialists

15 min, no strings
No sales pressure
Prototype in 7 days