Why it matters
Scaling production AI requires strict control over inference expenditure and runtime dependencies. Embracing open-weight models grants engineering teams greater architectural flexibility, lower operating costs, and enhanced ownership of their deployment pipelines.
Key points
- Open-weight model usage among US companies surged from 7% to 56% of total token share.
- Rising costs associated with frontier proprietary APIs are motivating the transition.
- Teams are favouring cost efficiency and infrastructure independence for ongoing production workloads.



