After years of one-upmanship on model size and benchmark scores, the AI contest is shifting to cost, control and where computation runs. Companies are increasingly using systems that route tasks to the most suitable—and cheapest—model, blending open-weight and proprietary options. Perplexity’s latest approach leans on a lower-cost Chinese model for most work, escalating only when needed, while investors like Benchmark’s Peter Fenton say open-weight models could soon account for the bulk of tokens generated. That threatens the margins of frontier labs and accelerates enterprise moves to run models locally or in controlled environments, aided by platforms like Ollama. The pivot also raises strategic questions for the U.S. as capable open models emerge from China, and could reshape data center demand as routine workloads migrate to devices and on-prem systems.
Related articles:
Llama 2: Open Foundation and Fine-Tuned Chat Models
Mixtral 8x7B
Meta Llama 3
NVIDIA H200 Tensor Core GPU
Ollama Model Library




























