“We love the world where we can use both”: How Nvidia thinks about local and frontier models
The New Stack Frederic Lardinois
Nvidia's Joey Conway describes a hybrid approach where organizations deploy both small local models and large frontier cloud models, routing simple tasks to local systems and complex ones to cloud models to reduce costs and latency. The DGX Spark, a $4,699 desktop machine with 128GB memory, can run models up to 200 billion parameters locally, with Nvidia's software stack handling routing and inference. This model-ensemble strategy keeps sensitive data on-premise while maintaining access to more powerful reasoning when needed, fundamentally shifting from single large models to specialized task-specific systems.
Why it matters
The models small enough to run on the box on your desk are getting good enough that the interesting question The post “We love the world where we can use both”: How Nvidia thinks about local and frontier models appeared first on The New Stack.