The number of users who want to run AI models not in the cloud but on their own computer or on hardware in their own office grows every month. NVIDIA has announced one of its most comprehensive update packages aimed at this demand. According to information reported by BabelTechReviews, the company brought together both clustering solutions and performance improvements in a single announcement spanning from the DGX Spark ecosystem to open models running on RTX cards. All of the announcements came right after the Nemotron 3.5 Lightning model introduced the same day; this shows that what is at hand is not a single product launch but an ecosystem strategy in which each piece feeds the others.
Why Has Local AI Become So Important?
When people hear local AI, the first thing that often comes to mind is 'models that run without the internet,' but there is much more behind it. In cloud-based services, every query has to travel to remote servers and back; this creates latency and means a recurring service fee. Add data privacy to that: for teams that do not want to send their own documents, code, or customer data outside, local execution becomes the only realistic option. Therefore, every move NVIDIA makes in this area enters the radar not only of curious users but also of enterprise developers.
On the other side, the hardware side had also been preparing for this scenario for a long time. The quality of open-weight models increases with every release, and capabilities that could previously be run only in data centers can now fit into desktop-level systems. As this balance changes, so does the answer users give to the question 'cloud or local?' NVIDIA's latest announcements should be read within this framework: the company aims to make local hardware more attractive through software-side improvements.

DGX Spark Ecosystem: From a Single System to a High-Speed Cluster
The most striking headline on the hardware side of the announcement is a serious capability increase for DGX Spark users. With the new NVIDIA Sync Cluster Assistant, users can connect two or more DGX Spark systems into a high-speed cluster. This assistant does not merely establish the connection; it also manages network configuration, workload routing, and automatic health monitoring from a single place. The automation of these steps, which previously required serious manual effort, makes it easier especially for small teams to set up their own mini data center.
Two additional innovations in the same package directly affect daily use. Google Chrome now officially arrives as a native ARM64 Linux build, which makes browser-based workflows on DGX Spark smoother. The accompanying NVIDIA Sync Resource Monitor shows CPU and GPU usage both in real time and historically. Moreover, this monitoring is not limited to a single system; it becomes possible to follow the entire cluster from one screen.
- NVIDIA Sync Cluster Assistant: Turns two or more DGX Sparks into a high-speed cluster; network configuration, workload routing, and automatic health monitoring are brought under one roof.
- Native Chrome (ARM64 Linux): Google Chrome's official ARM64 Linux build comes to DGX Spark and speeds up browser-based development workflows.
- NVIDIA Sync Resource Monitor: Provides a real-time and historical view of CPU and GPU usage; it works on a single system and across the entire cluster.
When we consider these three innovations together, the picture is clear: NVIDIA wants to turn DGX Spark from 'a powerful box on its own' into a scalable local infrastructure. It is no coincidence that monitoring and clustering tools arrive at the same time; they directly target the lack of visibility that is the biggest headache of multi-system setups. On the user side, this means a path to grow without making an expensive setup on day one.
Always-On Agents on RTX Cards: Muse Glimmer
Moving to the consumer side, we encounter Meta's Muse Glimmer model. This 30-billion-parameter, open-weight model is specifically optimized for 'always-on' local agents. The phrase 'always-on' matters here: assistants that run continuously in the background and can take on tasks without waiting for the user's request require a very different hardware profile from classic chatbots. In such scenarios, it becomes critical for the model to stay in memory and respond with low latency.
NVIDIA's figure stands out at this point: on a flagship card such as the GeForce RTX 5090, Muse Glimmer can be run locally at over 200 tokens per second. This represents the difference between 'trying out' a language model and making it part of a daily workflow. For the user, this speed makes scenarios such as code completion, document summarization, or a continuously running agent truly practical. It is also worth remembering that high token speed means shorter response times and a smoother experience.
LTX-2.5: A Local Leap in Video Generation
On the video side, LTX-2.5 stands out. The new version of the open video model supports multishot generation and generative editing. Multishot support makes it possible to generate sequences made up of consecutive shots instead of a single scene; generative editing opens the door to modifying an existing video without regenerating it from scratch. These two capabilities move local video generation from the 'experimentation' level to a level that can 'participate in the editing process.'
The optimizations NVIDIA announced concern the hardware side. The model is specifically optimized for NVIDIA RTX GPUs, DGX Spark, and DGX Station. As a result, users achieve a 2x performance increase while seeing a 40 percent reduction in memory usage. This decrease on the memory side is especially important because video models are generally among the workloads that strain VRAM the most. In other words, it becomes possible to produce longer clips and more complex projects with the same hardware.
Seven New Open Models Accelerated by NVIDIA
The broadest headline of the announcement is the new open models added to NVIDIA's acceleration list. The company announced seven models at once, addressing different use cases for developers to explore:
- Cosmos 3 Edge – This version of the Cosmos family has joined the list of accelerated models.
- MiniMax-H3 – This model bearing the MiniMax signature is also among those supported.
- Poolside's Laguna S 2.1 model – Poolside's new version is on the list of models offered with NVIDIA acceleration.
- DeepSeek-V4-Flash – DeepSeek's 'Flash'-labeled version is now among the accelerated models.
- Thinking Machines Lab's Inkling-Small model – This model from Thinking Machines Lab is one of the list's new members.
- Unsloth Desktop – This Unsloth version aimed at the desktop side has also been brought under acceleration.
- Wan-Animate-2 – The second version of the Wan-Animate family has also been added to the supported models.
NVIDIA has not shared detailed performance figures for each of the models in this announcement; the list answers more the question 'which models are ready now?' Still, this expansion itself is meaningful: the acceleration list growing with models from different categories shows that developers can experiment without being tied to a particular ecosystem. For teams that do not want to depend on a single provider's models, such lists serve as a guarantee that extends the life of a hardware investment.
What Do These Updates Actually Offer, and to Whom?
When we read the announcement as a whole, three separate user profiles stand out. A developer starting with a single DGX Spark finds the scaling path ready through the clustering assistant when the workload grows. An enthusiast with an RTX-based desktop can try local agents and video generation with models such as Muse Glimmer and LTX-2.5. On the enterprise side, the local infrastructure option becomes more realistic every day for teams that have to work without sending data outside.
The cost side should not be overlooked either. In cloud services, as usage increases, expenses grow linearly, whereas with local hardware the cost turns into an upfront investment. Therefore, the appeal of local solutions is fueled not only by privacy but also by the needs of teams that want a predictable budget. NVIDIA offering tools for both clusters and single systems aims to ease the transition between these two extremes. The continuous arrival of software optimizations creates an effect that extends the life of existing hardware.
What to Expect Next?
In NVIDIA's roadmap, local AI will apparently remain a long-term headline. The company's introduction of the Nemotron 3.5 Lightning model on the same day and its publication of accompanying technical blog posts show that investment on the software side continues. In the coming period, the accelerated model list can be expected to grow further, and existing optimizations can be expected to deepen with new versions. For users, the critical question is this: which workloads make more sense to keep in the cloud, and which locally?
The answer to this question will vary from user to user, but the balance is clearly shifting toward local. When clusterable desktop systems, local models exceeding 200 tokens per second, and video generation with improved memory efficiency come together, 'local AI' is no longer a hobby but a serious production option. NVIDIA's latest announcement is focused precisely on accelerating this transition.