NVIDIA, in collaboration with Microsoft and various software partners, has significantly enhanced the capability to run AI agents locally on NVIDIA hardware. This initiative coincides with the launch of new Windows PCs in the RTX Spark line, designed to facilitate a seamless experience for users.
Streamlined Software Installation
The latest announcements emphasize a simplified installation process for local agent software, improved inference speeds in popular open-source tools, and the introduction of NVIDIA PAIR—a new software layer that efficiently distributes inference tasks across multiple PCs within a local network.
Three notable agent applications are set to benefit from this streamlined setup on Windows systems equipped with NVIDIA GPUs: Hermes Agent, OpenClaw, and Perplexity Portable Computer. These enhancements aim to minimize the manual configuration previously required for model selection, server matching, and local settings adjustments.
- Hermes Agent: Developed by Nous Research, this application promises a one-click local setup for RTX and DGX systems on Windows. It automatically detects the installed NVIDIA GPU, selects an appropriate model and configuration, and runs it using an integrated version of llama.cpp, optimized for NVIDIA hardware. Once installed, Hermes is designed to maintain context across tasks, retain information between sessions, and develop reusable skills over time. Linux support for this simplified setup is anticipated in the near future.
- OpenClaw: This application is also introducing a simplified setup process for local models on RTX GPUs with a minimum of 24GB of VRAM. With a robust open-source community backing it, OpenClaw has collaborated with NVIDIA and Microsoft to ease the installation process for users.
- Perplexity Portable Computer: Already operational on Linux systems, including DGX Spark, this application is being extended to NVIDIA RTX GPUs with at least 24GB of VRAM. Windows support is forthcoming, which will broaden access to its bundled approach for models, orchestration, and tools.
This software allows users to maintain complete workflows on their devices while selectively sending tasks to cloud models when additional research or reasoning is necessary. It prioritizes user privacy by requesting permission before transmitting any content to the cloud, ensuring sensitive information remains secure on personal devices.
Performance Enhancements
NVIDIA is also focusing on boosting performance within the open-source inference stack. Collaborative efforts with the llama.cpp and vLLM communities have led to significant throughput improvements for local agent workloads across GeForce RTX, RTX Pro, and DGX systems.
Recent updates to llama.cpp have resulted in throughput increases of up to 1.9 times on a GeForce RTX 5090, thanks to kernel modifications, speculative decoding enhancements, and faster prefill processes. Similarly, vLLM has demonstrated throughput gains of 1.2 times on an RTX Pro 6000 Blackwell Workstation Edition and up to 1.4 times on two DGX Spark clusters. These performance enhancements are directly accessible in inference backends and applications such as LM Studio and Ollama, reflecting the growing trend among developers and enthusiasts to run language models locally rather than relying on hosted environments.
Leveraging Idle Computing Resources
The newly introduced NVIDIA PAIR software is tailored for users with multiple machines on the same network. PAIR, which stands for Personal AI Router, automatically identifies compatible PCs and efficiently routes inference requests to the system with available capacity.
This software is compatible with Ollama and LM Studio and is currently in beta for Windows, macOS, and Linux. It supports a range of hardware, including GeForce RTX 20 Series GPUs and newer models, RTX Pro workstation GPUs based on Turing architecture and later, DGX Spark systems, and Apple M4 or newer silicon. NVIDIA envisions PAIR as a means to optimize underutilized computing resources in homes and small offices, allowing agents to distribute tasks across several machines instead of relying on a single GPU to process requests sequentially.
Upcoming RTX Spark Systems
The push for local AI capabilities aligns with NVIDIA’s imminent release of RTX Spark Windows PCs, set to debut in October through hardware partners like Lenovo and Acer. The lineup includes compact desktop and laptop designs, with Lenovo unveiling the Yoga Pro 9n and Yoga 9n 2-in-1, while Acer has showcased a compact desktop concept. These systems are built around a Blackwell GPU and Grace CPU architecture, designed to cater to creators, gamers, and local AI agents alike.
Additionally, game publishers such as Electronic Arts, Embark, and Ubisoft are among those bringing titles to RTX Spark systems, further enhancing the appeal of this new hardware.
Creative Software Integration
CyberLink is one of the software companies poised to support the RTX Spark initiative. Its PhotoDirector AI PC Mode will incorporate image diffusion models into PhotoDirector 365, facilitating tasks such as image editing, enhancement, object removal, and background replacement. Users will have the flexibility to choose between local and cloud processing, with local AI processing leveraging TensorRT-RTX and FP8 on NVIDIA GPUs.
This integration highlights NVIDIA’s broader strategy to connect local AI applications not only to coding and agent software but also to consumer-focused creative applications. The overarching trend in the industry is a shift towards migrating generative AI workloads from cloud services to local devices, particularly for users concerned about privacy, latency, and ongoing costs. NVIDIA’s recent initiatives underscore its commitment to fortifying the software ecosystem surrounding its GPUs as competition intensifies over the management of AI workloads.
According to NVIDIA, more than half of U.S. households possess two or more PCs, much of which remains underutilized throughout the day, presenting a significant opportunity for maximizing computing power.