Nvidia used the opening day of IFA 2026 to outline a comprehensive local AI strategy, releasing a free open-source tool that turns idle household PCs into a shared AI compute cluster and confirming that its ARM-based RTX Spark Windows PCs will begin shipping in October.
PAIR Pools Home PCs for AI Workloads
The centerpiece of Thursday’s announcements is NVIDIA Personal AI Router, or PAIR, software that automatically discovers compatible PCs on a local network and routes AI inference requests to whichever machine has available capacity. More than half of U.S. households have two or more PCs, and Nvidia argues much of that computing power sits idle throughout the day.
PAIR works with the popular open-source inference applications Ollama and LM Studio and supports Windows, macOS, and Linux. Compatible hardware includes GeForce RTX 20-series GPUs and newer, RTX PRO workstation GPUs, DGX Spark, and Apple M4 or later chips. The tool does not fuse multiple GPUs into a single larger machine. Instead, it distributes independent inference requests across devices, reducing bottlenecks when AI agents split complex tasks into parallel sub-jobs. In one Nvidia test, five sub-agents working simultaneously took about 18 minutes on a single device but finished in under nine minutes when distributed across three machines using PAIR.
Simpler Setup and Faster Inference
Nvidia also addressed one of the persistent barriers to local AI adoption: setup complexity. Three agent applications — Hermes Agent from Nous Research, Perplexity’s Portable Computer, and OpenClaw — are getting streamlined one-click installation on RTX and DGX systems, each built on llama.cpp with Nvidia’s inference optimizations already applied. On the performance side, the llama.cpp backend now delivers up to 1.9x higher throughput on a GeForce RTX 5090, while vLLM shows gains of up to 1.4x on two-DGX Spark clusters.
RTX Spark PCs Arrive in October
Nvidia confirmed that RTX Spark Windows PCs will ship in October under the product name N1X. The platform pairs a Blackwell-based RTX GPU with a 20-core Grace CPU, supports up to 128GB of unified memory, and delivers up to one petaflop of AI compute. At IFA, Lenovo unveiled the Yoga Pro 9n and Yoga 9n 2-in-1, while Acer showed a compact desktop design. Asus, Dell, HP, and MSI are also preparing devices. Electronic Arts, Embark, and Ubisoft have joined earlier supporters including Krafton, NetEase, Riot Games, and Xbox in committing game titles to the platform.
A wave of locally deployable models rounds out the picture: Meta’s Muse Glimmer, DeepSeek v4 Flash, Qwen’s 3.8-Flash-Next, and Nvidia’s own Nemotron 3.5 Lightning are all optimized for RTX hardware.
Why Local AI Matters
By distributing AI workloads across existing consumer hardware, Nvidia is positioning itself to reduce reliance on centralized cloud data centers for everyday inference tasks. The approach could lower latency for time-sensitive applications, cut recurring cloud costs for power users, and give enterprises and developers more control over data privacy by keeping sensitive workloads on-premises. For journalists and content creators, faster local inference means quicker turnaround on AI-assisted research, translation, and summarization tasks without uploading drafts to external servers.
What Comes Next
Developers can download PAIR immediately from Nvidia’s GitHub repository, while consumers will need to wait until October to experience the full RTX Spark ecosystem on new N1X-class devices. As more models are optimized for local deployment and more households adopt multi-PC setups, the distributed inference model could become a standard pattern for running AI agents at home and in small offices.









