Your GPU. Your models. One router.
Fast & opinionated
local AI routing
One lightweight daemon. A single Zig binary. Generative models on your own GPU box, starting with NVIDIA DGX Spark.
View on GitHub
How it works
The right model. When you need it.
- 01
One connection
One MCP server for every agent. An OpenAI-style HTTP API for everything else.
- 02
Room to run
Models load on demand and unload when idle. Idle models make room by priority.
- 03
Keep favourites warm
Priorities keep your favourite models ready for the next job.
- 04
Every model type
Image and video today. Voice, music and more planned as adapters.
Local, from the kernel up
Built to generate.
Hand-written Zig engines with custom GPU kernels.
Qwen-Image 2.1
Text to image · 1024 × 1024 in about 12–17 seconds.
MiniMax H3
Text to video · 5 seconds at 768 × 448, with stereo audio.
Measured on a DGX Spark. Image editing and image to video are in progress.
Quick start
Install. Connect. Generate.
One command to install. One connection for your agent.
git clone https://github.com/jayleaton/localrouter && bash localrouter/tools/spark/install.sh
claude mcp add --transport http localrouter http://<host>:8190/mcpReplace <host> with your GPU box’s hostname or IP address.