Magnitude Wants Your Laptop to Tune Its Own AI Engine
Picture a local AI agent reviewing a codebase. It reads a long set of instructions, scans files, calls a shell tool, waits for the result, then sends another large prompt to the model. The model itself matters, but so does everything underneath it: memory movement, chip-specific instructions, caching, and the amount of work required for every generated word.
That is the problem Magnitude, a YC S25 startup, is tackling with an open-source inference engine for agents. Inference means running a trained artificial intelligence model to produce an answer. Magnitude profiles the computer you already own, recommends models that fit, and tunes the execution path for that particular machine instead of treating every laptop, graphics card, and processor as interchangeable.
Why can the same model feel fast on one machine and sluggish on another? The answer is often hiding below the model weights.
Self-optimizing does not mean retraining
The phrase self-optimizing can sound as if Magnitude changes a model’s intelligence while it runs. That is not what is happening. The model stays the same; the inference runtime, meaning the software that executes the model, adapts how the work is scheduled on your hardware.
Magnitude begins by examining the chip, available memory, and memory bandwidth, which is the rate at which data can move between parts of the computer. It then estimates whether each model and model variant will fit while leaving room for the conversation’s context window. A context window is the amount of prior text, tool output, and instructions the model can consider in one request.
The catalog also accounts for quantization. Quantization stores the model’s numerical weights with fewer bits, reducing download size and memory use. A Q4 version usually needs less memory than a Q8 version, although the smaller representation can lose some fidelity. Magnitude’s job is to find the highest-quality option that still leaves enough headroom for a real agent session.
The decision can be pictured with a small piece of illustrative pseudocode:
fitting_models = []
for model in catalog:
fits = model.memory + context_budget <= machine.free_memory
if fits:
model.score = (
0.45 * model.quality
+ 0.35 * model.speed
- 0.20 * model.memory_cost
)
fitting_models.append(model)
best_model = max(fitting_models, key=lambda model: model.score)
This is not a public Magnitude programming interface. It shows the idea: model selection becomes a hardware-aware decision rather than a guess based on a model’s name or parameter count.
The speedup happens below the model
A major part of the project’s approach involves kernels. A kernel is a small, highly optimized routine that performs a particular piece of mathematical work on a processor or graphics processor. General-purpose local model runners need kernels that work across many kinds of machines. Magnitude says it compiles and tunes those kernels on the device where the model will actually run.
That distinction matters because local inference has two different rhythms. During prefill, the engine processes the prompt and builds the internal state needed to answer. During decode, it generates the response one token at a time. A token is a small piece of text, such as a word fragment or punctuation mark. Agents often spend substantial time in both phases: they send long prompts during prefill, then generate tool calls and explanations during decode.
Magnitude’s current project page reports up to twice the speed of llama.cpp in its displayed comparisons, including a 92 percent decode improvement on Metal and a 19 percent decode improvement on CUDA. Metal is Apple’s graphics and compute framework; CUDA is NVIDIA’s platform for parallel computation. Those figures are project benchmarks rather than a guarantee for every model and computer, but they show where device-specific tuning can matter.
The comparison is also more nuanced than a race between two model runners. llama.cpp is a widely used open-source runtime with broad hardware support. Ollama and LM Studio make local models approachable too. Magnitude’s focus is narrower: choose a good model for the machine, then tune the entire path around that pairing.
Agents make memory and scheduling matter
A chatbot may handle one prompt at a time. An AI agent is a program that can plan work, call tools, inspect files, and continue across several steps. The application that coordinates those actions is often called an agent harness. This creates a different performance problem from ordinary chat.
An agent repeatedly resends instructions, conversation history, file excerpts, and tool results. Recomputing all of that from scratch wastes time, so inference engines maintain a key-value cache, commonly called a KV cache. It stores intermediate attention data from earlier tokens. When multiple sessions share the same opening instructions, a prefix cache can reuse the common beginning instead of processing it repeatedly.
Magnitude says concurrent sessions share prefix caches to avoid a large slowdown as more agents run at once. It also reports using less memory per agent and freeing model memory when agents stop. Models can be loaded on demand, then unloaded when the machine becomes crowded or a session goes idle.
Another technique is speculative decoding. A smaller draft model proposes several likely next tokens, while the larger model checks those guesses in batches. Correct guesses can reduce the number of expensive steps required from the larger model. For an agent that generates many short tool calls, these small savings can accumulate across an entire task.
The workflow is machine-first
The practical workflow is deliberately less technical than the machinery underneath it. Magnitude’s desktop app profiles the computer, presents recommended models, and downloads the chosen one. The bundled magnitude command-line interface, or CLI, is available without a separate installation.
The connection then looks roughly like this:
agent harness
│ one-click connection or OpenAI-compatible API
Magnitude inference engine
│ tuned kernels, caches, and model loading
local model on your computer
An application programming interface, or API, is a defined way for software programs to communicate. An OpenAI-compatible API uses a request format that many agent tools already understand, so the agent can send prompts to a local Magnitude model without needing a completely new integration.
The project currently targets macOS, Linux, and Windows, with support for Apple Silicon, NVIDIA and AMD graphics processors, and CPU-only systems. Its catalog is designed to cover both compact models for memory-constrained machines and larger models for systems with more available memory. Once a model is downloaded, Magnitude says prompts, files, and model execution remain on the computer, with no token billing or remote API key required.
Where the approach fits—and where it does not
Magnitude cannot create memory that a computer does not have. A model that exceeds the machine’s capacity will still be too large, and a small local model will not automatically match a much larger hosted model’s reasoning ability. Device tuning improves the route from model to output; it does not change the model’s underlying knowledge.
There is also a trade-off in optimizing around a catalog. A broad runner may support an unusual model immediately, while a hardware-aware engine may provide its best results for model families it has specifically tuned. The first run can involve downloads, setup, and tuning work before the fast path is ready.
Still, the central idea is compelling. Local AI agents are not only a question of which model you download. They are a systems problem involving memory, scheduling, caching, and the exact silicon under your desk. Magnitude’s bet is that an inference engine should understand that whole environment—and tune itself to it—before the agent begins its work.
Comments (0)
No comments yet. Be the first to respond!
Leave a Comment
Your comment will be visible after review.