Jeff: The Tiny Local Model That Makes Decisions in ~30 ms
Imagine a support ticket arriving with one practical question attached: should it go to refunds, damaged parcels, or account support? A human might explain the decision in a paragraph. Production software usually needs something smaller and more useful: one choice, a probability, and a safe fallback when confidence is low.
Jeff is built for that narrow moment. It is a family of fine-tuned models based on Qwen3.5 and Gemma 4 that perform zero-shot classification, meaning they choose among categories described in the request even when those categories were not named during training. The project uses the same request shape as Jev, but it is an independent implementation rather than an official companion.
The headline model is Jeff-Qwen3.5-0.8B. The 0.8B means the network contains roughly 800 million learned parameters, or numerical values adjusted during training. That is tiny beside many general-purpose language models, and it is the reason Jeff can behave more like a local decision component than a remote chatbot. On the project’s published measurements, it takes a median 22 milliseconds per decision on an NVIDIA RTX PRO 6000 and 28 milliseconds on an Apple M4 Max using MLX. That is where the roughly 30-millisecond claim comes from.
Jeff is a decision layer, not a chat window
A conventional language model generates text one token at a time. Jeff takes a different route. It runs one forward pass, which means one complete evaluation of the input through the network, and reads the result as probabilities over the available answers. There is no generated explanation to parse and no fragile search through a paragraph for the words yes or no.
A calibrated probability is meant to be more than a confident-looking number. It should roughly match reality over many decisions: predictions near 0.8 should be correct close to 80 percent of the time. That makes the output useful in software, where a confidence threshold can decide whether to automate a task or send it to a human review queue.
The request format is deliberately plain. A typical payload might look like this:
request = {
'state': 'The parcel arrived crushed and the customer wants a refund.',
'questions': {
'route': {
'type': 'choice',
'instructions': 'Which team should handle this?',
'criteria': {
'1': 'Refunds and payments',
'2': 'Damaged or lost parcels',
'3': 'Account and login problems'
}
},
'angry': {
'type': 'noul',
'instructions': 'Is the customer angry?'
}
}
}
The project supports choice questions with up to 255 options, noul questions for yes-or-no judgments, and score questions for placing something on a scale. Several independent questions can travel in one request, which is handy when a ticket needs both a route and a sentiment flag.
Why 0.8B is a useful size
The 0.8B model’s 16-bit weights take about 1.7 gigabytes. Weights are the learned values stored by the model, so this is a much more approachable footprint than the tens of gigabytes often associated with larger reasoning systems. It can run locally on an NVIDIA machine, and the project also provides an MLX path for Apple silicon.
The hardware still matters. The same measurements put the model at about 463 milliseconds on a 32-thread CPU, far slower than the listed GPU and laptop results. Local inference does not automatically mean instant inference; it means the network and model are close to your application, without a request crossing the internet. For a local service handling many small decisions, that difference can be more important than a polished conversational answer.
The prompt is part of the model
Jeff is not a planner. It will not reliably invent a multi-step strategy, simulate a world, and then select the best move. The stronger pattern is to let ordinary code do the reasoning that code is good at: calculate the current state, enumerate legal actions, and describe the consequences. Jeff then makes the final choice.
Wording matters because the model compares the descriptions it receives. Short option keys such as 1, 2, and 3 keep the output compact, while the text behind each key should use a consistent style. An option that says a move will result in being hit by a car is easier to compare with other consequence-based options than a bare command such as move forward.
That distinction explains why Jeff can perform well on a focused task while failing at an apparently related one. It is a fast classifier, not a hidden chain-of-thought engine. Treating it like a planner is likely to produce a disappointing result.
Benchmark numbers need some context
The project reports an overall score of 79.1 for Jeff-Qwen3.5-0.8B across five public benchmarks, compared with 83.0 for published Jev results. Those figures are informative but not a perfect head-to-head experiment because the published numbers were measured on a different sample of the same benchmark collections.
The split between tasks is more revealing. Jeff reaches 96.4 on Financial PhraseBank and 86.1 on RAGTruth, where classification and grounding are central. On more reasoning-heavy tests, its score drops to 64.0 on BIG-Bench Hard and 47.6 on the separate hard tier of JevBench. The larger Jev model remains far stronger in those settings. A small decision model can be excellent at choosing between clearly described options without becoming a miniature general reasoner.
Training at home changes the equation
Jeff is not only small at runtime. The project also describes a local training pipeline that avoids cloud GPUs. The 0.8B model reportedly takes about two hours to train on one RTX PRO 6000 workstation GPU, while the 2B version takes roughly three and a half hours.
The recipe uses full-weight fine-tuning, which means updating the model’s learned weights rather than attaching only a small adapter, for one pass through the training data. It then fits a temperature for calibration, a post-training adjustment that helps the probability scores line up better with observed accuracy. Synthetic examples are produced by an open model, filtered for possible benchmark leakage, and evaluated without training directly on the benchmark panel.
Domain-specific examples can matter more than another fraction of a benchmark point. In the project’s voice-navigation experiment, about 11,000 app-specific examples lifted held-out accuracy from 31.7 percent to 95.8 percent in around half an hour on one GPU. Held-out means those examples were kept separate from training so they could test whether the model learned a useful pattern rather than memorized the answers. That fine-tuned model still made decisions in about 40 milliseconds on an M4 Max.
Where Jeff fits
Jeff makes sense anywhere an application already knows the situation and needs a quick judgment: support routing, user-intent detection, moderation labels, voice commands, or game actions. It is English-only and text-only, and it works best when the options are concrete and their consequences are written in comparable language.
The larger lesson is not that a 0.8B model can replace a powerful reasoning system. It cannot. Jeff shows something more practical: a language model can be turned into a fast, local decision primitive. Let the surrounding application reason, calculate, and enforce rules; let Jeff choose among the options in front of it. At roughly 30 milliseconds, that boundary becomes small enough to disappear inside an ordinary software interaction.
Comments (0)
No comments yet. Be the first to respond!
Leave a Comment
Your comment will be visible after review.