Jeff is an open-source project that provides small, locally-runnable decision models for fast zero-shot classification tasks. The models—fine-tuned versions of Qwen3.5 and Gemma 4—operate without generating text, instead returning calibrated probabilities for user-defined categories in a single forward pass.

According to the source, the 0.8B model runs in approximately 22 milliseconds on an NVIDIA RTX PRO 6000 GPU and 28 ms on an Apple M4 Max (using MLX). The larger 2B model takes 24 ms and 60 ms respectively on the same hardware. This speed enables real-time decision-making for applications like support-queue routing, user-intent detection, moderation labeling, and voice commands. Importantly, the zero-shot capability means categories don’t need to appear in training data; users describe options in plain language, and the model assigns probabilities to each.
On standard benchmarks covering 4,599 questions from five public datasets, Jeff-Qwen3.5-0.8B achieved 79.1% accuracy overall, approaching the performance of Jev (83.0%), a larger proprietary model. The 2B version scored 83.1% on the same benchmarks. However, the source notes that on reasoning-heavy tasks like BBH and JudgeBench, Jeff underperforms larger models—an expected tradeoff at this model size.
The entire project was developed on local hardware. The 0.8B model trained in about two hours on a single RTX PRO 6000 workstation GPU; the 2B variant took approximately 3.5 hours. All synthetic training data came from Qwen3.8-Flash-Next, an open model, running on two DGX Spark instances. A closed model was used only to spot-check data quality, not to generate training examples.
Jeff uses the same request format as Jev but is independent and not affiliated with TypeSafe, Jev’s makers. The training methodology builds on the open-source AutoJev recipe. For cases where zero-shot accuracy is insufficient, fine-tuning on domain-specific examples can significantly improve performance—the source reports a voice-navigation fine-tune moving accuracy from 31.7% to 95.8% in under 30 minutes on a single GPU.
The project provides models via Hugging Face and includes setup instructions for deployment on NVIDIA GPUs, Apple silicon, and CPUs. Game-playing tests (Doom, Frogger, Pac-Man) demonstrated the model’s ability to make decisions in novel contexts described in natural language, with decision latency of 29–49 ms per move on M4 Max.
Key facts
- Jeff-Qwen3.5-0.8B makes decisions in ~22 ms on NVIDIA RTX PRO 6000 and ~28 ms on Apple M4 Max
- Zero-shot classification achieves 79.1% overall accuracy on benchmark suites, approaching 83.0% from larger Jev model
- Models trained entirely on local hardware: 0.8B in ~2 hours, 2B in ~3.5 hours on a single RTX PRO 6000 GPU
- Fine-tuning on domain examples can boost accuracy from 31.7% to 95.8% in under 30 minutes
- Independent open-source project using same request format as Jev but not affiliated with TypeSafe
