- New option, same privacy model.
PM_CLASSIFY_BACKEND=jevsendshiggs classifyto a Jev decision server you run yourself. Your mail still never leaves hardware you own. The default staysllm, meaning your usual local model. - Faster. On the same RTX 3080, Open-Jev-2B scored each email in 0.24 s against 0.62 s for
gemma4:e4b(about 2.6×). End to end on 59 real messages,higgs classifywent from 1.86 s to 0.35 s per email (about 5.3×). That comparison is between the Mac setup higgs was using and the GPU box, so it mixes hardware. - Better at bulk mail and ranking. Jev flagged 18 of 25 automated messages as bulk where Gemma flagged 14. On the 155-email attention test it scored an AUROC of 0.87 against 0.81 for Gemma on the same GPU.
- Labels on par. On 42 messages with an unambiguous category, both picked the right top label 27 times.
Classification never needed a writer
Until now, higgs classify worked like every other LLM classifier: hand a generative model the message, ask for a JSON object with labels and a confidence, and parse what comes back. It works, but you're paying for a lot of machinery you don't use. The model writes tokens one at a time, it can return JSON that doesn't parse, and a slow generation can hit a request timeout. The "confidence" is also just a number the model typed.
A Jev model is a decision model. You give it a question and the options, and in one forward pass it returns a calibrated probability for each option. Nothing is generated. It handles three typed question kinds: yes/no (noul), pick one (choice), and a 0–5 rating (score). For triage, that's the whole job.
How higgs uses it
With the Jev backend enabled, higgs sends the first 1,200 characters of each message body to the Jev server and asks two questions:
- Yes/no: is this automated or mass-sent? The answer sets the message's
is_mailing_listflag. - Choice: which of the 11 taxonomy labels fits? The top label goes into
suggested_labels, plus the runner-up when its probability is at least 0.3.
The confidence higgs records is the top label's probability. There's no JSON for the model to get wrong and no generation step to time out. Only classification changes: summarize, digest, ask, and extract still use your PM_LLM_BACKEND (Ollama or a self-hosted OpenAI-compatible server) exactly as before.
One thing we learned while tuning: the wording of the question matters. When we asked only about newsletters and marketing, the bulk question caught 8 of 18 automated emails in a tuning sample. Naming notifications and unsubscribe or notification-settings links, which matches the rule higgs already uses for is_mailing_list, caught 14 of 18 at an AUROC of 0.992.
The test
We didn't use a public benchmark. We used a real inbox. The test set is 155 emails from the maintainer's own Proton Mail account: 35 he had kept because they needed his personal attention, and 120 picked at random from mail he had archived. Every model got the same question:
"Does this email need the recipient's personal attention (a reply, decision, deadline, or action), as opposed to automated or bulk mail?"
This question isn't literally one of the two that higgs classify asks. It's a proxy, and we chose it because the maintainer's own keep and archive decisions give us real ground truth for it. It tests the thing triage is for: separating mail a person has to deal with from everything else.
The main metric is AUROC, which measures how well a model's probability ranks needs-attention mail above everything else. 0.5 is a coin flip and 1.0 is a perfect ranking. We use it because it doesn't depend on picking a threshold.
Results
| Model | Hardware | Body chars | Time / email | AUROC |
|---|---|---|---|---|
gemma4:e4b-mlx (Ollama, structured JSON) | Apple M4 Pro, 48 GB | 500 | 0.78 s | 0.78 |
gemma4:e4b-mlx (Ollama, structured JSON) | Apple M4 Pro, 48 GB | 1200 | 0.75 s | 0.76 |
| autotrust/JEV-9B (MLX) | Apple M4 Pro | 500 | 1.5 s | 0.85 |
gemma4:e4b (Ollama GGUF, 100% GPU) | RTX 3080 10 GB | 500 | 0.62 s | 0.81 |
gemma4:e4b (Ollama GGUF, 100% GPU) | RTX 3080 10 GB | 1200 | 0.64 s | 0.78 |
| Open-Jev-2B (PyTorch, batch 16, no conv kernel) | RTX 3080 10 GB | 500 | 0.32 s | 0.87 |
| Open-Jev-2B (+ causal-conv1d CUDA kernel) | RTX 3080 | 500 | 0.24 s | 0.87 |
| Open-Jev-2B (+ causal-conv1d) | RTX 3080 | 1200 | 0.41 s | 0.885 |
155 emails (35 needs-attention, 120 archived), measured 2026-09-29. "Body chars" is how much of each message body the model saw.
What that looks like in an inbox
AUROC is abstract, so here are concrete operating points. We took each Gemma run's yes/no answer as-is and put Open-Jev-2B's threshold at 0.3:
| Model · setup | Caught | False alarms |
|---|---|---|
gemma4:e4b-mlx · M4 Pro · 500 | 24 | 21 |
gemma4:e4b-mlx · M4 Pro · 1200 | 24 | 22 |
gemma4:e4b · RTX 3080 · 500 | 25 | 16 |
gemma4:e4b · RTX 3080 · 1200 | 23 | 13 |
| Open-Jev-2B · RTX 3080 · 500 · threshold 0.3 | 29 | 22 |
Same 155-email test. Caught is out of 35 needs-attention emails; false alarms are out of 120 archived ones.
Open-Jev-2B surfaced four to six more of the emails that mattered than any Gemma run. On the GPU, Gemma raised fewer false alarms (13–16 against 22), so this is a trade-off, not a clean sweep. Because Open-Jev-2B returns a probability rather than a bare yes/no, 0.3 is a choice rather than a fixed property of the model: a lower cut-off catches more, a higher one raises fewer false alarms.
Speed, honestly
It's tempting to compare Open-Jev-2B on the RTX 3080 (0.24 s) with Gemma on the Mac (0.78 s), but those are different machines, so that ratio isn't a model-vs-model result. Here are the comparisons that keep the hardware and input the same:
| Setup | Gemma e4b | Jev |
|---|---|---|
| RTX 3080, 500 chars | 0.62 s · 0.81 | 0.24 s · 0.87Open-Jev-2B · ≈2.6× faster |
| RTX 3080, 1200 chars | 0.64 s · 0.78 | 0.41 s · 0.885Open-Jev-2B · ≈1.6× faster |
| Apple M4 Pro, 500 chars | 0.78 s · 0.78 | 1.5 s · 0.85JEV-9B (MLX) · slower |
Cells show time per email · AUROC on the same 155-email test.
Gemma's time barely changes with input length here (0.62 s to 0.64 s), while Open-Jev-2B's grows (0.24 s to 0.41 s). The speed gap narrows with longer inputs, but Jev stays ahead. Its AUROC held up with more text (0.87 to 0.885), while Gemma's slipped (0.81 to 0.78).
End-to-end: higgs classify on a real mailbox
The scoring test sends each model a trimmed prompt. A real run also fetches over IMAP, parses, and writes results. We ran higgs classify INBOX --no-state on 59 messages from the maintainer's inbox. Wall-clock time includes the IMAP fetch, and every run finished with 0 errors.
| Setup | Workers | Total | Per email |
|---|---|---|---|
gemma4:e4b-mlx · M4 Prothe setup higgs was using | 2 | 109.8 s | 1.86 s |
| Open-Jev-2B · RTX 3080 | 1 | 28.6 s | 0.48 s |
| Open-Jev-2B · RTX 3080 | 4 | 20.4–21.1 s | 0.35 s |
59 messages, --no-state. Total is wall-clock time including the IMAP fetch. The Gemma and Jev rows ran on different machines.
That's about 5.3× faster end to end. Again, it compares the Mac setup higgs was actually using with Jev on a GPU box, so part of the gain is the hardware. The Jev server scores one request at a time, but 4 workers overlap the network round trips, which is where 28.6 s drops to 20.4 s.
Quality on the same 59 messages (34 personal, 25 automated):
- Bulk flag: Jev marked 18 of 25 automated messages as bulk, where Gemma marked 14.
- False bulk flags: Jev flagged 1 of 34 personal messages, a Google Calendar invitation (which is itself machine-sent). Gemma flagged none.
- Category labels: on 42 messages with an unambiguous category, both got the top label right 27 times. Jev's labels are on par with Gemma's, not better.
For longer-run context: the day before, a production run with gemma4:e4b-mlx sustained about 2.2 s per email (roughly 5,800 emails in about 3 hours with 2 workers). That run sent the full 3,000-character snippet and generated a JSON rationale for each message. The larger local MLX models we tried were worse. The 35B and 31B models were unusable for classify because requests timed out, and gemma4:12b-mlx took about 40 s per email.
Why not the 9B model on the Mac?
autotrust/JEV-9B (Apache-2.0, built on Qwen3.5-9B) runs on Apple Silicon through MLX once its LoRA adapter is merged. It's a real option if a Mac is the only machine you have: on the M4 Pro it ranked better than Gemma (0.85 vs. 0.78 AUROC). It was also about twice as slow (1.5 s vs. 0.78 s per email) and no more accurate than Open-Jev-2B. So the setup we recommend is Open-Jev-2B on an NVIDIA GPU.
Set it up
The server is the open-source Open-Jev loader (MIT) serving the Open-Jev-2B checkpoint (an Apache-2.0 LoRA adapter plus a scalar decision head on Qwen/Qwen3.5-2B). Follow the Open-Jev README to install it and fetch the checkpoint, then start the server bound to loopback:
# On the machine with the GPU. Binds to 127.0.0.1 only. $ python -m jev.server --checkpoint ./checkpoints/open-jev-2b/package/checkpoint \ --device cuda:0 --max-length 4096 --batch-size 16 --no-prefix-cache \ --host 127.0.0.1 --port 8791
Then point higgs at it. Only classification moves; the other AI commands keep using your usual local model:
# Route classify through Jev $ export PM_CLASSIFY_BACKEND="jev" $ export PM_JEV_BASE_URL="http://127.0.0.1:8791" # default # summarize / digest / ask / extract are unchanged $ export PM_LLM_BACKEND="ollama" PM_OLLAMA_MODEL="gemma4" # Preview first, apply when it looks right $ higgs classify --dry-run --limit 20 INBOX
higgs sends the Jev server the first 1,200 characters of each message body. If you install the optional causal-conv1d CUDA kernel in the server's environment, scoring gets faster (0.32 s down to 0.24 s per email on our 3080, at 500 characters).
Privacy note. By default the Jev server and higgs share a machine and talk over 127.0.0.1, so no mail leaves the box. If your GPU is in a different machine, PM_JEV_BASE_URL can point at it, but your mail will then cross your own network to get there. Keep that link on hardware and a network you control.
What we didn't measure
- Other inboxes. This is one person's mail: 155 messages for the scoring test and 59 for the end-to-end run. It's a small real-world test, not a benchmark suite, and your mix of newsletters, receipts, and human mail will differ.
- Statistical significance. With samples this small, a few emails either way can move these numbers.
- Energy or cost. We didn't measure power draw, so we're making no efficiency claims.
- Per-category accuracy. We only checked the top label on 42 unambiguous messages, not accuracy for each of the 11 labels.
- Other GPUs and Macs. The only hardware we measured was one RTX 3080 (10 GB) and one M4 Pro (48 GB). We didn't run Open-Jev-2B on the Mac or end to end with Gemma on the 3080.
Credits
Jev is originally a closed decision model from TypeSafe AI. The open models we tested are independent reproductions: Open-Jev and its Open-Jev-2B checkpoint, and autotrust/JEV-9B. Thanks to both projects for releasing their work openly. higgs has no affiliation with TypeSafe AI, Open-Jev, or autotrust.
higgs is open source (Apache-2.0) with zero telemetry. Install it, read the README, or browse the blog.