Blog · Classification

Jev decision models: faster triage, still on your hardware

higgs classify can now ask a local decision model instead of a text generator. On one real inbox it was faster, better at spotting bulk mail, and better at ranking what needs attention than the Gemma model higgs used before. Its category labels were on par, not better. Here's what we measured, what we didn't, and how to set it up.

TL;DR
  • New option, same privacy model. PM_CLASSIFY_BACKEND=jev sends higgs classify to a Jev decision server you run yourself. Your mail still never leaves hardware you own. The default stays llm, meaning your usual local model.
  • Faster. On the same RTX 3080, Open-Jev-2B scored each email in 0.24 s against 0.62 s for gemma4:e4b (about 2.6×). End to end on 59 real messages, higgs classify went from 1.86 s to 0.35 s per email (about 5.3×). That comparison is between the Mac setup higgs was using and the GPU box, so it mixes hardware.
  • Better at bulk mail and ranking. Jev flagged 18 of 25 automated messages as bulk where Gemma flagged 14. On the 155-email attention test it scored an AUROC of 0.87 against 0.81 for Gemma on the same GPU.
  • Labels on par. On 42 messages with an unambiguous category, both picked the right top label 27 times.

Classification never needed a writer

Until now, higgs classify worked like every other LLM classifier: hand a generative model the message, ask for a JSON object with labels and a confidence, and parse what comes back. It works, but you're paying for a lot of machinery you don't use. The model writes tokens one at a time, it can return JSON that doesn't parse, and a slow generation can hit a request timeout. The "confidence" is also just a number the model typed.

A Jev model is a decision model. You give it a question and the options, and in one forward pass it returns a calibrated probability for each option. Nothing is generated. It handles three typed question kinds: yes/no (noul), pick one (choice), and a 0–5 rating (score). For triage, that's the whole job.

How higgs uses it

With the Jev backend enabled, higgs sends the first 1,200 characters of each message body to the Jev server and asks two questions:

  1. Yes/no: is this automated or mass-sent? The answer sets the message's is_mailing_list flag.
  2. Choice: which of the 11 taxonomy labels fits? The top label goes into suggested_labels, plus the runner-up when its probability is at least 0.3.

The confidence higgs records is the top label's probability. There's no JSON for the model to get wrong and no generation step to time out. Only classification changes: summarize, digest, ask, and extract still use your PM_LLM_BACKEND (Ollama or a self-hosted OpenAI-compatible server) exactly as before.

One thing we learned while tuning: the wording of the question matters. When we asked only about newsletters and marketing, the bulk question caught 8 of 18 automated emails in a tuning sample. Naming notifications and unsubscribe or notification-settings links, which matches the rule higgs already uses for is_mailing_list, caught 14 of 18 at an AUROC of 0.992.

The test

We didn't use a public benchmark. We used a real inbox. The test set is 155 emails from the maintainer's own Proton Mail account: 35 he had kept because they needed his personal attention, and 120 picked at random from mail he had archived. Every model got the same question:

"Does this email need the recipient's personal attention (a reply, decision, deadline, or action), as opposed to automated or bulk mail?"

This question isn't literally one of the two that higgs classify asks. It's a proxy, and we chose it because the maintainer's own keep and archive decisions give us real ground truth for it. It tests the thing triage is for: separating mail a person has to deal with from everything else.

The main metric is AUROC, which measures how well a model's probability ranks needs-attention mail above everything else. 0.5 is a coin flip and 1.0 is a perfect ranking. We use it because it doesn't depend on picking a threshold.

Results

ModelHardwareBody charsTime / emailAUROC
gemma4:e4b-mlx (Ollama, structured JSON)Apple M4 Pro, 48 GB5000.78 s0.78
gemma4:e4b-mlx (Ollama, structured JSON)Apple M4 Pro, 48 GB12000.75 s0.76
autotrust/JEV-9B (MLX)Apple M4 Pro5001.5 s0.85
gemma4:e4b (Ollama GGUF, 100% GPU)RTX 3080 10 GB5000.62 s0.81
gemma4:e4b (Ollama GGUF, 100% GPU)RTX 3080 10 GB12000.64 s0.78
Open-Jev-2B (PyTorch, batch 16, no conv kernel)RTX 3080 10 GB5000.32 s0.87
Open-Jev-2B (+ causal-conv1d CUDA kernel)RTX 30805000.24 s0.87
Open-Jev-2B (+ causal-conv1d)RTX 308012000.41 s0.885

155 emails (35 needs-attention, 120 archived), measured 2026-09-29. "Body chars" is how much of each message body the model saw.

What that looks like in an inbox

AUROC is abstract, so here are concrete operating points. We took each Gemma run's yes/no answer as-is and put Open-Jev-2B's threshold at 0.3:

Model · setupCaughtFalse alarms
gemma4:e4b-mlx · M4 Pro · 5002421
gemma4:e4b-mlx · M4 Pro · 12002422
gemma4:e4b · RTX 3080 · 5002516
gemma4:e4b · RTX 3080 · 12002313
Open-Jev-2B · RTX 3080 · 500 · threshold 0.32922

Same 155-email test. Caught is out of 35 needs-attention emails; false alarms are out of 120 archived ones.

Open-Jev-2B surfaced four to six more of the emails that mattered than any Gemma run. On the GPU, Gemma raised fewer false alarms (13–16 against 22), so this is a trade-off, not a clean sweep. Because Open-Jev-2B returns a probability rather than a bare yes/no, 0.3 is a choice rather than a fixed property of the model: a lower cut-off catches more, a higher one raises fewer false alarms.

Speed, honestly

It's tempting to compare Open-Jev-2B on the RTX 3080 (0.24 s) with Gemma on the Mac (0.78 s), but those are different machines, so that ratio isn't a model-vs-model result. Here are the comparisons that keep the hardware and input the same:

SetupGemma e4bJev
RTX 3080, 500 chars 0.62 s · 0.81 0.24 s · 0.87Open-Jev-2B · ≈2.6× faster
RTX 3080, 1200 chars 0.64 s · 0.78 0.41 s · 0.885Open-Jev-2B · ≈1.6× faster
Apple M4 Pro, 500 chars 0.78 s · 0.78 1.5 s · 0.85JEV-9B (MLX) · slower

Cells show time per email · AUROC on the same 155-email test.

Gemma's time barely changes with input length here (0.62 s to 0.64 s), while Open-Jev-2B's grows (0.24 s to 0.41 s). The speed gap narrows with longer inputs, but Jev stays ahead. Its AUROC held up with more text (0.87 to 0.885), while Gemma's slipped (0.81 to 0.78).

End-to-end: higgs classify on a real mailbox

The scoring test sends each model a trimmed prompt. A real run also fetches over IMAP, parses, and writes results. We ran higgs classify INBOX --no-state on 59 messages from the maintainer's inbox. Wall-clock time includes the IMAP fetch, and every run finished with 0 errors.

SetupWorkersTotalPer email
gemma4:e4b-mlx · M4 Prothe setup higgs was using2109.8 s1.86 s
Open-Jev-2B · RTX 3080128.6 s0.48 s
Open-Jev-2B · RTX 3080420.4–21.1 s0.35 s

59 messages, --no-state. Total is wall-clock time including the IMAP fetch. The Gemma and Jev rows ran on different machines.

That's about 5.3× faster end to end. Again, it compares the Mac setup higgs was actually using with Jev on a GPU box, so part of the gain is the hardware. The Jev server scores one request at a time, but 4 workers overlap the network round trips, which is where 28.6 s drops to 20.4 s.

Quality on the same 59 messages (34 personal, 25 automated):

For longer-run context: the day before, a production run with gemma4:e4b-mlx sustained about 2.2 s per email (roughly 5,800 emails in about 3 hours with 2 workers). That run sent the full 3,000-character snippet and generated a JSON rationale for each message. The larger local MLX models we tried were worse. The 35B and 31B models were unusable for classify because requests timed out, and gemma4:12b-mlx took about 40 s per email.

Why not the 9B model on the Mac?

autotrust/JEV-9B (Apache-2.0, built on Qwen3.5-9B) runs on Apple Silicon through MLX once its LoRA adapter is merged. It's a real option if a Mac is the only machine you have: on the M4 Pro it ranked better than Gemma (0.85 vs. 0.78 AUROC). It was also about twice as slow (1.5 s vs. 0.78 s per email) and no more accurate than Open-Jev-2B. So the setup we recommend is Open-Jev-2B on an NVIDIA GPU.

Set it up

The server is the open-source Open-Jev loader (MIT) serving the Open-Jev-2B checkpoint (an Apache-2.0 LoRA adapter plus a scalar decision head on Qwen/Qwen3.5-2B). Follow the Open-Jev README to install it and fetch the checkpoint, then start the server bound to loopback:

# On the machine with the GPU. Binds to 127.0.0.1 only.
$ python -m jev.server --checkpoint ./checkpoints/open-jev-2b/package/checkpoint \
    --device cuda:0 --max-length 4096 --batch-size 16 --no-prefix-cache \
    --host 127.0.0.1 --port 8791

Then point higgs at it. Only classification moves; the other AI commands keep using your usual local model:

# Route classify through Jev
$ export PM_CLASSIFY_BACKEND="jev"
$ export PM_JEV_BASE_URL="http://127.0.0.1:8791"   # default

# summarize / digest / ask / extract are unchanged
$ export PM_LLM_BACKEND="ollama"  PM_OLLAMA_MODEL="gemma4"

# Preview first, apply when it looks right
$ higgs classify --dry-run --limit 20 INBOX

higgs sends the Jev server the first 1,200 characters of each message body. If you install the optional causal-conv1d CUDA kernel in the server's environment, scoring gets faster (0.32 s down to 0.24 s per email on our 3080, at 500 characters).

Privacy note. By default the Jev server and higgs share a machine and talk over 127.0.0.1, so no mail leaves the box. If your GPU is in a different machine, PM_JEV_BASE_URL can point at it, but your mail will then cross your own network to get there. Keep that link on hardware and a network you control.

What we didn't measure

Credits

Jev is originally a closed decision model from TypeSafe AI. The open models we tested are independent reproductions: Open-Jev and its Open-Jev-2B checkpoint, and autotrust/JEV-9B. Thanks to both projects for releasing their work openly. higgs has no affiliation with TypeSafe AI, Open-Jev, or autotrust.


higgs is open source (Apache-2.0) with zero telemetry. Install it, read the README, or browse the blog.