A typed decision model. Give it a state and named questions; it returns choices, scores and probabilities, never free-form text.
GitHub Weights HF Space Training history
Each answer option maps to a single label token. Jet reads the next-token logits for those labels only, divides by a temperature fitted on held-out data, and applies softmax. One forward pass per question, no sampling. Answers always follow the requested type, but decisions can still be wrong.
| Type | Criteria | Answer |
|---|---|---|
choice | 2–255 named options | Selected key, probability per key |
score | 2–10 ordered levels | Fractional score, level, probabilities |
noul | None | Probability that the answer is yes |
Difference from the archived Decision Index 0.1 reference score on 21 benchmarks where the metric and case count match.
Development evaluation of the v6.1 step-2,000 adapter before its BF16 merge, not v6.2. Kev figures are archived published scores, not a new local run. Matching counts do not prove identical cases. No official overall Decision Index has been measured for Jet. Full report
914 fixed cases, full merged BF16 weights.
Small changes; not statistically significant. Financial sentiment measures SEntFiN transfer, not FinEntity.
Real outputs recorded in the model's reference file (Jet V5, Qwen3-0.6B, calibrated). Pick a case to see the request and the typed answer.
For live inference, run the model locally or start the Hugging Face Space, which serves the earlier 0.6B release and may be paused on free hosting.
The v6.2 release is self-contained: merged bf16 weights plus the CUDA runtime.
hf download michaljach/jet --revision v6.2.0 --local-dir jet
cd jet && python -m pip install -r requirements.txt
echo '{"state":"I was charged twice this month.",
"questions":{"billing":{"type":"noul",
"instructions":"Is this a billing issue?"}}}' | python jet.py