Quaedra Research

Jet

Jet-4B is the v6 line. Versions up to V5 were smaller 0.6B Jet models; Jet-4B continues their prompt format and data. Those runs use Qwen3-0.6B unless stated otherwise. Results depend on the specified splits and inference backend; partial benchmark samples are not official Decision Index scores. Historical reports use “released Jet” to mean the model published at the time of that experiment, not necessarily today's release.

Initial experiments and V2

Public-source training introduced choice, score and yes/no decisions, followed by varied-scale ordinal questions. V2 used two epochs and 2,910 steps, selecting step 2,750 by validation NLL. The fused release reached about 67.7% accuracy on the 3,683-row general test. A Qwen3-1.7B comparison was evaluated but not retained.

The complete original tables, earlier 1,500-row experiment, ordinal evaluation, and Kev transfer evaluation are preserved in initial results.

V3 — broader decision training

Added relevance, entailment, stance, sarcasm, tool selection, Boolean rules, arithmetic and code behavior. Trained on 47,313 rows for two epochs / 4,888 steps with rank-16 LoRA and learning rate 1e-4. New-family accuracy reached 80.6%; general-test accuracy regressed to 65.2%.

Experiment report

Warm start and V4 — transfer and routing

A lower-learning-rate warm start was followed by transfer and routing training. V4 reached 66.4% general-test accuracy and 93.3% on domain-held-out tool routing. It improved several partial language/relevance diagnostics but regressed on code. A CPU-fused bf16/q8 release package was prepared separately.

Training and comparisons · Export preparation

V5 — wider intent and adversarial language tasks

Continued from V4 with 7,997 retained examples and 2,000 each from native training partitions of CLINC150, ANLI, HellaSwag and When2Call preference data. The completed R2 run used 15,997 rows, one epoch / 3,570 steps, rank 16, learning rate 1e-5, and selected step 3,500. Calibration was separate.

The first attempt was stopped after identifying retained utterances shared with native evaluation content. R2 removes three such utterances from continuation training, but does not erase inherited cross-dataset overlap in its initialization. The checkpoint is not certified decontaminated.

Held-out accuracy V4 V5
General test 66.44% 66.66%
New families 59.00% 73.50%
Transfer tasks 65.97% 67.50%
Domain-held-out tool routing 93.33% 85.00%

Selection-only evaluation chose original/reversed option-order averaging for the benchmark adapter. This increases inference work and is not part of standard serving. Gains and regressions on fixed public benchmark samples are recorded separately; no overall leaderboard result or win has been established.

Results · Protocol · Completion audit

V5 release — 2026-09-24

Superseded the same day by v6. The V5 R2 selected checkpoint was released fused with native MLX CUDA, with a calibrated bf16 model and validated q8 ONNX export. Its files remain at model-repository revision 25ccbd9e09c75643b3c2214e2b2522bec39171a7. Export and publication receipts are recorded in release notes.

V6 — Jet-4B — 2026-09-24

Moved the backbone to Qwen3.5-4B with a fresh rank-16 LoRA (learning rate 1e-4) trained with PyTorch/PEFT on the same 15,997-row train_v5_r2 mixture, one epoch / 4,000 updates, selecting step 3,750. The adapter was merged into full bf16 weights. Selection accuracy was 87.93% (1,400 rows) and independent local test accuracy 94.00% (600 rows). These splits differ from the V2–V5 tables above, so the numbers are not directly comparable. The official Decision Index was not measured.

Superseded by v6.1 below.

Release notes · Model card

V6.1 — targeted continuation — 2026-09-25

The released model continues from the full merged v6 backbone with a correction LoRA on 22,643 examples. Validation selected step 2,000 of 5,661. Two later repair trials were rejected by their sarcasm/retention guards and were not released. The release is a complete merged BF16 model in the existing quaedra/jet repository, with the product name Jet. The previous release is archived as v6.0.0.

The broad development evaluation covers 25 benchmarks and 67,459 requests, including two sampled retrieval datasets. These are pre-merge adapter results; 144 paired merge checks preserved all selected answers, with probability changes up to 7.04 points. No official overall index score is claimed.

Release card · Benchmark report · Repair-trial report

v6.2.0 — focused continuation, 2026-09-26

Full merged continuation from v6.1, selected at step 250 of the 3e-6 trial. The 8e-6 trial kept its step-zero fallback and is not released. Both ran 1,000 updates on 4,000 examples: 800 banking, 800 entity sentiment, 400 sarcasm/literal, and 2,000 retention. Protocol, data hashes, environment and trial records are in experiments/jet-focused-20260925/.

Local holdout Jet v6.1 Jet v6.2 merged
Banking accuracy 74.68 74.68
Entity sentiment macro-F1 71.16 71.21
Sarcasm F1 46.81 50.00
Retention accuracy 93.40 93.40

These 914-case holdout measurements were repeated on the full merged weights. Focus holdouts are new local splits; retention is reused. They are not Decision Index scores. Financial sentiment measures SEntFiN transfer, not FinEntity. The adapter's 71.72 financial macro-F1 became 71.21 after BF16 merging.

The 157 merge-verification cases had no argmax flips and at most 4.76 percentage points of probability drift. The standalone runtime now accepts 16,384-token complete prompts and passed the longest API-Bank input (11,495 tokens). Calibration is inherited from v6.1, not refitted. Previous benchmark charts remain explicitly attributed to v6.1.

Decision Index 0.3 (published 2026-10-07) measured v6.2 officially: 40.01, rank 52 of 113; public 42.17, same skills 40.93, new domains 34.46. See the Decision Index notes.

Edit this page on GitHub