Welp-35B-A3B is Qwen3.6-35B-A3B in 13.21 GB. It runs on a 16 GB GPU with the full 262k context, and scores higher than Unsloth's UD-IQ3_XXS at exactly the same size.
Welp is the same model, quantized more carefully: no fine-tuning and no new training data. It keeps the per-tensor format mix of Unsloth's UD-IQ3_XXS and re-quantizes only the expert weights. Each expert matrix is quantized one 256-weight block at a time with llama.cpp's own quantizer, and each block's rounding error is pushed into the weights not yet quantized (GPTQ, applied block-wise so standard formats can be used). Layers go in order, each fed the output of the already quantized layers.
Quantized from Qwen3.6-35B-A3B (35B total, 3B active, 256 experts, top-8). Experts gate/up IQ2_S, down IQ3_XXS (IQ4_XS in 3 layers), everything else Q6_K, as in UD-IQ3_XXS. Standard llama.cpp formats.
164 problems, thinking off, temperature 0. Same harness for every model, RTX 4080 Super. Every model here fits on a 16 GB GPU.
Tokens per second, llama-bench tg128.
Tokens per second, llama-bench pp4096.
HumanEval pass@1 against GGUF size, same harness and card as Fig I. Points to the upper left are better. The dashed line joins the best Qwen3.6-35B-A3B quantization at each size; Welp moves it up at 13.21 GB.
| Model | Base | GB | HumanEval |
|---|---|---|---|
| Ternary experts (research) | Qwen3.6-35B-A3B | 8.64 | 78.7% |
| IQ2_XXS | Qwen3.6-35B-A3B | 9.50 | 84.8% |
| Ternary experts, Q2_K down (research) | Qwen3.6-35B-A3B | 9.82 | 90.9% |
| Welp-35B-A3B | Qwen3.6-35B-A3B | 13.21 | 95.7% |
| UD-IQ3_XXS (Unsloth) | Qwen3.6-35B-A3B | 13.21 | 93.9% |
| Q4_K_M | Qwen3.6-35B-A3B | 21.17 | 95.7% |
| Bonsai 2 27B | Qwen3.8-27B | 5.95 | 91.5% |
| UD-Q3_K_XL | Qwen3.8-27B | 13.15 | 96.3% |
Welp ties Q4_K_M (157 of 164) at 62% of its size; Q4_K_M does not fit on a 16 GB card. The two ternary builds are Quaedra research quantizations from the Welp log and are not released. One HumanEval problem is 0.6 points.
| Model | GB | HumanEval | Long-exact | Context | Decode tok/s | Prefill tok/s |
|---|---|---|---|---|---|---|
| Welp-35B-A3B | 13.21 | 95.7% | 23/37 | 262k | 164 | 5,111 |
| UD-IQ3_XXS (Unsloth) | 13.21 | 93.9% | 21/37 | 262k | 165 | 5,165 |
| Qwen3.8-27B (UD-Q3_K_XL) | 13.15 | 96.3% | 26/37 | 131k | 47 | 2,155 |
| Bonsai 2 27B | 5.95 | 91.5% | 15/37 | 262k | 85 | 2,278 |
Single runs on an RTX 4080 Super (16 GB). Context is the largest that fits entirely on the card with a q4_0 KV cache. Long-exact is the bonsai-ada-surgery suite (37 long tool-using tasks, thinking on). Bonsai 2 scores 15/37 here against 17/37 in the bonsai-ada-surgery report (RTX 4070). Qwen3.8-27B is the dense model Bonsai 2 is built from; Welp and UD-IQ3_XXS are Qwen3.6-35B-A3B, so perplexity is only comparable between those two: 5.762 for Welp against 5.873 for UD-IQ3_XXS on wikitext-2 (1.880 against 1.900 on code).
Welp is one problem behind the dense Qwen3.8-27B on HumanEval (157 against 158 of 164) and three tasks behind on long-exact (23 against 26 of 37), at 3.5 times its decode speed and twice its context. It beats UD-IQ3_XXS of its own base model on every quality measure at the same size. The differences against UD-IQ3_XXS on HumanEval and long-exact are small enough to be run-to-run noise on their own; the perplexity gain is consistent.
With a q4_0 KV cache, Welp runs the full 262k context entirely on a 16 GB card (14.8 GB, including about 0.8 GB used by the desktop). After a 214k-token prompt it still decodes at 107 tok/s, and prefill runs at 2,043 tok/s. A passcode hidden at 10%, 50% and 90% of that prompt was retrieved every time.
Welp uses only standard llama.cpp formats, the same as UD-IQ3_XXS. The weights go up on Hugging Face as quaedra/Welp-35B-A3B-GGUF; the upload is still pending. Recipe, scripts and evals are on GitHub.
llama-server -m Welp-35B-A3B.gguf -ngl 99 -fa on -c 262144 \ -ctk q4_0 -ctv q4_0 -ub 256 --jinja
For a more precise KV cache at half the context, use -c 131072 -ctk q8_0 -ctv q8_0.
Thanks to the Qwen team for the base model, Unsloth for the UD-IQ3_XXS format mix, and llama.cpp for the formats.