Quaedra Research

Welp

Welp-35B-A3B is Qwen3.6-35B-A3B in 13.21 GB. It runs on a 16 GB GPU with the full 262k context, and scores higher than Unsloth's UD-IQ3_XXS at exactly the same size.

13.21 GBGGUF file
262kcontext on 16 GB
95.7%HumanEval pass@1

How it works

Welp is the same model, quantized more carefully: no fine-tuning and no new training data. It keeps the per-tensor format mix of Unsloth's UD-IQ3_XXS and re-quantizes only the expert weights. Each expert matrix is quantized one 256-weight block at a time with llama.cpp's own quantizer, and each block's rounding error is pushed into the weights not yet quantized (GPTQ, applied block-wise so standard formats can be used). Layers go in order, each fed the output of the already quantized layers.

Quantized from Qwen3.6-35B-A3B (35B total, 3B active, 256 experts, top-8). Experts gate/up IQ2_S, down IQ3_XXS (IQ4_XS in 3 layers), everything else Q6_K, as in UD-IQ3_XXS. Standard llama.cpp formats.

Benchmarks

Fig I: HumanEval pass@1

164 problems, thinking off, temperature 0. Same harness for every model, RTX 4080 Super. Every model here fits on a 16 GB GPU.

Fig II: Decode speed

Tokens per second, llama-bench tg128.

Fig III: Prefill speed

Tokens per second, llama-bench pp4096.

Fig IV: Performance vs size (log scale)

HumanEval pass@1 against GGUF size, same harness and card as Fig I. Points to the upper left are better. The dashed line joins the best Qwen3.6-35B-A3B quantization at each size; Welp moves it up at 13.21 GB.

Show as table
ModelBaseGBHumanEval
Ternary experts (research)Qwen3.6-35B-A3B8.6478.7%
IQ2_XXSQwen3.6-35B-A3B9.5084.8%
Ternary experts, Q2_K down (research)Qwen3.6-35B-A3B9.8290.9%
Welp-35B-A3BQwen3.6-35B-A3B13.2195.7%
UD-IQ3_XXS (Unsloth)Qwen3.6-35B-A3B13.2193.9%
Q4_K_MQwen3.6-35B-A3B21.1795.7%
Bonsai 2 27BQwen3.8-27B5.9591.5%
UD-Q3_K_XLQwen3.8-27B13.1596.3%

Welp ties Q4_K_M (157 of 164) at 62% of its size; Q4_K_M does not fit on a 16 GB card. The two ternary builds are Quaedra research quantizations from the Welp log and are not released. One HumanEval problem is 0.6 points.

ModelGBHumanEvalLong-exactContextDecode tok/sPrefill tok/s
Welp-35B-A3B13.2195.7%23/37262k1645,111
UD-IQ3_XXS (Unsloth)13.2193.9%21/37262k1655,165
Qwen3.8-27B (UD-Q3_K_XL)13.1596.3%26/37131k472,155
Bonsai 2 27B5.9591.5%15/37262k852,278

Single runs on an RTX 4080 Super (16 GB). Context is the largest that fits entirely on the card with a q4_0 KV cache. Long-exact is the bonsai-ada-surgery suite (37 long tool-using tasks, thinking on). Bonsai 2 scores 15/37 here against 17/37 in the bonsai-ada-surgery report (RTX 4070). Qwen3.8-27B is the dense model Bonsai 2 is built from; Welp and UD-IQ3_XXS are Qwen3.6-35B-A3B, so perplexity is only comparable between those two: 5.762 for Welp against 5.873 for UD-IQ3_XXS on wikitext-2 (1.880 against 1.900 on code).

Welp is one problem behind the dense Qwen3.8-27B on HumanEval (157 against 158 of 164) and three tasks behind on long-exact (23 against 26 of 37), at 3.5 times its decode speed and twice its context. It beats UD-IQ3_XXS of its own base model on every quality measure at the same size. The differences against UD-IQ3_XXS on HumanEval and long-exact are small enough to be run-to-run noise on their own; the perplexity gain is consistent.

Full context

With a q4_0 KV cache, Welp runs the full 262k context entirely on a 16 GB card (14.8 GB, including about 0.8 GB used by the desktop). After a 214k-token prompt it still decodes at 107 tok/s, and prefill runs at 2,043 tok/s. A passcode hidden at 10%, 50% and 90% of that prompt was retrieved every time.

Run it

Welp uses only standard llama.cpp formats, the same as UD-IQ3_XXS. The weights go up on Hugging Face as quaedra/Welp-35B-A3B-GGUF; the upload is still pending. Recipe, scripts and evals are on GitHub.

llama-server -m Welp-35B-A3B.gguf -ngl 99 -fa on -c 262144 \
  -ctk q4_0 -ctv q4_0 -ub 256 --jinja

For a more precise KV cache at half the context, use -c 131072 -ctk q8_0 -ctv q8_0.

Thanks to the Qwen team for the base model, Unsloth for the UD-IQ3_XXS format mix, and llama.cpp for the formats.