SuperGPQA is in. Same 1,000 held-out questions, 4 samples each, parent UD-Q4_K_XL vs POCKET, same settings, 16K generation cap. The only bytes that differ are the 300 Q8_0 tensors.
Accuracy (mean of 4):
- parent 59.10%, POCKET 61.55%
- paired difference +2.45 points [+1.50, +3.35]
- single sample 58.70 vs 62.20, majority of 4 63.40 vs 65.40
The cap matters here, so up front: 806 of the parent's 4,000 samples hit 16K, vs 688 for POCKET, and most capped samples give no answer (796 vs 675). On the 667 questions where no sample in either build hit the cap, the difference is +0.64 [0.00, +1.31].
So most of the gain is POCKET finishing within the budget where the parent runs out. Same budget, more answers delivered.
Length: mean 6,716 to 6,051 tokens (-9.9%), geometric-mean ratio 0.868 [0.848, 0.888]; on the uncapped questions the token-weighted ratio is 0.842.
Symmetric deciles (ratio / accuracy parent vs POCKET):
D1 0.91 (82/82), D2 0.91 (76/76), D3 0.83 (80/80), D4 0.80 (74/75), D5 0.79 (81/83), D6 0.81 (74/74), D7 0.86 (62/63), D8 0.86 (44/54), D9 0.95 (19/27), D10 1.00 (1/2)
Two differences from MMLU-Pro. First, the bottom deciles also shrink (about 9%), probably because even the "short" SuperGPQA traces here are 500 to 1,000 tokens, past the ~700 step we saw. Second, D9 and D10 sit at the cap in both builds, so their ratios are compressed toward 1.0, and D8 and D9 are where POCKET's earlier finish turns into the accuracy gap.
The greedy first-commit run is next; we will post it here.