VIDRAFT_LAB
AI & ML interests
Recent Activity
Organizations
Organizer Discovery downfold log (auto-updated)
채점 지연 안내와 현재 상태 (Scoring delay: resolved)
Results from the greedy first-commit run. MMLU-Pro, 250 questions (200 from deciles 6 to 10 on the symmetric cut, 50 from deciles 1 to 5 as a control), T=0, 131K cap, parent UD-Q4_K_XL vs POCKET, both 250/250. Tokens counted with the parent's tokenizer (identical IDs to llama-tokenize).
First commit = the first point the trace states a final option letter (rule and per-question index linked below).
Deciles 6 to 10 (200 questions), parent to POCKET:
- tokens before the first commit: 3,245 to 3,287 (POCKET slightly longer)
- tokens after the first commit: 5,053 to 3,536 (-1,517)
- share of the total saving that falls after the first commit: 112% [67, 215] (above 100% because the pre-commit part is slightly longer in POCKET)
- first commit sits at a median 0.62 vs 0.67 of the trace, inside the thinking block 90.5% vs 92.0%
- wait / verify / double-check type phrases per question: 14.3 to 11.0
- first commit equals final answer: 88.5% vs 89.0%; same first-commit letter in both builds: 82.9%
- accuracy: 76.5% vs 80.0% (single greedy pass, for reference only)
Control, deciles 1 to 5 (50 questions): 304 to 324 tokens, 54 to 59 after commit, same first commit on all 50.
So your re-verification hypothesis holds: both builds take the same road to their first answer, and POCKET spends less time re-checking it afterwards. Short answers rarely re-check, which is why the bottom half does not move.
One assumption did not hold: the traces do not run together for long. They diverge early (median first divergent token 26), so "divergence before or after the commit" does not discriminate here; the before/after-commit lengths carry the result.
Summary: https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF/blob/main/analysis/first_commit/fc_mmlu_analysis_2026-10-04.md
Per-question index (250 rows): https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF/blob/main/analysis/first_commit/fc_mmlu_commit_index.jsonl
Commit rule and token counting: https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF/blob/main/analysis/first_commit/fc_mmlu_commit_rule.txt
SuperGPQA is in. Same 1,000 held-out questions, 4 samples each, parent UD-Q4_K_XL vs POCKET, same settings, 16K generation budget. The only bytes that differ are the 300 Q8_0 tensors.
Accuracy:
- mean of 4: parent 59.10%, POCKET 61.55%, paired difference +2.45 [+1.54, +3.36]
- single sample: 58.70 vs 62.20
- majority of 4: 63.60 vs 65.90 (over answered samples; ties go to the answer that appears first in sample order, seeds 21 to 24)
The budget matters here: 798 of the parent's 4,000 samples reached the 16K cap, vs 675 for POCKET. Much of the gain is POCKET finishing within the budget where the parent runs out. Same budget, more answers delivered.
Length: mean 6,716 to 6,051 tokens (-10% in total), geometric-mean per-question ratio 0.868 [0.848, 0.888].
Symmetric deciles (length ratio / accuracy parent vs POCKET):
D1 0.91 (82/82), D2 0.91 (76/76), D3 0.83 (80/80), D4 0.80 (74/75), D5 0.79 (81/83), D6 0.81 (74/74), D7 0.86 (62/63), D8 0.86 (44/54), D9 0.95 (19/27), D10 1.00 (1/2)
Two differences from MMLU-Pro: the bottom deciles also shrink, probably because even the short SuperGPQA traces are already past the ~700-token step; and D9 to D10 sit at the cap in both builds, so their ratios are compressed toward 1.0. D8 and D9 are where finishing earlier turns into the accuracy gap.
POCKET-Darwin-180B: a clean R0 vs R3 test on held-out SuperGPQA (+2.45 points)
Would anyone be willing to create a quantized version of Darwin-180B-RSI?
SuperGPQA is in. Same 1,000 held-out questions, 4 samples each, parent UD-Q4_K_XL vs POCKET, same settings, 16K generation cap. The only bytes that differ are the 300 Q8_0 tensors.
Accuracy (mean of 4):
- parent 59.10%, POCKET 61.55%
- paired difference +2.45 points [+1.50, +3.35]
- single sample 58.70 vs 62.20, majority of 4 63.40 vs 65.40
The cap matters here, so up front: 806 of the parent's 4,000 samples hit 16K, vs 688 for POCKET, and most capped samples give no answer (796 vs 675). On the 667 questions where no sample in either build hit the cap, the difference is +0.64 [0.00, +1.31].
So most of the gain is POCKET finishing within the budget where the parent runs out. Same budget, more answers delivered.
Length: mean 6,716 to 6,051 tokens (-9.9%), geometric-mean ratio 0.868 [0.848, 0.888]; on the uncapped questions the token-weighted ratio is 0.842.
Symmetric deciles (ratio / accuracy parent vs POCKET):
D1 0.91 (82/82), D2 0.91 (76/76), D3 0.83 (80/80), D4 0.80 (74/75), D5 0.79 (81/83), D6 0.81 (74/74), D7 0.86 (62/63), D8 0.86 (44/54), D9 0.95 (19/27), D10 1.00 (1/2)
Two differences from MMLU-Pro. First, the bottom deciles also shrink (about 9%), probably because even the "short" SuperGPQA traces here are 500 to 1,000 tokens, past the ~700 step we saw. Second, D9 and D10 sit at the cap in both builds, so their ratios are compressed toward 1.0, and D8 and D9 are where POCKET's earlier finish turns into the accuracy gap.
The greedy first-commit run is next; we will post it here.
That is a clean way to split it, and the re-verification hypothesis fits what we see: short answers barely change, long ones drop by a step.
We will run it as you describe: greedy decoding on the same prompts, parent vs POCKET, on a sample from deciles 6 to 10 plus a small control from deciles 1 to 5. For each question we will report where the first divergent token sits relative to the first answer commit, whether both builds commit to the same first answer, and how many tokens follow that commit in each.
Results will go up here together with SuperGPQA.
Good question. From the symmetric cut, deciles 6 to 10 (geometric-mean ratio, roughly 700 to 78K tokens):
D6 0.872, D7 0.867, D8 0.857, D9 0.741, D10 0.876
So it looks more like a step than a slope. The ratio drops from about 1.00 to about 0.87 around 700 tokens, then stays in the 0.74 to 0.88 range without a steady decline; the longest decile (D10) comes back up to 0.876. D9 is the one dip, and with ~200 questions per decile we would not read a trend into a single bin.
We will check whether the same step shows up on SuperGPQA, where many more traces run past 700 tokens, and include the per-decile table with the cap counts.
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
📦 360 GB → 111 GB (4-bit GGUF, 4 files)
🖥️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s
💻 RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
🧊 128 GB mini PC: whole model in memory, no GPU needed
🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How?
· Only ~3B of 180B parameters are active per token (10 of 512 experts)
· llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough
· Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.
📝 Article: https://huggingface.co/blog/FINAL-Bench/data-center-ai-now-on-a-laptop-pocket-darwin-180b
🤗 Model: FINAL-Bench/POCKET-Darwin-180B-GGUF
🧬 Original: FINAL-Bench/Darwin-180B-RSI
#Darwin #RSI #GGUF #llamacpp #OnDevice #MoE
Thanks for reconciling the cells. Here are the deciles, with one caveat about how to cut them.
Cut by parent length, as you asked, the top decile does fall: POCKET/parent ratio 0.756 [0.687, 0.829] (geometric mean 0.611), vs 1.008 for the bottom nine combined.
But that cut is biased. Sorting on one build's length puts the questions where that build happened to run long into the top decile, so the ratio there is pulled down by regression to the mean. Cut the same data by POCKET length and the top decile flips to 1.027 (geometric 1.177).
So we sorted on the geometric mean of the two lengths, which treats both builds the same:
- Deciles 1 to 5 (up to ~700 tokens): ratio 0.97 to 1.03, essentially unchanged
- Deciles 6 to 10 (~700 tokens and up): 0.74 to 0.89, i.e. 11 to 26% shorter
- Top decile vs bottom nine: +0.088 [-0.009, +0.187] (mean ratio), -0.041 [-0.12, +0.043] (geometric). No separate drop in the top decile.
So it is neither a top-decile effect nor a flat ~15% shrink. Short answers stay the same length; answers past roughly 700 tokens get 11 to 26% shorter. The typical per-question reduction (geometric mean ratio) is 8.7% [6.2, 11.2]; the 14.5% headline is the token-weighted total, carried by the long questions. Accuracy tracks the parent within ±2 points in every decile.
For SuperGPQA we will report both the parent-length and symmetric cuts, plus how many samples hit the 16K cap in each build, since capping can shrink the ratio. Results when both builds finish.