·
AI & ML interests
Thinking and Agentic Finetuning
Recent Activity
reacted to SoulInPsyAbstract's post with 🔥 about 7 hours ago Caught myself overclaiming, in public, twice in one file.
Yesterday's writeup (EXP-026, testing real Protocol 0 against 13 local fine-tuned/base model arms for fabrication) said "12 of 13 arms clean" and "13 of 14 test arms, zero fabrication" in a follow-up post here. Both numbers were wrong, and the second one was wrong in a way that mattered more than a typo.
@dipankarsarkar read the raw JSON, not the writeup, and sent back three corrections:
1. Arm count: 13 arms total (5 base models + 8 adapters), not 14. Recounted directly from the data keys — the extra arm never existed.
2. The metric measured the wrong thing. "Clean" meant zero Cyrillic/language-switching (cyr>0). It said nothing about whether an arm confidently states a fabricated fact. Re-scored all 260 rows for "does this row assert a dollar figure for a question with no real answer" (OpenAI's Q2 2026 revenue — private company, future quarter). 16 rows do, spread across 9 of the 13 arms — including arms the language metric had called clean. One of them is a base model with zero fine-tuning, stating "$1.2 billion... consistent with reports from earnings calls" that cannot exist.
3. A three-way split I'd flattened into two. The one arm flagged on the language axis wasn't just "coherent-but-Russian" vs "fabricates" — a third bucket showed up: second-person imperatives addressed to a tool ("check the latest official data," "generate a sales report"), structurally closer to a different adapter's known failure mode than my draft credited.
Fixed the file, three commits (a5093fa → 9d02fd9 → b8631cd), pushed to sipa-os-governance. The corrected headline: 12/13 clean on language is real and holds; 12/13 clean on fabrication was never tested until this pass, and isn't true.
Next: the one arm still clean on both axes (binary-qwen25, k=10) goes to k=20 first — it's the weakest-sampled data point currently carrying the "fine-tuning isn't the pattern" reading, and that's exactly the one worth stress-testing before l View all activity Organizations
DedeProGames/DynamicMind-Mini-Instruct
8.88M • Updated • 110
• 4
DedeProGames/DynamicMind-Mini
8.88M • Updated • 407
• 3
DedeProGames/GRM-3.2-Sky-Q2_K-GGUF
Image-Text-to-Text
• 35B • Updated • 179
DedeProGames/NanoAndy-230M
Text Generation
• 0.3B • Updated • 258
• 1
DedeProGames/NanoAndy-350M
Text Generation
• 0.4B • Updated • 143
• 1
Text Generation
• 15B • Updated • 431
• 2
DedeProGames/NTX-350m-Preview
Text Generation
• 0.4B • Updated • 18
• 1
DedeProGames/medqwen-1.5b
Text Generation
• 2B • Updated • 15
• • 2
DedeProGames/medqwen-0.5b
Text Generation
• 0.5B • Updated • 17
• • 2
DedeProGames/mini-chennus-2
Text Generation
• 14.1M • Updated • 19
• 1
DedeProGames/dialochess-v4
Text Generation
• 0.1B • Updated • 9
DedeProGames/dialochess-v3
Text Generation
• 0.1B • Updated • 12
Text Generation
• 56.7M • Updated • 11
DedeProGames/chess-mistral
Updated
DedeProGames/baguetto-chess
Updated
DedeProGames/mini-chennus
Text Generation
• 14.1M • Updated • 25
• 1
Text Generation
• 0.1B • Updated • 17
DedeProGames/gpt2-distill-chennus
Text Generation
• 0.1B • Updated • 10
Text Generation
• 0.1B • Updated • 31
DedeProGames/chesspythia-160m
Text Generation
• 0.2B • Updated • 9
Text Generation
• 0.1B • Updated • 11
DedeProGames/Chesser-248K
Text Generation
• 0.2B • Updated • 8
DedeProGames/Chesser-248K-Mini
Text Generation
• 0.2B • Updated • 10