Ran Le
leran1995
AI & ML interests
LLM Pretraining
Recent Activity
updated a collection about 11 hours ago
Nanbeige4.2-3B updated a collection about 12 hours ago
Nanbeige4.2-3B updated a collection about 12 hours ago
Nanbeige4.2-3BOrganizations
Tested on the Artificial Analysis Intelligence Index?
👀 1
1
#25 opened 4 days ago
by
ATHENIS
Pretrain recipe questions for small-model compute budgeting (not data)
👀 1
1
#26 opened 3 days ago
by
Infatoshi
Awesome!
👍 1
1
#24 opened 7 days ago
by
AsThirtyThree
Amazing for its size, but spirals into questionable solutions once it fails something
1
#21 opened 10 days ago
by
Mk2Oracle
Massive activation at layer 21, and three follow-ups on loop depth + LoopSplit
1
#20 opened 12 days ago
by
an0nya
Suitability for 8 GB graphics cards
👀 4
3
#18 opened 12 days ago
by
thermi6
Undocumented config params + two questions on loop depth and length-control RL
1
#16 opened 13 days ago
by
an0nya
Some questions about training
1
#14 opened 14 days ago
by
mobinx
the modal is tiny but kv cache exploding!
👍 6
2
#10 opened 16 days ago
by
rosspanda0
preserve_thinking
5
#9 opened 16 days ago
by
owao
Please specify max context size in model's card.
2
#8 opened 16 days ago
by
Reverger
Any plans for FP8?
2
#5 opened 16 days ago
by
spanspek
Add community evaluation results for CLAW-EVAL, GPQA, HLE, HMMT_FEB_2026, SWE-BENCH_PRO, SWE-BENCH_VERIFIED, TERMINAL-BENCH-2.0
#7 opened 16 days ago
by
nielsr
"Weimplify is asked:"
4
#4 opened 17 days ago
by
owao
Q on the model
4
#3 opened 17 days ago
by
TomLucidor
gguf
➕ 5
4
#1 opened 17 days ago
by
Tapka
damn i was in the process of doing something very similar
1
#2 opened 17 days ago
by
nraxl1
Any plans for a larger scale up? (e.g., 7B - 12B version)
1
#46 opened about 2 months ago
by
rpopreapovle
Update README.md
#38 opened 5 months ago
by
kerasakit