I appreciate your merges, @nightmedia - even if we have to fix the MTP 😃
Nick M
veldierin
AI & ML interests
None yet
Recent Activity
new activity 2 days ago
peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF-MTP:Agentic Coding using OpenCode facing the issueOrganizations
None yet
Keep us updated on this recipe, it looks interesting
❤️ 1
6
#2 opened 2 days ago
by
veldierin
Agentic Coding using OpenCode facing the issue
4
#9 opened 2 days ago
by
SolutionsDealer
New activity in DavidAU/Qwen3.8-27B-Cold-Fable-Fusion-GAIN-V1.1-732-Heretic-Uncensored-stage1 3 days ago
Gemini trace analysis
12
#1 opened 4 days ago
by
nightmedia
MTP enabled: unbounded VRAM growth during inference until CUDA abort (V100)
4
#4 opened 4 days ago
by
fozosan
Model stops answering prematurely
➕ 1
2
#1 opened 6 days ago
by
theblackcat
Great model so far
❤️🔥 3
6
#17 opened 8 days ago
by
mythrime
request for re-quantization w/fixed MTP heads
👍 1
4
#3 opened 5 days ago
by
veldierin
How it compares to Ornith 1.5 35B?
2
#2 opened 5 days ago
by
MartinPatterson
Mtp quality? 🤔
5
#7 opened 13 days ago
by
LinkuStarto
APEX quant
11
#1 opened 10 days ago
by
benoe
Waiting for Ornith-1.5-35B-Heretic-MTP-APEX-GGUF
3
#2 opened 8 days ago
by
atfa
VERSION NVFP4 : Would a nvfp4 version possible ?
➕ 1
10
#14 opened 11 days ago
by
crazyhenres
Testing it on greenboost-cli with good results
❤️👍 2
8
#15 opened 9 days ago
by
hyphaed
replied to nightmedia's post 9 days ago
reacted to nightmedia's post with 👍 10 days ago
Post
4177
Qwen3.8-27B metrics
It's hard to track all model cards where I post these, so I figured people would get more value out of seeing these in the open.
The performance is as measured on a M4 MBP 128GB, speed may vary depending on your platform.
These are all instruct metrics, generated by including this line in the jinja template:
Then run the test suite to generate the metrics:
This will generate the file:
This is a JSON containing all gathered metrics; for example the q4-hi:
I use the value of acc_norm for metrics, rounded to 3 decimals.
As I get more quants tested, I will add them here.
A complete test run for a single quant takes 7-9 hours depending on quant size, 10-12 hours for BF16 depending on perplexity: this is why you see on my model cards that I usually post the first three, that only take 2-3 hours :)
-G
It's hard to track all model cards where I post these, so I figured people would get more value out of seeing these in the open.
quant arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.591,0.782,0.896,0.746,0.448,0.801,0.711
q8-hi 0.602,0.779,0.896,0.747,0.446,0.793,0.703
q6-hi 0.602,0.775,0.895,0.748,0.448,0.795,0.710
q4-hi 0.604,0.780,0.898,0.744,0.454,0.795,0.708
mxfp4 0.581,0.771,0.889,0.738,0.442,0.798,0.713
1M
mxfp8 0.590,0.787,0.897,0.744,0.446,0.801,0.709
Quant Perplexity Peak Memory Tokens/sec
mxfp8 6.090 ± 0.054 34.74 GB 138
mxfp4 5.952 ± 0.051 21.30 GB 148The performance is as measured on a M4 MBP 128GB, speed may vary depending on your platform.
These are all instruct metrics, generated by including this line in the jinja template:
{%- set enable_thinking = false %}Then run the test suite to generate the metrics:
mlx_lm.evaluate --model MODEL --tasks winogrande boolq arc_challenge arc_easy hellaswag openbookqa piqaThis will generate the file:
eval_MODEL_0.4.9_winogrande_boolq_arc_challenge_arc_easy_hellaswag_openbookqa_piqaThis is a JSON containing all gathered metrics; for example the q4-hi:
"arc_challenge": {
"alias": "arc_challenge",
"acc,none": 0.5819112627986348,
"acc_stderr,none": 0.014413988396996116,
"acc_norm,none": 0.6040955631399317,
"acc_norm_stderr,none": 0.01429122839353657
},I use the value of acc_norm for metrics, rounded to 3 decimals.
As I get more quants tested, I will add them here.
A complete test run for a single quant takes 7-9 hours depending on quant size, 10-12 hours for BF16 depending on perplexity: this is why you see on my model cards that I usually post the first three, that only take 2-3 hours :)
-G
Introducing Unsloth Dynamic v3 Qwen3.8
👍🔥 52
34
#74 opened 10 days ago
by
danielhanchen
Brainwaves MTP Head: Recipe-Reconstruction Approach
🧠❤️ 3
3
#2 opened 10 days ago
by
veldierin
Status on repo - ready to test now
🚀❤️ 1
10
#1 opened 10 days ago
by
veldierin