Ivan Baldo
ibaldonl
AI & ML interests
MLOps, Scalability, Performance, OnPremises
Recent Activity
new activity 6 days ago
Qwen/Qwen3.8-27B:NVFP4 Shootout (Quality and Speed) liked a dataset 12 days ago
llamaindex/ExtractBench liked a model 22 days ago
bartowski/apodex_Apodex-1.1-mini-GGUFOrganizations
None yet
NVFP4 Shootout (Quality and Speed)
❤️➕ 7
10
#192 opened 22 days ago
by
Rieker
Missing vLLM FP8 KVCache calibration
#3 opened 23 days ago
by
ibaldonl
Maybe add vLLM FP8 KVCache calibration scales?
4
#3 opened 28 days ago
by
ibaldonl
FP8 vLLM KVCache calibration is missing
5
#2 opened 29 days ago
by
ibaldonl
Maybe add vLLM FP8 KVCache calibration scales
3
#2 opened 29 days ago
by
ibaldonl
vLLM KVCache FP8 calibration data missing
➕ 1
1
#7 opened 29 days ago
by
ibaldonl
Doesn't have FP8 VLLM KVCache calibration scales
➕ 1
3
#3 opened 29 days ago
by
ibaldonl
Plans to upstream this patch in vLLM?
1
#2 opened 29 days ago
by
ibaldonl
Maybe run the exact same eval setup comparing against nvidia/Qwen3.6-27B-NVFP4?
#8 opened about 2 months ago
by
ibaldonl
16gb card owners: post your numbers
2
#4 opened 4 months ago
by
deucebucket
Github repo doesn't exist anymore
1
#1 opened 4 months ago
by
ibaldonl
hoping to add multi-language support
👀 1
2
#2 opened 6 months ago
by
TIN-HF
Qwen/Qwen3.6-35B-A3B-GPTQ-Int4?
➕ 8
4
#9 opened 6 months ago
by
sujithr
30.3 GB?
👀 4
3
#6 opened 7 months ago
by
pedalnomica
Is it good for languages other than Chinese?
#1 opened 7 months ago
by
ibaldonl
The one done with GPTQ seems better than this one.
3
#1 opened 8 months ago
by
ibaldonl
Add Qwen3-VL-30B-A3B-Instruct-NVFP4 and -quantized.w4a16
#2 opened 11 months ago
by
ibaldonl
Rerank (Score) API
👍 5
2
#1 opened over 1 year ago
by
alexhyzheng
Unable to add model.
1
#1 opened over 1 year ago
by
Sakalti
meta-llama/Meta-Llama-3-8B-Instruct model links are mixed up
4
#1101 opened over 1 year ago
by
SandInTheDunes