Non-uniform GGUF quantizations via GSQ + RCO: per-tensor mixed precision in standard GGUF form
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
11:04 AM
RCO: Model Compression with Exact Budget Constraints via Riemannian Manifolds, https://huggingface.co/papers/2605.00649
High-quality QAT FP4 models to use with the fp_quant vLLM/Transformers integration on Blackwell NVIDIA GPUs. See https://arxiv.org/abs/2509.23202
MXFP4 and NVFP4 quantized models
GPTQ-quantized Gemma3 models
Models prequantized with [HIGGS](https://arxiv.org/abs/2411.17525) zero-shot quantization. Requires the latest `transformers` to run.
-
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Paper • 2411.17525 • Published • 6 -
ISTA-DASLab/Llama-3.3-70B-Instruct-HIGGS-GPTQ-4bit
19B • Updated • 9 • 7 -
ISTA-DASLab/Llama-3.1-8B-Instruct-HIGGS-GPTQ-4bit
Text Generation • 3B • Updated • 9 -
ISTA-DASLab/Llama-3.1-8B-Instruct-HIGGS-GPTQ-3bit
Text Generation • 2B • Updated • 8
AQLM quantized LLMs
-
Extreme Compression of Large Language Models via Additive Quantization
Paper • 2401.06118 • Published • 15 -
ISTA-DASLab/Meta-Llama-3-70B-Instruct-AQLM-2Bit-1x16
Text Generation • 11B • Updated • 34 • 20 -
ISTA-DASLab/Meta-Llama-3-70B-AQLM-2Bit-1x16
Text Generation • 11B • Updated • 33 • 14 -
ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16
Text Generation • 2B • Updated • 33 • 12
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling, https://huggingface.co/papers/2604.18556
-
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Paper • 2604.18556 • Published • 11 -
ISTA-DASLab/Kimi-K2.6-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 24 -
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 65 • 1 -
ISTA-DASLab/Llama-3.1-70B-Instruct-2Bit-GSQ
Text Generation • 7B • Updated • 21
MatGPTQ quantized models
DASLab support for GGUF
https://arxiv.org/abs/2502.05003
Official AQLM quantizations for "PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression": https://arxiv.org/abs/2405.14852
-
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Paper • 2405.14852 • Published • 2 -
ISTA-DASLab/Meta-Llama-3.1-70B-Instruct-AQLM-PV-2Bit-1x16
Text Generation • 11B • Updated • 24 • 45 -
ISTA-DASLab/Mistral-Nemo-Instruct-2407-AQLM-PV-2Bit-1x16-hf
3B • Updated • 8 • 3 -
ISTA-DASLab/Meta-Llama-3.1-8B-Instruct-AQLM-PV-2Bit-1x16-hf
Text Generation • 2B • Updated • 23 • 8
Non-uniform GGUF quantizations via GSQ + RCO: per-tensor mixed precision in standard GGUF form
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling, https://huggingface.co/papers/2604.18556
-
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Paper • 2604.18556 • Published • 11 -
ISTA-DASLab/Kimi-K2.6-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 24 -
ISTA-DASLab/Kimi-K2.5-2Bit-GSQ
Image-Text-to-Text • 84B • Updated • 65 • 1 -
ISTA-DASLab/Llama-3.1-70B-Instruct-2Bit-GSQ
Text Generation • 7B • Updated • 21
11:04 AM
RCO: Model Compression with Exact Budget Constraints via Riemannian Manifolds, https://huggingface.co/papers/2605.00649
High-quality QAT FP4 models to use with the fp_quant vLLM/Transformers integration on Blackwell NVIDIA GPUs. See https://arxiv.org/abs/2509.23202
MatGPTQ quantized models
MXFP4 and NVFP4 quantized models
DASLab support for GGUF
GPTQ-quantized Gemma3 models
https://arxiv.org/abs/2502.05003
Models prequantized with [HIGGS](https://arxiv.org/abs/2411.17525) zero-shot quantization. Requires the latest `transformers` to run.
-
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Paper • 2411.17525 • Published • 6 -
ISTA-DASLab/Llama-3.3-70B-Instruct-HIGGS-GPTQ-4bit
19B • Updated • 9 • 7 -
ISTA-DASLab/Llama-3.1-8B-Instruct-HIGGS-GPTQ-4bit
Text Generation • 3B • Updated • 9 -
ISTA-DASLab/Llama-3.1-8B-Instruct-HIGGS-GPTQ-3bit
Text Generation • 2B • Updated • 8
Official AQLM quantizations for "PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression": https://arxiv.org/abs/2405.14852
-
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Paper • 2405.14852 • Published • 2 -
ISTA-DASLab/Meta-Llama-3.1-70B-Instruct-AQLM-PV-2Bit-1x16
Text Generation • 11B • Updated • 24 • 45 -
ISTA-DASLab/Mistral-Nemo-Instruct-2407-AQLM-PV-2Bit-1x16-hf
3B • Updated • 8 • 3 -
ISTA-DASLab/Meta-Llama-3.1-8B-Instruct-AQLM-PV-2Bit-1x16-hf
Text Generation • 2B • Updated • 23 • 8
AQLM quantized LLMs
-
Extreme Compression of Large Language Models via Additive Quantization
Paper • 2401.06118 • Published • 15 -
ISTA-DASLab/Meta-Llama-3-70B-Instruct-AQLM-2Bit-1x16
Text Generation • 11B • Updated • 34 • 20 -
ISTA-DASLab/Meta-Llama-3-70B-AQLM-2Bit-1x16
Text Generation • 11B • Updated • 33 • 14 -
ISTA-DASLab/Meta-Llama-3-8B-Instruct-AQLM-2Bit-1x16
Text Generation • 2B • Updated • 33 • 12