Instructions to use astroanand/astro-code-switch-0.5_7.9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use astroanand/astro-code-switch-0.5_7.9B with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir astro-code-switch-0.5_7.9B astroanand/astro-code-switch-0.5_7.9B
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
astro_code_switch0.5_7.9B
4-bit decoder-only quantization of Gemma 4 E4B, for on-device code-switched (Hindi–English) dictation. Built for Astro.
The audio encoder is left at bf16 on purpose. Only the language decoder is quantized — it is 94% of the weights, so it is the whole lever, and quantizing the encoder costs transcription quality for almost no size saving.
| component | params | share | precision |
|---|---|---|---|
language_model (decoder) |
7.46 B | 94.1% | 4-bit, group 64, affine |
audio_tower (encoder) |
0.305 B | 3.8% | bf16 |
vision_tower (unused for dictation) |
0.167 B | 2.1% | bf16 |
| total | 7.93 B |
Measured results
40 clips, same eval script and prompt, run one at a time on an M5 Pro (48 GB).
| variant | on disk | WER | CER | median latency | RTF |
|---|---|---|---|---|---|
| bf16 baseline | 15.9 GB | 3.59% | 1.41% | 749 ms | 7.3× |
| this model (4-bit decoder) | 4.82 GB | 3.76% | 2.05% | 386 ms | 14.8× |
| 3-bit decoder | 3.96 GB | 260% | 296% | 373 ms | 12.6× |
3.3× smaller and ~2× faster for +0.17pp WER.
The 3-bit row is the useful part: quantization does not degrade gently here, it falls off a cliff. At 3 bits the model enters runaway repetition loops and is unusable. 4-bit is the frontier, not a midpoint.
Why this exists
Parakeet is faster and more accurate on English (1.35% WER, 90 ms) but cannot speak Hindi or Gujarati. Gemma 4's audio encoder is genuinely cross-lingual and holds code-switched speech intact — English stays Latin, Hindi stays Devanagari, in the same utterance:
Paneer stock में available है, या नहीं confirm करो।
That property is what this model is for. The bf16 original is too heavy to keep resident (15.9 GB, 749 ms); this makes it practical on a laptop.
Usage
from mlx_vlm import load, generate
model, processor = load("astroanand/astro_code_switch0.5_7.9B")
Script policy (Devanagari ↔ Roman) is handled after transcription by a deterministic normalizer, not by prompting. Prompt steering was measured and does not work on this family: the 0.6B is steerable but unstable, the 1.7B ignores prompts entirely. Romanization is only safe in the Indic→Latin direction; Latin→Indic is destructive.
Provenance
Quantized from mlx-community/gemma-4-e4b-it-bf16 with mlx_vlm. Reproduction scripts
(01_inspect_split.py, 02_quantize.py, 03_eval.py) and full results are in the
Astro research tree. Original Gemma terms of use apply.
- Downloads last month
- 53
4-bit
Model tree for astroanand/astro-code-switch-0.5_7.9B
Base model
google/gemma-4-E4B