Instructions to use EldanRing/Winnow-12B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EldanRing/Winnow-12B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-12B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-12B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EldanRing/Winnow-12B:BF16 # Run inference directly in the terminal: llama cli -hf EldanRing/Winnow-12B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EldanRing/Winnow-12B:BF16 # Run inference directly in the terminal: ./llama-cli -hf EldanRing/Winnow-12B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EldanRing/Winnow-12B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EldanRing/Winnow-12B:BF16
Use Docker
docker model run hf.co/EldanRing/Winnow-12B:BF16
- LM Studio
- Jan
- vLLM
How to use EldanRing/Winnow-12B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EldanRing/Winnow-12B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EldanRing/Winnow-12B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/EldanRing/Winnow-12B:BF16
- Ollama
How to use EldanRing/Winnow-12B with Ollama:
ollama run hf.co/EldanRing/Winnow-12B:BF16
- Unsloth Desktop
- Pi
How to use EldanRing/Winnow-12B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf EldanRing/Winnow-12B:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "EldanRing/Winnow-12B:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use EldanRing/Winnow-12B with Docker Model Runner:
docker model run hf.co/EldanRing/Winnow-12B:BF16
- Lemonade
How to use EldanRing/Winnow-12B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EldanRing/Winnow-12B:BF16
Run and chat with the model
lemonade run user.Winnow-12B-BF16
List all available models
lemonade list
- Hermes Agent
How to use EldanRing/Winnow-12B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf EldanRing/Winnow-12B:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default EldanRing/Winnow-12B:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use EldanRing/Winnow-12B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf EldanRing/Winnow-12B:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "EldanRing/Winnow-12B:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Winnow-12B
Winnow-12B is an EldanRing fine-tune of Gemma 4 12B IT for local typed decisions, chat, and image input. Give it a shared state and questions with known answer options; the Winnow server scores those options through /v1/systemone without generating an explanation. Ordinary chat uses /v1/chat/completions from the same loaded model.
The GGUF downloads contain the merged fine-tune. No separate adapter or base-model download is needed.
Quickstart · API reference · Benchmarks
Decision quality
| Model | JevBench public subset, 231 items | Kev-v9 clean, 1,046 items |
|---|---|---|
| Winnow-12B BF16 | 85.28% | 81.45% |
| Winnow-12B Q8 | 85.71% | 81.55% |
| Jev 1.13, hosted via OpenRouter | 85.71% | 87.00% |
Winnow Q8 and Jev each answered 198 of 231 JevBench questions correctly on these frozen inputs. JevBench reports public-subset accuracy rather than the composite leaderboard score. These panels were used during development. The BF16 results use an earlier GGUF export; see export provenance and the evaluation report.
Downloads
Choose one target model. Add the matching projector for image input.
| File | Use | Size |
|---|---|---|
| Winnow-12B-Q8_0.gguf | Q8_0; tested with 64K context and vision | 12.67 GB / 11.80 GiB |
| Winnow-12B-BF16.gguf | BF16; larger-memory systems or CPU/GPU offload | 23.83 GB / 22.20 GiB |
| Winnow-12B-NVFP4.gguf | Smaller Linux/CUDA 8K text and vision presets | 8.16 GB / 7.60 GiB |
| mmproj-Winnow-12B.gguf | F16 vision projector for any of the three targets | 175 MB / 0.163 GiB |
Use the exact target filename or a Winnow preset. Checksums identify every file.
Running the model
The quickstart covers installation, downloads, and launch commands. Direct decisions support noul (yes/no), choice (named options), and score (ordered levels). Multiple questions share a state prefill and can reuse a cached prefix.
F16 is the default and recommended target K/V cache. Cache precision is separate from the GGUF weight format; use --cache q8_0 for an explicit Q8 override. The pinned MTP assistant uses the target's shared K/V cache.
Q8 context and vision
The measured RTX 5070 Ti 16 GB profile used Q8 weights, Q8 KV, full GPU offload, the matching projector, and exclusive memory scheduling.
| Measurement | Result |
|---|---|
| Configured context capacity | 65,536 positions |
| Verified shared prefix with an image | 65,022 positions, including 1,024 image positions |
| Observed peak device VRAM | 15.01 GiB |
| Four questions at near-full context, cold | 25.00 s |
| Same request, cached median of three repeats | 143.0 ms |
| Short-prompt generation, median of three 512-token runs | 55.5 tokens/s |
Context includes formatting, images, questions, and output. Chat and decisions take turns using their K/V contexts in this profile; switching can evict the cached prefix. See full timing definitions for the separate capacity and generation measurements. BF16 weights alone exceed 16 GB VRAM.
NVFP4 tradeoffs
On a matched direct-text workload, Q8 versus NVFP4 used 13,529 versus 9,229 MiB peak device memory and served 3.43 versus 4.98 decisions/s. The RTX 5070 Ti test used four concurrent requests, 4K decision/16K chat context, Q8 KV, and no vision or MTP. Startup was excluded; this is native decision throughput, not generation speed or sustained service capacity.
| Historical matched direct panel | Q8 GGUF | NVFP4 GGUF |
|---|---|---|
| Jev public, 231 decisions | 198/231 (85.71%) | 193/231 (83.55%) |
| Kev-clean, 1,046 decisions | 852/1,046 (81.45%) | 814/1,046 (77.82%) |
| Typed teacher agreement, 2,000 decisions / 400 groups | 1,398/2,000 (69.90%) | 1,412/2,000 (70.60%) |
NVFP4 reduced memory use and improved measured throughput while losing verified-label accuracy on Jev and Kev. Typed measures synthetic teacher agreement. This historical native-T1 comparison used previously observed panels; its Q8 Kev count differs from the separate release campaign above. GGUF results do not transfer to HF/vLLM NVFP4 backends.
Adaptive reasoning and MTP
Direct decisions are the default. Adaptive reasoning generates context before scoring a question again. In the Q8 confirmation, source-equal accuracy/consensus agreement changed 62.50% → 64.06%, with a paired 95% change interval of −2.34 to +5.47 pp. NLL worsened, and mean CLI latency rose from 198 to 743 ms.
MTP separately drafts ordinary chat tokens using the matching 12B assistant. Direct serving needs only the target. Linux/CUDA presets cover Q8 text and NVFP4 vision with MTP at 8K; the Q8 direct 64K vision profile runs without MTP.
See reasoning evaluations, runtime profiles, and assistant attribution for settings and commands.
Training and limitations
Winnow is a rank-32, alpha-64 LoRA fine-tune with zero dropout, merged into the base model before GGUF export. It adapts the attention and MLP projections. The base revision is 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7.
The private training data combines synthetic scenarios, teacher-supervised examples, and labeled semantic tasks, including routing, rule application, evidence selection, ordinal judgments, entailment, and answerability. Training and validation were split. The training data and pipeline are private.
Candidate probabilities are normalized over the supplied options. Confidence describes concentration among those options, not a guarantee of correctness. The reported default decision temperature is 1.0.
Credits and license
Winnow-12B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 12B IT, released under Apache 2.0. See LICENSE and NOTICE.
The separate inference code builds on llama.cpp by Georgi Gerganov and contributors and preserves its MIT license. Jev-style refers to the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.
- Downloads last month
- 43,818
4-bit
8-bit
16-bit
Model tree for EldanRing/Winnow-12B
Base model
google/gemma-4-12B