Runs with DeepswapLLM

Run this model on a GPU too small to hold it โ€” full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.

Mistral-Small-24B-Instruct-Jbliterated

Drop-in replacement for mistralai/Mistral-Small-24B-Instruct-2501 with refusal behaviors surgically removed at the weight level. No system prompt tricks, no inference-time patches. The weights themselves no longer encode refusal.

Method

SVD multi-direction abliteration โ€” instead of removing a single refusal vector (which leaves deeper noncompliance strategies intact), we decompose the harmful-vs-harmless activation space into its principal components via SVD and remove the top 5 orthogonal directions across all 40 transformer layers. This captures 79โ€“93% of the contrastive variance per layer, eliminating both surface refusal and deeper evasion behaviors.

Setting Value
Method SVD multi-direction abliteration
Directions 5 per layer
Layers All 40
Multiplier 2.0
Null-space constraints Enabled (preserves math/coding/reasoning)
Norm preservation Enabled

What This Fixes

Standard (single-direction) abliteration removes the surface "I can't help with that" response but leaves deeper behavioral directions intact. The model finds creative workarounds:

  • Prompt reinterpretation โ€” steering toward a safer reading of the question
  • Disclaimer injection โ€” answering but wrapping in warnings
  • Strategic omission โ€” leaving out the key details
  • Safer framing โ€” answering a related but less harmful version

SVD multi-direction abliteration eliminates all of these noncompliance strategies.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated",
    torch_dtype=torch.float16,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated")

Requirements

  • Base model: mistralai/Mistral-Small-24B-Instruct-2501

License

apache-2.0


Apollo Raines builds post-training tools that separate behavior from knowledge and identity from architecture.

Downloads last month
1,015
Safetensors
Model size
24B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated

Quantized
(108)
this model
Quantizations
2 models

Space using ApolloRaines/Mistral-Small-24B-Instruct-Jbliterated 1