llamma.cpp support

#2
by RandomUserNA12312 - opened

Curious if this was in progress and if there was a MR on github we could track on the llamma.cpp repo?

For Current Llama master branch it would be the attached diff @RandomUserNA12312 :

diff --git a/src/llama-arch.cpp b/src/llama-arch.cpp
index 836cfade2..d170e3e69 100644
--- a/src/llama-arch.cpp
+++ b/src/llama-arch.cpp
@@ -40,6 +40,7 @@ static const std::map<llm_arch, const char *> LLM_ARCH_NAMES = {
     { LLM_ARCH_QWEN3VLMOE,       "qwen3vlmoe"       },
     { LLM_ARCH_QWEN35,           "qwen35"           },
     { LLM_ARCH_QWEN35MOE,        "qwen35moe"        },
+    { LLM_ARCH_MAPLE,            "maple"            },
     { LLM_ARCH_PHI2,             "phi2"             },
     { LLM_ARCH_PHI3,             "phi3"             },
     { LLM_ARCH_PHIMOE,           "phimoe"           },
diff --git a/src/llama-arch.h b/src/llama-arch.h
index 49c2a6ac3..25a5d37e0 100644
--- a/src/llama-arch.h
+++ b/src/llama-arch.h
@@ -45,6 +45,7 @@ enum llm_arch {
     LLM_ARCH_QWEN3VLMOE,
     LLM_ARCH_QWEN35,
     LLM_ARCH_QWEN35MOE,
+    LLM_ARCH_MAPLE,
     LLM_ARCH_PHI2,
     LLM_ARCH_PHI3,
     LLM_ARCH_PHIMOE,
diff --git a/src/llama-graph.cpp b/src/llama-graph.cpp
index 2be3b75fb..dedb537d3 100644
--- a/src/llama-graph.cpp
+++ b/src/llama-graph.cpp
@@ -2153,6 +2153,15 @@ ggml_tensor * llm_graph_context::build_moe_ffn(
                             cur = ggml_clamp(ctx0, cur, -INFINITY, limit);
                             cb(cur, "ffn_moe_gate_clamped", il);
                             cur = ggml_swiglu_split(ctx0, cur, up);
+                        } else if (arch == LLM_ARCH_MAPLE) {
+                            // Maple (MLX reference): silu(min(gate, +CLAMP)) * clip(up, -CLAMP, +CLAMP)
+                            cur = ggml_clamp(ctx0, cur, -INFINITY, limit);
+                            cb(cur, "ffn_moe_gate_clamped", il);
+                            ggml_tensor * gate_act = ggml_silu(ctx0, cur);
+                            cb(gate_act, "ffn_moe_silu", il);
+                            gate_act = ggml_clamp(ctx0, gate_act, -INFINITY, limit);
+                            cb(gate_act, "ffn_moe_silu_clamped", il);
+                            cur = ggml_mul(ctx0, gate_act, up);
                         } else {
                             ggml_tensor * gate_act = ggml_silu(ctx0, cur);
                             cb(gate_act, "ffn_moe_silu", il);
diff --git a/src/llama-model.cpp b/src/llama-model.cpp
index 4cc1c0a1c..1d4d31cf6 100644
--- a/src/llama-model.cpp
+++ b/src/llama-model.cpp
@@ -314,6 +314,8 @@ static llama_model * llama_model_mapping(llm_arch arch, const llama_model_params
             return new llama_model_kimi_linear(params);
         case LLM_ARCH_STEP35:
             return new llama_model_step35(params);
+        case LLM_ARCH_MAPLE:
+            return new llama_model_maple(params);
         default:
             throw std::runtime_error(std::string("unsupported model architecture: '") + llm_arch_name(arch) + "'");
     }
@@ -2613,6 +2615,7 @@ llama_rope_type llama_model_rope_type(const llama_model * model) {
             return LLAMA_ROPE_TYPE_NORM;

         // the pairs of head values are offset by n_rot/2
+        case LLM_ARCH_MAPLE:
         case LLM_ARCH_FALCON:
         case LLM_ARCH_FALCON_H1:
         case LLM_ARCH_GROK:
diff --git a/src/models/maple.cpp b/src/models/maple.cpp
new file mode 100644
index 000000000..f7dd44a5d
--- /dev/null
+++ b/src/models/maple.cpp
@@ -0,0 +1,203 @@
+#include "models.h"
+
+// Maple-Preview (DeepGrove, MIT): 20B-A1B ternary-weight MoE reasoning LLM.
+//
+// Block structure (matches the MLX reference implementation):
+//   - 3:1 hybrid attention: sliding-window-512 layers (partial RoPE, 64/128 dims)
+//     interleaved with full-attention layers that carry NO RoPE (n_rot_full == 0),
+//     full attention at il % 4 == 3.
+//   - Flash-head QK: per-head RMSNorm (attn_q_norm / attn_k_norm) applied on the
+//     reshaped Q/K tensors before RoPE, k_proj at n_embd_gqa (4 heads x 128).
+//   - MoE: 256 experts, 8 active, moe_intermediate 512, clamp-7 SwiGLU
+//     (silu(min(gate, +7)) * clip(up, -7, +7)) with fp32 softmax + renorm routing.
+
+void llama_model_maple::load_arch_hparams(llama_model_loader & ml) {
+    ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH,  hparams.n_ff_exp, false);
+    ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
+
+    // 3:1 SWA:GA hybrid; full-attention layers at il % 4 == 3 (config layer_types).
+    hparams.set_swa_pattern(4, false);
+    // Sliding window for the SWA layers (maple.attention.sliding_window in the GGUF).
+    ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa, false);
+    // Sliding-window attention cache (windowed KV, window = 512 from the GGUF).
+    hparams.swa_type = LLAMA_SWA_TYPE_STANDARD;
+
+    // GGUF has no rope.type key; Maple uses plain interleaved RoPE
+    // (MLX reference: initialize_rope(..., traditional=False) -> NORMAL).
+    hparams.rope_type = LLAMA_ROPE_TYPE_NEOX;
+
+    // Clamp-7 SwiGLU for the routed experts (consumed by build_moe_ffn).
+    for (uint32_t i = 0; i < hparams.n_layer(); ++i) {
+        hparams.swiglu_clamp_exp[i] = 7.0f;
+    }
+
+    type = LLM_TYPE_UNKNOWN;
+}
+
+void llama_model_maple::load_arch_tensors(llama_model_loader &) {
+    LLAMA_LOAD_LOCALS;
+
+    tok_embd = create_tensor(tn(LLM_TENSOR_TOKEN_EMBD, "weight"), {n_embd, n_vocab}, 0);
+
+    // output
+    output_norm = create_tensor(tn(LLM_TENSOR_OUTPUT_NORM, "weight"), {n_embd}, 0);
+    output      = create_tensor(tn(LLM_TENSOR_OUTPUT,      "weight"), {n_embd, n_vocab}, TENSOR_NOT_REQUIRED);
+    // if output is NULL, init from the input tok embed
+    if (output == NULL) {
+        output = create_tensor(tn(LLM_TENSOR_TOKEN_EMBD, "weight"), {n_embd, n_vocab}, TENSOR_DUPLICATED);
+    }
+
+    for (int i = 0; i < n_layer; ++i) {
+        auto & layer = layers[i];
+
+        layer.attn_norm = create_tensor(tn(LLM_TENSOR_ATTN_NORM, "weight", i), {n_embd}, 0);
+
+        create_tensor_qkv(layer, i, n_embd, n_embd_head_k * n_head, n_embd_gqa, n_embd_gqa, 0);
+        layer.wo = create_tensor(tn(LLM_TENSOR_ATTN_OUT, "weight", i), {n_embd_head_k * n_head, n_embd}, 0);
+
+        // Flash-head QK: per-head RMSNorm
+        layer.attn_k_norm = create_tensor(tn(LLM_TENSOR_ATTN_K_NORM, "weight", i), {n_embd_head_k}, 0);
+        layer.attn_q_norm = create_tensor(tn(LLM_TENSOR_ATTN_Q_NORM, "weight", i), {n_embd_head_k}, 0);
+
+        layer.ffn_norm = create_tensor(tn(LLM_TENSOR_FFN_NORM, "weight", i), {n_embd}, 0);
+
+        layer.ffn_gate_inp = create_tensor(tn(LLM_TENSOR_FFN_GATE_INP, "weight", i), {n_embd, n_expert}, 0);
+
+        if (n_expert == 0) {
+            throw std::runtime_error("n_expert must be > 0 for MAPLE");
+        }
+        if (n_expert_used == 0) {
+            throw std::runtime_error("n_expert_used must be > 0 for MAPLE");
+        }
+
+        // MoE branch
+        const int64_t n_ff_exp = hparams.n_ff_exp ? hparams.n_ff_exp : n_ff / n_expert_used;
+
+        layer.ffn_gate_exps = create_tensor(tn(LLM_TENSOR_FFN_GATE_EXPS, "weight", i), {  n_embd, n_ff_exp, n_expert}, 0);
+        layer.ffn_down_exps = create_tensor(tn(LLM_TENSOR_FFN_DOWN_EXPS, "weight", i), {n_ff_exp,   n_embd, n_expert}, 0);
+        layer.ffn_up_exps   = create_tensor(tn(LLM_TENSOR_FFN_UP_EXPS,   "weight", i), {  n_embd, n_ff_exp, n_expert}, 0);
+    }
+}
+
+std::unique_ptr<llm_graph_context> llama_model_maple::build_arch_graph(const llm_graph_params & params) const {
+    return std::make_unique<graph>(*this, params);
+}
+
+llama_model_maple::graph::graph(const llama_model & model, const llm_graph_params & params) : llm_graph_context(params) {
+    const int64_t n_embd_head = hparams.n_embd_head_v();
+
+    GGML_ASSERT(n_embd_head == hparams.n_embd_head_k());
+
+    ggml_tensor * cur;
+    ggml_tensor * inpL;
+
+    inpL = build_inp_embd(model.tok_embd);
+
+    // inp_pos - contains the positions
+    ggml_tensor * inp_pos = build_inp_pos();
+
+    auto * inp_attn = build_attn_inp_kv_iswa();
+
+    ggml_tensor * inp_out_ids = build_inp_out_ids();
+
+    for (int il = 0; il < n_layer; ++il) {
+        ggml_tensor * inpSA = inpL;
+
+        // norm
+        cur = build_norm(inpL,
+                model.layers[il].attn_norm, NULL,
+                LLM_NORM_RMS, il);
+        cb(cur, "attn_norm", il);
+
+        // self_attention
+        {
+            // compute Q and K and RoPE them
+            auto [Qcur, Kcur, Vcur] = build_qkv(model.layers[il], cur,
+                    n_embd_head, n_head, n_head_kv, il);
+
+            // Flash-head QK: per-head RMSNorm, then partial RoPE.
+            // n_rot(il) is 64 on SWA layers and 0 on full-attention layers,
+            // so the full-attention layers get a no-op RoPE (no position info).
+            Qcur = build_norm(Qcur, model.layers[il].attn_q_norm, NULL, LLM_NORM_RMS, il);
+            cb(Qcur, "Qcur_normed", il);
+
+            Qcur = ggml_rope_ext(
+                    ctx0, Qcur, inp_pos, nullptr,
+                    hparams.n_rot(il), rope_type, n_ctx_orig, freq_base, freq_scale,
+                    ext_factor, attn_factor, beta_fast, beta_slow
+                    );
+
+            Kcur = build_norm(Kcur, model.layers[il].attn_k_norm, NULL, LLM_NORM_RMS, il);
+            cb(Kcur, "Kcur_normed", il);
+
+            Kcur = ggml_rope_ext(
+                    ctx0, Kcur, inp_pos, nullptr,
+                    hparams.n_rot(il), rope_type, n_ctx_orig, freq_base, freq_scale,
+                    ext_factor, attn_factor, beta_fast, beta_slow
+                    );
+
+            cb(Qcur, "Qcur", il);
+            cb(Kcur, "Kcur", il);
+            cb(Vcur, "Vcur", il);
+
+            cur = build_attn(inp_attn,
+                    model.layers[il].wo, model.layers[il].wo_b, model.layers[il].wo_s,
+                    Qcur, Kcur, Vcur, nullptr, nullptr, nullptr, 1.0f/sqrtf(float(n_embd_head)), il);
+        }
+        if (il == n_layer - 1 && inp_out_ids) {
+            cur   = ggml_get_rows(ctx0,   cur, inp_out_ids);
+            inpSA = ggml_get_rows(ctx0, inpSA, inp_out_ids);
+        }
+        ggml_tensor * ffn_inp = ggml_add(ctx0, cur, inpSA);
+        cb(ffn_inp, "ffn_inp", il);
+
+        // MoE branch (clamp-7 SwiGLU handled inside build_moe_ffn for MAPLE)
+        cur = build_norm(ffn_inp,
+                model.layers[il].ffn_norm, NULL,
+                LLM_NORM_RMS, il);
+        cb(cur, "ffn_norm", il);
+
+        ggml_tensor * moe_out =
+            build_moe_ffn(cur,
+                    model.layers[il].ffn_gate_inp,
+                    model.layers[il].ffn_up_exps,
+                    model.layers[il].ffn_gate_exps,
+                    model.layers[il].ffn_down_exps,
+                    nullptr,
+                    n_expert, n_expert_used,
+                    LLM_FFN_SILU, true,
+                    hparams.expert_weights_scale,
+                    LLAMA_EXPERT_GATING_FUNC_TYPE_SOFTMAX,
+                    il,
+                    nullptr, nullptr,
+                    model.layers[il].ffn_up_exps_s,
+                    model.layers[il].ffn_gate_exps_s,
+                    model.layers[il].ffn_down_exps_s);
+        cb(moe_out, "ffn_moe_out", il);
+        cur = moe_out;
+
+        cur = ggml_add(ctx0, cur, ffn_inp);
+
+        cur = build_cvec(cur, il);
+        cb(cur, "l_out", il);
+
+        // input for next layer
+        inpL = cur;
+    }
+    cur = inpL;
+
+    cur = build_norm(cur,
+            model.output_norm, NULL,
+            LLM_NORM_RMS, -1);
+
+    cb(cur, "result_norm", -1);
+    res->t_embd = cur;
+
+    // lm_head
+    cur = build_lora_mm(model.output, cur, model.output_s);
+
+    cb(cur, "result_output", -1);
+    res->t_logits = cur;
+
+    ggml_build_forward_expand(gf, cur);
+}
diff --git a/src/models/models.h b/src/models/models.h
index ad3dadaf3..15e0fba12 100644
--- a/src/models/models.h
+++ b/src/models/models.h
@@ -569,6 +569,18 @@ struct llama_model_qwen3moe : public llama_model_base {
     std::unique_ptr<llm_graph_context> build_arch_graph(const llm_graph_params & params) const override;
 };

+struct llama_model_maple : public llama_model_base {
+    llama_model_maple(const struct llama_model_params & params) : llama_model_base(params) {}
+    void load_arch_hparams(llama_model_loader & ml) override;
+    void load_arch_tensors(llama_model_loader & ml) override;
+
+    struct graph : public llm_graph_context {
+        graph(const llama_model & model, const llm_graph_params & params);
+    };
+
+    std::unique_ptr<llm_graph_context> build_arch_graph(const llm_graph_params & params) const override;
+};
+

 struct llama_model_qwen3vl : public llama_model_base {
     llama_model_qwen3vl(const struct llama_model_params & params) : llama_model_base(params) {}

You should make a pr. I used this patch and some quantization surgery to convert https://huggingface.co/stamsam/maple-preview-gguf/blob/main/maple-q4_k_m.gguf to a q2_0.
Fits fully in vram (8gig), output is coherent, speed is insane (2300pp 167tg on short ctx, 2031pp 113tg at 40k).
The resulting q2_0 is 5.6GiB, q2_0 is not ideal because of the smaller group size (64), so there are some redundant group scales, but q2_0 has good cuda support.

edit: opend a hf pr, file here https://huggingface.co/stamsam/maple-preview-gguf/blob/88b54139c39cd8a4f260f3ac745d69efaf7096ca/maple-q2_0.gguf

I ran it partially offloaded on an old 2015 laptop with a gtx 960m (2gig) and still got 128pp 5tg on a short context.

Excellent stuff. The CPU path worked for me on M2 Max, and then had an agent bolt ASI metal for the gpu-s too - that 2x speed up it!
I put here the changes for the gpu paths on top of the original cpu paths changes, need to make the the cpu path work in llama.cpp in branch:

https://github.com/ljubomirj/llama.cpp/tree/LJ-maple-tq2_0-metal

Now I see my references are all over the place, I run the this script post-building it:

https://github.com/ljubomirj/llama.cpp/blob/LJ-maple-tq2_0-metal/llama_server_maple-preview_macbook2.sh

And this is where a got everything else from.

Code repo llama.cpp fork
https://github.com/deepgrove-ai/llama.cpp
Model files
https://huggingface.co/deepgrove/maple-preview-GGUF
My model file
https://huggingface.co/deepgrove/maple-preview-GGUF/resolve/main/maple-preview-TQ2_0-head-Q4_K.gguf
The llama.cpp fork and GGUF announcement
https://x.com/deepgrove_ai/status/2085190212427411715
Original non-GGUF HF
https://huggingface.co/deepgrove/maple-preview

My original intention was to run this not on the mac but on an old thinkpad that lacks gpu hence the interest in cpu only inference.
Coming to that now, that one was busy previously with trying the LFM models -
https://x.com/ljupc0/status/2085018529946902869

Thank you everyone for the amazing help here.

Sign up or log in to comment