AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP

Checkpoint Tier 1 certified on df-macbookpro-m5 (2026-08-09) for Hub commit 85b0a78a1484. Measured size vs matched uniform baseline, quality retention ≥0.98, and conversion integrity. Tier 1 is not a speed claim: MTP acceleration is not certified. See the certificate.

AXQuant mixed-precision (AXQ) Gemma 4 target with multi-token prediction (MTP) — named like the Qwen AXQ-MTP line (…-MLX-AXQ-*-MTP).

Property Value
Product class AXQ 4bit
Target size 26B-A4B
Measured BPW ~4.90
Hub id AutomatosX/AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP
Hub commit 85b0a78a14843a818d403f9a2525efa2f081c2a4
Engine pair gemma-4-26b-a4b-itgemma-4-26b-a4b-it-assistant
MTP layout assistant/ + ax_gemma4_assistant_mtp.json (Gemma assistant-MTP)
Certification host df-macbookpro-m5

Layout

./                          # AXQ target weights (+ vision sidecar when present)
assistant/                  # gemma4_assistant drafter
ax_gemma4_assistant_mtp.json
ax_composite_pack_manifest.json
axquant_*.json

Naming follows Qwen AXQ MTP packs: the -MTP suffix means MTP assets are in-repo.

Certification

Claim Status
Checkpoint Tier 1 (size + quality + load) Certified on df-macbookpro-m5
MTP acceleration Tier 2 Not certified (assistant-MTP is present for product completeness only)
Vision / multimodal quality Not claimed

Certificate: gemma4-26b-a4b-axq4-tier1.md

Runtime (AX Engine 7.1.5)

Download the complete snapshot and serve its local root:

hf download AutomatosX/AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP --local-dir ./AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP
ax-engine serve ./AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP --port 31418

AX Engine validates ax_gemma4_assistant_mtp.json and the exact-paired assistant/ bundle before attaching the drafter. Assistant-MTP is enabled by default in AX Engine 7.1.5 with a maximum draft depth of two. Set AX_MLX_GEMMA4_ASSISTANT_MTP=0 to force direct decode, or set AX_MLX_GEMMA4_ASSISTANT_MTP_MAX_DEPTH=1 to cap drafting at one token.

Stock MLX-LM does not consume the assistant/ bundle. This is not a Qwen sidecar, so the oMLX/MTPLX Qwen import workflow does not apply. The checkpoint card still makes no Tier 2 MTP acceleration claim.

Published / card updated: 2026-08-21.

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

Modality Claim Supported Reason
Vision present-not-certified true vision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP.json
Audio not-applicable false audio not supported (no tower config and no sidecar weights)
Downloads last month
73
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collections including AutomatosX/AX-gemma-4-26b-a4b-MLX-AXQ-4bit-MTP