Ling 3.0 Flash — CIRU IU4
Ciru builds for Kairic.ai to bring the best inference optimizations to production.
Ling-3.0-Flash-CIRU-IU4 is CIRU's Ling-specific INT4 checkpoint and AMD Strix Halo serving build for InclusionAI's 124B-parameter, 5.1B-active Ling 3.0 Flash model.
Custom runtime required for native IU4: use the
ciru/ling3-gfx1151-iu4branch ofciru-ai/vllm. Stock vLLM can load these weights through W4A16 compatibility, but it does not provide the IU4 path or reproduce the performance figures below.
This release combines three complementary changes:
- Better S4/G32 weights. CIRU recalibrated all 61,440 routed-expert projections in layers 2–41 using a Ling activation-importance objective. The checkpoint remains symmetric signed INT4, group size 32, in the same compressed-tensors format as the official model.
- Native S4 × S4 prefill on AMD IU4. Prefill uses a two-plane residual activation representation and the native
gfx1151integer WMMA lane. The packed checkpoint bytes remain canonical and safe for the normal W4 path. - Production W4 decode with native MTP K1. Sparse decode stays on the optimized W4 path, where Ling's 512-expert/top-8 routing is better matched to small token batches.
The result is one production profile: native IU4 where it earns throughput, optimized W4 where sparse decode wins, and no user-facing mode selection.
Weight fidelity
Against the unchanged official INT4 checkpoint on CIRU's locked 92 × 4096-token BF16-reference corpus:
| Checkpoint | Mean KL | Top-token agreement |
|---|---|---|
| Official Ling INT4 | 0.040795 | 95.7063% |
| CIRU IU4 weights | 0.035464 | 95.8784% |
| Improvement | −13.07% | +0.172 percentage points |
The recalibration changed weights, not the model architecture or tensor schema. The final next-token-prediction layer remains byte-identical to the official checkpoint because it was outside the calibration coverage.
Measured serving performance
Validated on one Ryzen AI MAX+ 395 / Radeon 8060S Strix Halo APU with 128 GB unified memory, native Linux, ROCm/HIP 7.15, Python 3.12, Torch 2.13 ROCm, Triton 3.8, and the CIRU gfx1151 vLLM runtime.
Six simultaneous 4K requests
Matched warmed run: six unique uncached 4096-token prompts, 256 forced output tokens each, MTP K1, 16K scheduler budget, prefix reuse disabled.
| Metric | Official W4 control | CIRU IU4 | Change |
|---|---|---|---|
| Aggregate prefill | 504.89 tok/s | 644.94 tok/s | +27.74% |
| Aggregate decode | 70.92 tok/s | 81.90 tok/s | +15.48% |
| Median per-request decode | 11.92 tok/s | 14.51 tok/s | +21.73% |
| Median TTFT | 48.67 s | 38.10 s | −21.72% |
| Batch wall time | 70.25 s | 55.71 s | −20.69% |
| MTP acceptance | 98.19% | 98.32% | +0.13 pp |
Both arms completed all requests with zero cached prompt tokens, zero prefix hits, zero preemptions, and no faults.
Full HumanEval and HumanEval+
The production profile completed a six-worker full-suite run with workers replenished continuously:
| Evaluation | Result |
|---|---|
| HumanEval pass@1 | 159 / 164 — 96.95% |
| HumanEval+ pass@1 | 156 / 164 — 95.12% |
| Aggregate generated throughput | 64.01 tok/s |
| Completion tokens | 200,550 |
| Generation wall time | 3,133.06 s |
| MTP acceptance | 91.88% |
| Runtime faults / preemptions | 0 / 0 |
Sampling used InclusionAI's recommended live settings (temperature=0.6, top_p=0.95, top_k=20). The harness imposed a 4096-token ceiling per task; seven tasks reached that ceiling, so these are the measured capped scores rather than an extrapolated score.
Model details
| Property | Value |
|---|---|
| Architecture | BailingMoeV3ForCausalLM / hybrid KDA–MLA MoE |
| Parameters | 124B total, 5.1B active |
| Experts | 512 routed experts, top-8 |
| Native context | 262,144 tokens |
| Weight format | compressed-tensors pack-quantized |
| Quantization | symmetric signed INT4, group size 32, BF16 scales |
| Source checkpoint | inclusionAI/Ling-3.0-flash-int4 |
Usage
The custom runner, gfx1151 kernel sources, extension builder, and production launcher are published in tools/ling_iu4. The branch is pinned to public release commit 40027ba1d.
1. Build the runner
Ubuntu 24.04 / Debian-family setup:
sudo apt update
sudo apt install -y build-essential git libnuma-dev libdrm-dev libelf-dev \
ninja-build pkg-config python3.12 python3.12-dev python3.12-venv xxd
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
git clone --branch ciru/ling3-gfx1151-iu4 \
https://github.com/ciru-ai/vllm.git "$HOME/vllm-ling-ciru-iu4"
cd "$HOME/vllm-ling-ciru-iu4"
git checkout 40027ba1d
bash tools/ling_iu4/install_runtime.sh
source .venv/bin/activate
The installer rebuilds the wheel-pinned ROCr userspace library with the host
compiler. This prevents the distributed ROCm 7.15 runtime from leaving a CPU
core continuously busy while the gfx1151 server is idle. The launcher
checksum-verifies and preloads the corrected library automatically.
Existing installs made from commit 92ccee356 can apply only the idle-runtime
hotfix without rebuilding vLLM, the IU4 extension, or the model:
git fetch origin ciru/ling3-gfx1151-iu4
git checkout 40027ba1d
bash tools/ling_iu4/install_rocr_idle_fix.sh "$PWD"
Restart the server afterward. CIRU_DISABLE_ROCR_IDLE_FIX=1 is an emergency
diagnostic rollback to the wheel runtime and may restore the idle CPU spin.
2. Download the model
hf download jcbtc/Ling-3.0-Flash-CIRU-IU4 \
--local-dir "$HOME/models/Ling-3.0-Flash-CIRU-IU4"
3. Launch the production profile
From the CIRU vLLM checkout:
source .venv/bin/activate
MODEL_PATH="$HOME/models/Ling-3.0-Flash-CIRU-IU4" \
bash tools/ling_iu4/run_ling_ciru_iu4.sh
The launcher provides the measured configuration: native 262,144-token context, six scheduler slots, a 16K batched-token budget, chunked prefill, prefix caching, Ling tool/reasoning parsers, optimized W4 sparse execution, native IU4 execution for occupied expert work, and MTP K1. Override HOST, PORT, MAX_MODEL_LEN, MAX_NUM_SEQS, MAX_NUM_BATCHED_TOKENS, or GPU_MEMORY_UTILIZATION only when needed.
The kernel build and arithmetic smoke were reproduced from the clean public branch on CIRU's Radeon 8060S (gfx1151). The performance measurements were validated on CIRU's native-Linux NixOS host. The Ubuntu commands above are the supported mainstream build path, but are not presented as a separate clean-host benchmark validation. Native Windows vLLM is unsupported; Windows users require WSL2 with AMD ROCDXG.
Request example
curl -fsS http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Ling-3.0-Flash-CIRU-IU4",
"messages": [{"role": "user", "content": "Write a Python function that merges overlapping intervals."}],
"chat_template_kwargs": {"enable_thinking": true},
"temperature": 0.6,
"top_p": 0.95,
"top_k": 20,
"max_tokens": 1024
}'
Intended use and limitations
This release targets local, OpenAI-compatible text generation and coding-agent workloads on high-memory AMD Strix Halo systems. It preserves Ling's native reasoning and tool-call behavior.
- The native IU4 acceleration is hardware-specific to
gfx1151; other devices use their supported W4A16 kernels. - The checkpoint is approximately 77 GB before runtime caches and KV cache. A 128 GB unified-memory system is the validated target.
- Benchmark results are single-host measurements and should not be generalized to different ROCm, vLLM, firmware, power, or memory configurations.
- Generated code must be reviewed and sandboxed before production execution.
Credits and license
- Base model, tokenizer, architecture, and official INT4 checkpoint: InclusionAI, MIT license.
- Inference engine: vLLM, Apache-2.0.
- AMD ROCm, HIP, and
gfx1151integer WMMA support: AMD and their respective upstream projects. - Weight calibration, native IU4 kernels, runtime integration, profiling, and validation: CIRU.
The original MIT license governs the model weights and accompanying model code. Review the upstream model card before deployment.
- Downloads last month
- 273
Model tree for jcbtc/Ling-3.0-Flash-CIRU-IU4
Base model
inclusionAI/Ling-3.0-flash-int4