Core AI model zoo

Qwen3.6-35B-A3B (text decoder) — Core AI

The first MoE model on the zoo, and the answer to Apple’s most-requested model (coreai-models issue #1). Source: Qwen/Qwen3.6-35B-A3B (an image-text-to-text checkpoint; this card is the text decoder).

Architecturally this is Qwen3.5’s hybrid decoder + a sparse Mixture-of-Experts FFN: the token mixers are exactly Qwen3.5 — 40 layers on a 3:1 interleave of GatedDeltaNet linear-attention mixers and gated full attention (head_dim 256, GQA 16/2, partial mRoPE θ=1e7) — but here the GatedDeltaNet runs 32 value heads over 16 key heads (GVA: each k/q head is shared across two value heads) and every FFN is a 256-expert top-8 MoE with a shared expert:

35B parameters, ~3B active per token → a frontier-quality model that decodes at Mac-class speed because only 8 of 256 experts fire per token.

⬇️ Converted .aimodel bundle: qwen3_6_35b_a3b_decode_sym8_gather/ (35 GB, the gather_qmm kernel build — 2.1× faster, same clean int8 quality; full LanguageBundle incl. tokenizer; decode-only loop-free for the pipelined engine). Convert with conversion/export_qwen3_6_moe_metal_decode_pipelined.py — recipe: recipe.toml.

Use it

One line — run the kit’s task op on this model (import CoreAIOps; no session, no model plumbing, downloads on first use):

let tldr = try await CoreAI.summarize(text, options: .model("qwen3.6-35b-a3b"))

Twenty ops, one shape — Cookbook.

▶️ Run it (source) — the ChatDemo runner (GUI + CLI, one app for every chat model in the catalog):

git clone https://github.com/john-rocky/coreai-kit
open coreai-kit/Examples/ChatDemo/ChatDemo.xcodeproj
# → Run, then pick "Qwen3.6-35B-A3B (MoE)" in the model picker

# agents / headless (macOS):
cd coreai-kit/Examples/ChatDemo
swift run chat-cli --model qwen3.6-35b-a3b --prompt "What can you do, offline?"

💻 Build with it — complete; the glue is kit API, copy-paste runs:

import CoreAIKit

let chat = try await ChatSession(catalog: "qwen3.6-35b-a3b")
let reply = try await chat.respond(to: prompt)
// reply: the answer, generated fully on-device

Also runs behind Apple’s FoundationModels API — CoreAIKit’s KitLanguageModel plugs this bundle into the system LanguageModelSession; capabilities (tool calling, guided generation) auto-detect per model.

The take-home is Examples/ChatDemo/Sources/QuickStart.swift — this exact code as one typed function, no UI; the CLI is an argument shell over it, and the GUI drives the same ChatSession across turns for its transcript. Multi-turn? Hold the ChatSession and call respond(to:) per turn — it keeps the conversation history; streamResponse(to:) yields tokens as they decode.

Integration checklist

Measured (macOS 27 beta, M4 Max 128 GB, release llm-benchmark, COREAI_CHUNK_THRESHOLD=1)

config bundle prefill tok/s decode tok/s numerics (teacher-forced vs bf16 HF oracle)
sym8 gather kernel (reads 8/256 experts) = SHIP 35 GB 65.3 64.9 CLEAN — 0 introduced flips/18 vs fp16 (sym8 = the int8lin recipe via a bit-exact gather)
int8 linear (GatherMM, reads all 256 experts) 35 GB 30.8 30.9 14/16 (1 flip, pos 6 margin 0.31) — same int8 quality, but the dense over-read
fp16 control (4 layers) 8 GB 167 175 — (truncated; engine-path smoke)

Numerics in full (the 35B is too large for an fp32 oracle on 128 GB, so the oracle is the checkpoint’s native bf16; gate = teacher-forced single-step argmax under the oracle-margin≥0.1 rule):

Speed: the expert-gather 2× is now CLOSED (gather_qmm kernel)

Update: the ~2× MoE-expert-gather penalty below is FIXED. A custom gather_qmm Metal kernel (models/macos/moe_metal.py) reads only the 8 routed experts (8/256) instead of GatherMM’s dense all-256 read — decode 30.9 → 64.9 tok/s (2.1×), and CLEAN (the sym8 scheme = the same symmetric-linear int8 recipe, read via a bit-exact gather: 0 introduced flips/18 vs fp16). That is exactly the “~2× (MoE expert-gather)” term predicted below, now realized. The remaining gap to MLX 4-bit is the int8-vs-int4 bytes term — which is a hard wall here (int4 fails this model’s numerics, below). The original accounting, for the record:

30.9 tok/s (the GatherMM path) is ~4× slower than MLX 4-bit on the same M4 Max (mlx-community/Qwen3.6-35B-A3B-4bit: 125 tok/s, measured), characterized as:

The real fix is a custom Metal gather-matmul kernel (Core AI exposes TorchMetalKernelcoreai.metal4_kernel), but that API is beta-experimental and the integration with the pipelined decode path is unverified — a high-variance multi-day spike we are deliberately deferring until the Core AI MoE path / kernel API matures. So this card ships at frontier quality with honest decode speed, and we expect the number to rise on OS/runtime updates without any change to the bundle. For a “fast on Mac today” model, the dense Gemma / Qwen ports are the better fit; this is the frontier-capability slot.

A Core AI GPU bug this port surfaced (worked around at the engine surface)

A MoE decode graph raw-loaded via AIModel.load(from_preferred_compute_unit_kind(.gpu)) makes MPSGraph try to lower the 256-expert GatherMM to the Neural Engine, which fails pathologically: a 2-layer slice aborts with MPSGraphANEUtils.mm:939: failed assertion 'ANE compilation writeToFile failed!', and a 4-layer slice explodes the MPSGraph temp dir past 100 GB (it crashed the host once). Bisection (_smoke/isolate_moe_ane_blowup.py, single-layer random-weight probes under a hard temp guard): a dense MLP layer loads in 0.2 s / 1 GB temp; swapping in the 256-expert MoE triggers the ANE assertion. So it’s the GatherMM→ANE path, not the GatedDeltaNet or the states.

The real pipelined engine never hits it. It loads dynamic models with SpecializationOptions(preferredComputeUnitKind: .gpu) plus expectFrequentReshapes = true (ModelStructure.swift), and that flag steers the graph off the ANE path. The Python runtime binding doesn’t expose expectFrequentReshapes, so raw Python load can’t replicate it — but the engine is the deployment path anyway, and Apple’s own gpt_oss-20B MoE runs the same way (78 tok/s). Proof: llm-benchmark on the real bundle is rc=0, peak temp 1 GB, 30.9 tok/s. Filed as Core AI feedback (ANE assertion / temp blowup on a raw-loaded GPU-preferred GatherMM graph).

How to reproduce

The published bundle is the sym8 gather-kernel build. Its recipe is recipe.toml (zoo_convert.py show qwen3.6-35b-a3b) — note the one open question recorded there: --head-sym changes the lm_head spec but does not appear in the bundle name, so which head configuration shipped is not recoverable from the artifact.

cd coreai-models   # with the qwen3_5 + qwen3_5_moe model overlay (see ../conversion)
# convert the SHIPPED bundle (gather kernel; ~70 GB fp16 load, mmap quantize; ~35 GB bundle)
.venv/bin/python ../coreai-models-community/conversion/export_qwen3_6_moe_metal_decode_pipelined.py \
    sym8
# the pre-gather-kernel build, kept for the comparison table above (NOT what is published):
.venv/bin/python ../coreai-models-community/conversion/export_qwen3_6_decode_pipelined.py \
    int8hu --head-sym
# bench (needs the coreai-pipelined-extra-states engine patch + COREAI_CHUNK_THRESHOLD=1)
COREAI_CHUNK_THRESHOLD=1 .build/release/llm-benchmark \
    --model exports/qwen3_6_35b_a3b_decode_int8hu_block32_sym -p 64 -g 128 -n 3

Model overlay: models/macos/qwen3_5_moe.py (MoE FFN on SwitchGLU + the packed-expert checkpoint loader) on top of the qwen3_5.py hybrid decoder (which gained the GVA query/key head-repeat for this model). Same engine patch stack as the other pipelined riders; decode-only loop-free because the GatedDeltaNet while_loop doesn’t lower on the GPU delegate.