Core AI model zoo

Kernel campaign — session handoff (2026-07-02)

One-page state + next-steps so the next session(s) pick up cleanly. Detail: dense-int4km-flagship- session-findings.md, spec-decode-design.md, flagship-full-tuning-stack.md, tensorops-quantized- kernels.md, memory project_accel_levers_campaign.

Net result of this session

Honest strategic picture

Quality-safe dramatic flagship speedup does NOT come from 4-bit (int4/fp4 both degrade, capped). It comes from:

  1. spec-decode (LOSSLESS, ~3×) — the real quality-safe lever. Verify K tokens/forward, distribution-exact.
  2. QAT-int4 — the only quality-safe route to the 4-bit bandwidth win (train to recover the 104→~32 flip gap).
  3. int8-dense bandwidth ≈ ~1.6× (but experts-int4 also degrades → partial).

NEXT per lever — START HERE

Coordination / rules (non-negotiable)

Artifacts (committed, coreai-models-community, committer john-rocky)

bab5fa7 LFM #2 export · 200cb2b flagship plan · 3eb4a5f spec-decode design · 5b790f5 Qwen3.6 export+findings · 69c65ea fp4-framing fix · 6436d8a prefill-Mac-open fix + m4 scaffold · e591499 tree-attn verify kernel. (§8/§8b fp4-disproven added by Stream D/user.) Uncommitted by design: coreai-models macos/ arsenal (Apple clone — incl. the MetalInt4KMLinear.weight fix + gemma4_metal_mlp_fp4.py), ondevice/ scripts (non-git: _qwen36_mac_bench.py, _dense_int4km_microbench.py, _flagship_dense_coverage_audit.py, _prefill_sdpa_baseline.py), coreml bundles in exports/.