Core AI model zoo

Zero ANE regions is a failure mode, not a rounding error

--preferred-compute neural-engine is a preference. When it cannot be honoured the compile still succeeds and the bundle still runs — on the GPU. The cheap signal is the region count, and it is easy to measure wrong in two different ways.

Measured on a base M4 (Mac mini, 16 GB), macOS 27.0 build 26A428, Xcode 27.0, coreai-build 3600.83.1, compile architecture h16g; contributed by @4rg0naut in PR #36.

Count it with the right glob

find <bundle>.aimodelc -name '*ANE_region_*.mlir.bc' | wc -l    # 0 = silent GPU fallback

Use that pattern, not *ANE_region*. The regions live inside the MPSGraph package (…/mpsExecutable.mpsgraphpackage/binary_0.llir.bundle/<fn>_ANE_region_0_0.bc/<arch>/…mlir.bc), so the loose pattern matches the region directory and the file inside it and double-counts every region. On the bundle measured here, *ANE_region* reported 14 / 2 / 2 / 2 where the correct counts are 13 / 1 / 1 / 1. The ordering survives the error, so it hides behind a correct story — which is exactly when a count gets reused in four places before anyone checks it.

Two failure modes, and they are not the same

The repo’s fp32 CV exports do reach the ANE — CLIP at fp32 lands in a GPU/ANE tie (apple-models-bench.md). A text encoder (Granite-Embedding-97M, ModernBERT, S=128) at fp32 instead forms zero regions: the compile emits an MPSGraph delegate and nothing else. A tie and a total fallback look the same in a latency table and are different problems.

build regions warm median load
fp32 0 4.31 ms (GPU) 797 ms
fp16 13 13.19 ms —
fp16 + two fp32-op removals 1 2.14 ms 257 ms

The two ops are the ones compute-units-and-authoring.md §29 warns about, named for this graph: F.softmax(scores, dim=-1, dtype=torch.float32) in the attention, and pooled = pooled.float() in the pooling head.

13 regions → 1 region is 13.19 ms → 4.60 ms for the same graph at the same precision, measured in the same pass: ~0.72 ms per boundary. That is the compile-side view of the delivery cost compute-units-and-authoring.md reports from Instruments (25 submissions, ~1.1 ms median inter-submission gap, 28.7% idle) — same phenomenon, different model, different device, and it supports the same conclusion: the fix is fewer, larger regions, and the region count is the metric that tells you whether you got them.

Reading the count honestly

It is a shape metric, not a quality one: it says how finely the graph is cut, not how much ran where. Two instruments that pair with it:

Two more traps in the same family

Compression, for this graph: four negatives

variant ANE regions gate min cosine clear pair flips
w8 + fp16 13 PASS 0.9993834 0
w6 + fp16 14 FAIL 0.9983865 1
w4 + fp16 0 FAIL 0.9621623 16

An observation left open

The ANE shows two throughput states on this graph: burst ~1.85–2.25 ms (~31 GB/s effective) and sustained ~4.75 ms (~14.5 GB/s). The transition is one-way within a run (samples 0–830 fast, 831–20,000 slow in a 20,000-inference run), short 120-inference bursts stay fast, and 4 minutes of idle does not restore it. The machine is far too efficient to throttle in 2 s, pmset -g therm raises nothing, the GPU path does not show the same magnitude, and the ANE power rail was not trustworthy enough to separate a clock drop from a bandwidth ceiling. Reported in case others have seen it.