Core AI model zoo

Kev: a LoRA + pointer-head decision model on Core AI, at 0.8B and 4B

2026-10-04. jaredpalmer/kev-0.8b and kev-4b (tag v1.0, Apache-2.0) answer typed questions about a state with a probability for every option: a rank-16 LoRA on Qwen3.5-0.8B-Base / 4B-Base and a pointer head that scores each option’s closing token against the question’s last token. The port is the author’s merge, the overlay’s text decoder returning hidden rows, and the head on the host. Both published graphs take 128 tokens a call and run each Gated DeltaNet recurrence in the overlay’s fp32 Metal kernel; the earlier graphs took 16 tokens a call, unrolled. Cards: models/kev-0.8b/README.md, models/kev-4b/README.md. Transcripts below are in models/kev-<size>/; records outside the repo are marked “lane” and live under $ZOO_WORK_ROOT/_kev/. Mac M4 Max, macOS 27.0 26A428, Xcode 27.0 RC; iPhone 18 Pro, iOS 27.0 24A437.

What was reused and what is new

part source change
merged weights the author’s scripts/merge_lora_checkpoint.py (tag kev-1.0) none: run as is, its output read through a local HF-cache id
decoder the overlay’s Qwen3_5StatefulForCausalLM (forward_stateful; the Gated DeltaNet recurrence unrolled, chunked in the graph, or in the overlay’s fp32 Metal kernel) subclass conversion/kev/qwen3_5_kev_decoder.py: lm_head = identity, the final-norm hidden state at every position out, one static-S function; no code change between 0.8B and 4B
oracle the author’s package (Checkpoint.load, kev.model.admit, row-form forward) oracle_kev.py runs it unchanged and asserts the readout on every record
head the author’s PointerHead (head.pt) export_head.py writes head.safetensors + kev_head.json; the hosts run it in float64
host clef-flash’s host.py / decide.py form new conversion/kev/host.py (the author’s to_record, render, user_tokens, encode, to_answers without the package) and decide.py
gates clef-flash’s readout gate, bisect and Swift scorer new readout_gate.py, parity_*.py, int8_bisect_torch.py, gate_swift.py
Swift host ClefFlash / DeciderVision’s low-level runtime pattern new apps/Kev (library + CLI), apps/KevGate (device gate)

The readout contract, and the token ids not to trust

One causal row per question: [<|fim_prefix|>] + state + [<|fim_middle|>] + instructions + for each option ([<|box_start|>] + option + [<|box_end|>]) + [<|fim_suffix|>], positions 0..L − 1, fresh states. The head reads the hidden state at <|fim_suffix|> and at every <|box_end|>: z = (k(h_opt) · q(h_decide)) / 16, p = softmax(z / T) within the question (T 2.3510958125672174 / 2.406050072164233).

The merge is the author’s script, and the proof is bit equality

scripts/merge_lora_checkpoint.py folds W + (B @ A) · α / r (α / r = 2) in fp32 and writes a full-weight checkpoint the author’s loader reads (head.pt weights “full”). Through that loader the merged checkpoint gives the adapter checkpoint’s logits bit for bit on 434/434 questions for both sizes, and three tensors per size equal the formula bit for bit (the adapter moves them by 2.8–3.6 % at 0.8B, 1.4–2.0 % at 4B). The red arms show the check can fail: the base alone moves p by up to 0.789 (25/53 argmax), the base’s Gated DeltaNet projections put back by up to 0.282 (gate-kev-0.8b-merge.json, gate-kev-4b-merge.json).

The merge is reproducible. Kev-0.8B’s merged checkpoint was deleted after round 15 and written again by the same script in round 16: model.safetensors (b5ebf92a…) and head.pt (c33e2288…) came out byte-identical to round 1’s. Only merge.json differs, by its phase times, so its sha256 is not the one gate-kev-0.8b-merge.json lists.

The oracle runs on one thread

transformers’ reference causal conv (no causal_conv1d package) is a depthwise F.conv1d with 6,144 groups (8,192 at 4B), which the CPU runs as many tiny calls; their thread start-up dominates. The decoder module’s probe on the same conv: F.conv1d at 1 / 4 / 12 threads = 14.8 / 86.9 / 262.7 ms per token at 0.8B with bit-identical hidden rows; the same conv as one bmm over the windows = 7.5 ms at 1 thread (gate-kev-0.8b-torch-parity.json probe). At 4B, 12 threads also change the bits (gate-kev-4b-torch-parity.json probe). The oracle runs with one thread (oracle_summary.json threads: 1); the parity harness uses the bmm form and checks it against F.conv1d at the last chunk of every row (7,812 checks at 0.8B, 10,416 at 4B, all bit-identical).

The decoder is the overlay’s text graph without a head

Qwen3_5KevDecoder is the overlay’s stateful text decoder with lm_head replaced by the identity: load_report reads 320 / 426 keys and leaves none, and the checkpoint has no lm_head key (tied). Making the head an identity also keeps the tied embedding out of any int8 pass that targets linears. In fp32 torch, driven like the graph, it gives the oracle’s p within 3.5e-6 (0.8B) / 6.6e-6 (4B) on all 434 questions, and the overlay’s plain forward_stateful bit for bit (3 rows, every chunk and layer).

The scan form sets the cost of a call

Each Gated DeltaNet layer runs its recurrence inside the call. The overlay has three forms of it, chosen with --gdn-scan in conversion/kev/export_decoder.py. The graph’s inputs, outputs and states are the same in all three.

form --gdn-scan what runs inside a call
unrolled, US unroll the S steps, one after another, fp32
in-graph chunk, CS chunk the overlay’s _gated_delta_chunk: the S tokens at once, the triangular inverse as ⌈log2 S⌉ doublings, fp32
Metal kernel, KS metal the overlay’s fp32 Metal chunk kernel (qwen3_5_gdn_metal), one dispatch per layer for the whole call

A call’s time fits a + b·S, with S the tokens in the call. The table gives a and b on the M4 Max (macOS 27.0, 2026-10-04; gate-kev-<size>-forms.json call_cost), for Kev-0.8B unless a row names Kev-4B. A window marked shared ran beside other GPU work.

form window a, ms b, ms per token
unrolled round 11 screen, shared 1.7 1.04
in-graph chunk round 11 screen, shared, two points 9.2 0.10
Metal kernel round 11 screen, shared 7.9 0.17
unrolled round 11 lock window, two points 0.34 1.387
Metal kernel round 11 lock window 8.84 0.166
Metal kernel, static S round 14 8.52 0.169
Metal kernel, dynamic length, cap 128 / 256 / 512 round 14 7.99 / 7.34 / 7.44 0.184 / 0.181 / 0.182
Metal kernel, iPhone 18 Pro round 13 bench, S = 16 to 128 12.81 0.181
unrolled, Kev-4B round 12 screen, shared 16.24 2.509
Metal kernel, Kev-4B round 12 screen, shared 22.84 0.700

A dynamic query length: measured, not shipped

--dynamic-query exports the Metal-kernel form with a query length of 2..cap (input_ids [1, -1]); the kernel is built for the cap. Round 14 measured caps 128, 256 and 512 on the M4 Max; round 15 gated and timed the cap-512 graph.

int8: measured, not shipped

Kev-0.8B (gate-kev-0.8b-int8.json): int8 per block of 32 over all 186 linears (symmetric_with_clipping) fails (0.0431 / mean 0.00326; held out 0.0273). The fp32 bisect with the exporter’s own int8 weights reproduces it (0.0420, per-row correlation 0.972). Kept fp16 alone: one layer at best 0.0360 (layer 3); all MLP linears (53 % of the weights) 0.0308; all Gated DeltaNet projections (38 %) 0.0359; all attention (9 %) 0.0390. The rule (worst < 0.010 with at most 50 % fp16, by single-layer then single-kind rank, then unions) found no set (best 0.0141 at 45.5 %). Layers 0–11 fp16 (50 %, outside the rule) give 0.0046 on the bisect rows, 0.0089 on the fixture and 0.0043 on the held-out set, at 85 % of fp16’s bytes.

Kev-4B (gate-kev-4b-int8.json): int8lin fails (0.0601; held out 0.0285). The instrument (40 rows, a rule written first, at most 25 % fp16, target worst ≤ 0.006 and mean ≤ 0.0012):

int8 variant, fp16 kept share worst mean
all, block 32, symmetric_with_clipping 0 % 0.0710 0.00689
all, block 16 0 % 0.0470 0.00453
the embedding table only (body fp32) 100 % 0.0052 0.00056
asymmetric block 32, in_proj_qkv + out_proj 21.2 % 0.0163 0.00208
asymmetric block 32, in_proj_qkv + out_proj + v_proj 21.7 % 0.0164 0.00191

No set met the target; the rule took the two best by mean. On the Mac GPU both pass the 434 rows (0.0174 / 0.0175); the one run on the held-out set fails on one score row (0.0277, tv4sh_11, fp16’s held-out worst too).

What the compiled asset holds

First calls and caches

An AOT cache entry is named by the function’s type

A dynamic-length graph’s footprint grows while the call length keeps changing

Round 15 probed the cap-512 dynamic graph through the .aimodel specialized by the Swift runtime, on the M4 Max (lane results/r15_leak_0_8b.json, phys_footprint of one process):

Host: float64 head, Python’s sum, Python’s text

The shared prefix is exact on a static graph

Every row of a request starts with the state’s Ls tokens, so with a static call of S tokens and k = ⌊Ls / S⌋ the first k calls are identical in every row. Run them once, copy the four states after them (the KV axis up to S·k, the conv and recurrent states as they are), and run each row’s remaining tokens from position S·k in the same call grid. At S = 16 the hidden rows equal the direct run’s bit for bit on 564/564 rows of each size, and the fixture’s 20 multi-question records take 662 calls instead of 1,830 (gate-kev-<size>-host.json shared_prefix). With the published K128 graph the Swift host’s shared rows equal its direct rows on 434/434 and 130/130 rows (gate-kev-0.8b-swift.json). A graph with a dynamic query length cuts the calls elsewhere (above). coreai.runtime.NDArray(array) copies its array, and state[name].numpy() reads what the calls wrote: read the states back once, then build fresh arrays per question.

Swift host notes

Timing on a shared machine

Other lanes shared this Mac’s GPU while the forms were timed. A time in the cards counts only if it was measured under this rule, fixed before the timing runs:

iPhone 18 Pro

Kev-0.8B through apps/KevGate (Release, no increased-memory-limit entitlement, 3,529 MB available at launch) (gate-kev-0.8b-iphone.json). Round 8 gated the earlier upload’s graph (U16), round 13 the Metal-kernel forms and the dynamic-length graph, round 15 the published K128 graph:

The rest of this section is round 8’s, on the U16 graph:

The fixture’s invented names

The 20 records written for the port use ten invented stems that had no DNS A record (.com/.net/.io/.co.uk/.app), no App Store result and no Companies House company. An exact-word web search on 2026-10-04 found four of them in use: a Kindle book series, a retail product, places in two games, a fantasy wiki and a fan-wiki character. The nine records using those four stems are withheld from the published fixture (their ids, request_sha256 and numbers stay); the oracle was not re-run with new names. A screen of invented names needs the exact-word web search as well as the registries.

Agreement with the fixture’s gold labels (our subset)

The author’s fp32 oracle against the fixture’s labels: our subset of 434 questions, not the author’s evaluation suites or tables (oracle_summary.json, oracle_summary_4b.json, lane).

group questions Kev-0.8B Kev-4B
transfer-v4 development, first 60 (MMLU) 60 25 45
transfer-v4, 20 of each other source 140 101 119
transfer-v4 score, first 20 20 8 14
SemIf authored144 144 104 129
written for the port 70 52 66

Numbers of record