Core AI model zoo

Kev-4B — Core AI

🤗 mlboydaisuke/Kev-4B-CoreAI · Apache-2.0 · source jaredpalmer/kev-4b (tag v1.0, commit 6cfce5c) · base Qwen/Qwen3.5-4B-Base (revision 1001bb4)

A decision model. Give it a state (text, or a JSON object or array) and typed questions: noul (yes / no), choice (named options) or score (ordered levels). It returns a probability for every option of every question. It never generates text. Requests and responses use the SystemOne-compatible request shape: {model, state, questions} in, {model, answers, usage} out, one typed decision per question.

Jared Palmer trained it as a rank-16 LoRA adapter and a pointer head on Qwen3.5-4B-Base (32 layers: 24 Gated DeltaNet and 8 full attention, hidden size 2,560). The author’s card: “It is intended for developers who classify, route, triage or check documents and who need probabilities that can be thresholded, for example to send uncertain cases to human review.” Each question is its own row: the state, then the question and its options. The head scores each option’s closing token against the question’s last token. The author’s card reports benchmark results; none of them are re-measured here.

This port merges the adapter into the base with the author’s own merge script and exports the text backbone as one Core AI graph. The graph takes 128 token ids per call and returns the final-norm hidden state at every position; it has no vocabulary head. Each Gated DeltaNet layer runs its recurrence in an fp32 Metal kernel, one GPU dispatch per layer per call. Weights are fp16 (8.41 GB). The host runs the pointer head from the author’s head weights. The gate is probability parity with the author’s own fp32 code, on every option of every question, on a fixture of 434 questions and on a held-out set of 130.

It runs on the Mac GPU. This release is for the Mac: the one load of an iPhone AOT asset of an earlier graph on an iPhone 18 Pro crashed (below). Kev-0.8B, the same port at 0.8B, is measured on the iPhone 18 Pro.

Readout contract

The contract is Kev-0.8B’s, with this checkpoint’s head and temperature; the full description is in the Kev-0.8B card. In short:

[<|fim_prefix|>] + render(state) + [<|fim_middle|>] + render(instructions)
  + for each option: [<|box_start|>] + option_text + [<|box_end|>]
  + [<|fim_suffix|>]

The fixture and the held-out set are Kev-0.8B’s (434 and 130 questions), with this checkpoint’s own fp32 oracle run (the author’s package, one CPU thread): fixtures-kev-4b.json. It has twelve near-ties on the fixture and four on the held-out set.

Core AI shape

The same module as Kev-0.8B, conversion/kev/qwen3_5_kev_decoder.py, with no code change: the overlay’s stateful Qwen3.5 text decoder on the merged weights, no vocabulary head. One function, main, at a static 128 tokens. Each Gated DeltaNet layer runs its recurrence in the overlay’s fp32 Metal chunk kernel (qwen3_5_gdn_metal): one GPU dispatch per layer per call, the recurrent state kept in fp32 through the call.

  name shape, type
inputs input_ids [1, 128] int32
  position_ids [1, seq] int32, the ramp 0 .. seq − 1
states keyCache, valueCache [8, 1, 4, ctx, 256] fp16, ctx up to 4,096
  convState, recState [24, 1, 8192, 3], [24, 1, 32, 128, 128] fp16
output hidden [1, 128, 2560] fp16, every position

Per row of T ids, from zeroed states: ⌈T / 128⌉ calls, the last padded with <|endoftext|> (248044) and its padded rows dropped. metadata.json says kind: decision-backbone and language.prefill_chunk: 128, and carries the readout contract. The host (conversion/kev/decide.py, apps/Kev) is Kev-0.8B’s: the head in float64 from head/head.safetensors, p rounded to fp32 once, the SystemOne-compatible response, the shared prefix and the prepared state. No engine, no runtime patch.

Measured (Apple M4 Max, macOS 27.0 26A428, 2026-10-03 – 10-04)

The bar is Kev-0.8B’s, fixed before any graph ran: the argmax equal to the oracle’s on every question whose oracle top-2 margin is above 0.02 (near-ties listed apart), max |Δp| ≤ 0.02 over every option of every question, the mean over rows of each row’s mean |Δp| ≤ 0.002, every process’s re-run bit-equal. The Python gates load the AOT .aimodelc (h16c, --expect-frequent-reshapes) with SpecializationOptions.default(); their times are on a GPU shared with other work.

Before any graph: the merge and fp32 torch

The graph alone on the Mac GPU

set questions argmax (margin > 0.02) near-ties agreeing max |Δp| mean of row means bar
fixture 434 422/422 10/12 0.0153 0.00077 PASS
held out 130 126/126 3/4 0.0096 0.00069 PASS

The near-ties that flip have oracle margins of 0.0005–0.0024 and move by at most 0.0029. The worst fixture row is a SemIf item (0.0153), the worst held-out row a score item (0.0096). The lowest per-position hidden cosine on the 21 recorded rows is 0.9921. Every value is finite and every process’s re-run is bit-equal. The red arms move p: a swapped state moves 5 of 5 argmaxes (max |Δp| 0.988), a grammatical “not” 3 of 5 (0.957). Transcript: gate-kev-4b-readout.json.

From the request: Python and Swift

Time per decision on the Mac

Swift, Release CLI, the AOT asset, in one machine-wide GPU lock window (2026-10-04 14:05–14:30 JST). Two processes per graph counted only when no other GPU job ran and the one-minute load average was at most 12; each made 6 decisions per item after one warm-up, and the table gives the median (p10–p90) over the 12. The earlier unrolled graph (16 tokens per call) ran in the same window:

request tokens calls ms p10–p90 unrolled S = 16, ms
one question 94 1 100.2 99.3–101.2 269.5
one question 380 3 298.0 296.1–304.5 1,075.4
one question 1,518 12 1,189.0 1,182.9–1,194.9 4,341.9
one question 1,802 15 1,488.1 1,484.3–1,498.5 5,200.2
five questions on one 137-token state, each row from zero 240 10 986.7 983.0–990.2 2,367.9
the same five questions, the state run once (shared) 240 6 615.6 613.3–624.5 894.7
eight questions on that state, each row from zero 320 16 1,578.7 1,574.2–1,591.6 3,776.0
the same eight questions, shared 320 9 920.0 916.9–930.6 1,262.4
four questions on one 1,477-token state, each row from zero 1,622 49 4,851.7 4,839.2–4,870.6 17,087.7
the same four questions, shared 1,622 16 1,614.5 1,607.5–1,647.8 4,713.7

Loading the AOT asset took 9.90 / 9.91 s at the start of each process and 1.30 / 1.22 s when loaded again. The two processes ended at 0.89 and 0.95 GB of footprint. Transcript: gate-kev-4b-timing-mac.json.

iPhone 18 Pro: not in this release

An earlier graph of this port (the recurrence unrolled, 16 tokens per call) compiled for the iPhone 18 Pro’s GPU (h19p, --expect-frequent-reshapes) is 15,557,759,365 bytes. Loaded once in the headless gate app (iOS 27.0 24A437, .default options, no increased-memory-limit entitlement), the app died inside AIModel(contentsOf:) with EXC_BAD_ACCESS (SIGSEGV) in the on-device compile for delegates (-[MPSGraphAICodeCompilerDelegate getInitializedAICodeBytecodeWithPayloadPrefix:delegateId:]). Its last memory sample, 0.32 s in, read a footprint of 175.9 MB with 3,364.1 MB available. The cause is not isolated and the load was not retried; the graph of this release was not loaded on the phone (../kev-0.8b/gate-kev-0.8b-iphone.json round8.device.summary.load_4b, round8.load_4b_memory_samples).

Forms measured

Every form below was exported and compiled the same way and read against the same oracle: the full gate, or the 130-row subset where the table says so. The times come from different windows, so compare forms only within one window. Full records: gate-kev-4b-forms.json.

Mac (Swift Release CLI, the 94-token question, median of the counted processes):

form 94-token decision, ms window fixture max |Δp| note
Gated DeltaNet unrolled, S = 16 269.5 round 12 0.0154  
unrolled, S = 32 / 64 — — 0.0078 / 0.0070 (130 rows) not timed in a lock window (104.8 / 174.1 ms per call on a shared GPU)
Metal kernel, S = 32 — — 0.0074 (130 rows) not timed in a lock window (48.4 ms per call on a shared GPU)
Metal kernel, S = 64 118.0 round 12 0.0170 five questions shared: 435.9 ms (S = 128: 615.6)
Metal kernel, S = 128 (this release) 100.2 round 12 0.0153  
Metal kernel, query length 2..512, host calls of at most 512 ids 83.3 round 12 0.0132 not shipped: memory grows (below)
in-graph chunk scan — — — not run at 4B; on Kev-0.8B it fails from S = 32
int8, every linear (unrolled S = 16) — — 0.0601 fails the bar (Precision)
any form on the Neural Engine — — — not run: the recurrence’s fp32 state is not an ANE element type (qwen3.5-static-ane.md)

S = 64 for several short questions on one state. In the same window, five questions on one 137-token state, shared, took 435.9 ms with S = 64 and 615.6 ms with S = 128; one 94-token question took 118.0 and 100.2 ms. This release ships S = 128.

Why the dynamic-length graph does not ship. It passed the gate and decided the 94-token question in 83.3 ms in the same window. Its two timing processes ended at 6.31 and 6.32 GB of footprint, against 0.89 and 0.95 GB for S = 128. On Kev-0.8B the same form keeps growing while the call length keeps changing, until the AIModel is created again (Kev-0.8B card).

Precision

int8 was measured and is not shipped. It was measured on the unrolled S = 16 graph; the Metal-kernel graphs were not measured in int8. int8 per block of 32 over every decoder linear fails the bar:

decoder (unrolled S = 16) fp16 kept main.mlirb bytes fixture max |Δp| / mean held out max |Δp| / mean bar
fp16 100 % 8,414,007,636 0.0154 / 0.00078 0.0109 / 0.00072 PASS
int8lin: every linear int8 (symmetric_with_clipping) 0 % 5,068,103,201 0.0601 / 0.00192 0.0285 / 0.00169 FAIL
asymmetric int8, in_proj_qkv + out_proj + v_proj fp16 21.7 % 5,882,832,315 0.0175 / 0.00135 0.0277 / 0.00123 FAIL
asymmetric int8, in_proj_qkv + out_proj fp16 21.2 % 5,863,831,482 0.0174 / 0.00138 not run —

The two asymmetric sets are the ones an fp32 torch bisect chose under a rule written before any result (at most 25 % of the linear weights fp16, the target worst row ≤ 0.006 met by none; the two best by mean taken). The one with v_proj fp16 passes the fixture and fails the held-out set on one score row, the row that also breaks int8lin and is fp16’s held-out worst. A set built before the rule’s last stage (layers 0–3 and out_proj fp16) fails both (0.0272, 0.0293). The quantizer’s qscheme accepts symmetric, asymmetric and symmetric_with_clipping; block 16 and an int8 embedding table also compile (gate-kev-4b-int8.json).

What the compiled asset holds. --expect-frequent-reshapes adds an fp16 copy of every linear weight: for the unrolled graph the h19p asset of the fp16 bundle is 15,557,759,365 bytes with it and 8,413,701,860 without, and the int8 sets above compile to 13,026,670,589 and 13,007,670,689 bytes with it. This release’s Mac h16c asset is 15,552,786,962 bytes against an 8,412,500,461-byte main.mlirb. Without the flag the runtime specializes the unrolled graph again for every new position length: 9.6–14.6 s per new length on the Mac GPU (calls at a length already seen: 46.4 ms median, 33 rows); an int8 asset without the flag had not finished specializing its opening length after 12 minutes. A Mac gate worker on the efr fp16 asset of the unrolled graph showed up to 16.6 GB rss, up to 16.4 GB of it clean mapped file, and a phys_footprint of at most 1.15 GB.

⬇️ Bundle

mlboydaisuke/Kev-4B-CoreAI, one folder under gpu-pipelined/:

file what bytes sha256
kev_4b_decode_fp16_metal_pf128.aimodel/main.mlirb the decoder, fp16 8,412,500,461 da529dd4…413eb17f
metadata.json kind: decision-backbone, the readout contract, the gate 12,643 e641e44f…c9039b08
tokenizer/tokenizer.json the base model’s, verbatim 12,807,196 fe000e3e…d50d2927
tokenizer/tokenizer_config.json the base model’s, verbatim 16,713 3891e840…b6b9f89c
head/head.safetensors the pointer head: q / k weight [256, 2560] and bias, fp32 5,245,304 36c392f7…54157106
head/kev_head.json head size, scale, temperature, delimiter ids, provenance 2,038 5ff071ff…cfba7416

The repository root carries the base model’s config.json, the adapter’s adapter_config.json and the Apache-2.0 LICENSE, verbatim. SHA256SUMS lists every file. Its Hub revision is ee978c7 (2026-10-04).

No AOT asset ships: the Swift runtime’s specialization of the .aimodel equals the AOT asset bit for bit (above). To compile one for the Mac anyway (15.55 GB, the flag is required):

xcrun coreai-build compile kev_4b_decode_fp16_metal_pf128.aimodel --output aot --preferred-compute gpu \
    --platform macOS --architecture h16c --expect-frequent-reshapes

Use it

Swift, with the Kev package (macOS 27; the system CoreAI framework, Accelerate and swift-transformers’ tokenizer), on a download of the repository:

import Kev

let bundle = URL(filePath: "Kev-4B-CoreAI/gpu-pipelined/kev_4b_decode_fp16_metal_pf128")
let kev = try await KevDecider(bundle: bundle)              // the .aimodel, specialized here (GPU, frequent reshapes)
let response = try await kev.decide(requestJSON: requestData, shared: true)   // shared: the state's whole calls run once

// questions that arrive later, on the same state
let prepared = try await kev.prepare(state: try JSONParser.parse(stateData))       // the state's whole calls, once
let later = try await kev.decide(prepared: prepared, questionsJSON: questionsData)  // only the questions' rows run

Loading the bundle specializes the graph once (16.6 s on the M4 Max) and caches it; later loads read the cache. The Mac CLI: kev run --bundle Kev-4B-CoreAI/gpu-pipelined/kev_4b_decode_fp16_metal_pf128 --asset jit --request req.json --shared --out resp.json (swift build -c release --package-path apps/Kev). The request shape and the Python read-out (conversion/kev/decide.py run --model kev-4b …, AOT assets only) are as in the Kev-0.8B card.

Reproduce

Environment: the zoo overlay venv (coreai-core 1.0.0b2, coreai-torch 0.4.1, torch 2.9.0, transformers 4.57.6), Xcode 27.0 RC; the author’s code (the oracle and the merge) in its own venv from the author’s uv.lock at tag kev-1.0. Order and flags: conversion/kev/README.md.

(cd $ZOO_WORK_ROOT/_kev/kev-src && python scripts/merge_lora_checkpoint.py \
    --lora jaredpalmer/kev-4b@591dcb5bd6d05eb0b5131ea6608f93f10243335c --out $ZOO_WORK_ROOT/_kev/merged/kev-4b-v1.0)
# the decoder (--aot adds the h16c .aimodelc the Python gates load)
python conversion/kev/export_decoder.py fp16 --model kev-4b --gdn-scan metal --prefill-chunk 128 --aot
# the published bundle: the gated .aimodel bytes and their metadata, without exporting again
python conversion/kev/export_decoder.py fp16 --model kev-4b --gdn-scan metal --prefill-chunk 128 --metadata-only \
    --from-aimodel <bundles>/kev_4b_decode_fp16_metal_pf128/kev_4b_decode_fp16_metal_pf128.aimodel \
    --bundle-dir <ship>/kev_4b_decode_fp16_metal_pf128

Recipe: recipe.toml. Port notes: knowledge/kev-port.md.

Other formats

On the Hub (2026-10-04), not run here:

No other Core AI conversion of Kev was listed on the Hub on 2026-10-04.

License

Apache-2.0: the adapter and the head (jaredpalmer/kev-4b) and the base (Qwen/Qwen3.5-4B-Base); the bundle inherits it, and the Hugging Face repository carries the base’s LICENSE. The author’s package runs only in the oracle and the merge, at gate time. The fixture file’s terms are Kev-0.8B’s: SemIf authored144 under MIT with its notice, transfer-v4 records by reference, 11 of the 20 records written for the port published and 9 withheld (an invented name in each was found in use on the web).

Limits

From the author’s card: English only; text generation, chat and fully automated consequential decisions about people are out of scope; questions that need facts that are neither in the state nor general knowledge are out of scope; the temperature was fitted on the author’s development rows (the card explains how to measure and refit it). This port adds: