Core AI model zoo

Kev-0.8B — Core AI

🤗 mlboydaisuke/Kev-0.8B-CoreAI · Apache-2.0 · source jaredpalmer/kev-0.8b (tag v1.0, commit bf75a6a) · base Qwen/Qwen3.5-0.8B-Base (revision dc7cdfe)

A decision model. Give it a state (text, or a JSON object or array) and typed questions: noul (yes / no), choice (named options) or score (ordered levels). It returns a probability for every option of every question. It never generates text. Requests and responses use the SystemOne-compatible request shape: {model, state, questions} in, {model, answers, usage} out, one typed decision per question.

Jared Palmer trained it as a rank-16 LoRA adapter and a pointer head on Qwen3.5-0.8B-Base (24 layers: 18 Gated DeltaNet and 6 full attention, hidden size 1,024). The author’s card: “It reads one document (the state) and a set of typed questions about it, and returns a calibrated probability distribution over the options supplied with each question, in a single forward pass and without generating text.” Each question is its own row: the state, then the question and its options. The head scores each option’s closing token against the question’s last token. The author’s card reports benchmark results; none of them are re-measured here.

This port merges the adapter into the base with the author’s own merge script and exports the text backbone as one Core AI graph. The graph takes 128 token ids per call and returns the final-norm hidden state at every position; it has no vocabulary head. Each Gated DeltaNet layer runs its recurrence in an fp32 Metal kernel, one GPU dispatch per layer per call. Weights are fp16 (1.51 GB). The host runs the pointer head from the author’s head weights. The gate is probability parity with the author’s own fp32 code, on every option of every question, on a fixture of 434 questions and on a held-out set of 130.

It runs on the Mac GPU and on an iPhone 18 Pro. On both, the shipped .aimodel is specialized where it runs.

Readout contract

One row per question, every row from zeroed states. The ids are the author’s kev.model.encode in its row form (rows_of), with the base model’s tokenizer:

[<|fim_prefix|>] + render(state) + [<|fim_middle|>] + render(instructions)
  + for each option: [<|box_start|>] + option_text + [<|box_end|>]
  + [<|fim_suffix|>]

conversion/kev/oracle_kev.py runs the author’s package unchanged (Checkpoint.load("cpu", fp32), kev.model.admit, the row-form forward, one CPU thread) and records the reference: fixtures-kev-0.8b.json. The fixture is 384 records and 434 questions: lines 1–60 of the author’s transfer-v4 development file (evals/v4/transfer-v4/development.jsonl, tag kev-1.0; MMLU, 4 options), records 1–20 of each of its seven other sources (140, choice and noul), its score records 1–20, the 144 choice items of SemIf authored144, and 20 records written for the port (70 questions: noul, choice with 2–10 options, score with 2–7 levels, JSON states, three states of 1,431–1,746 tokens, one record of 8 questions). The longest row is 1,802 tokens. Fourteen questions have an oracle top-2 margin of 0.02 or less (near-ties) and are listed apart. The held-out set is the next 130 transfer-v4 records (mmlu 61–100, records 21–30 of each other source, score records 21–40), with its own oracle run.

Core AI shape

conversion/kev/qwen3_5_kev_decoder.py: the overlay’s stateful Qwen3.5 text decoder (Qwen3_5StatefulForCausalLM) on the merged weights, the vocabulary head replaced by the identity, the output the final-norm hidden state at every position. One function, main, at a static 128 tokens. Each Gated DeltaNet layer runs its recurrence in the overlay’s fp32 Metal chunk kernel (qwen3_5_gdn_metal): one GPU dispatch per layer per call, the recurrent state kept in fp32 through the call.

  name shape, type
inputs input_ids [1, 128] int32
  position_ids [1, seq] int32, the ramp 0 .. seq − 1
states keyCache, valueCache [6, 1, 2, ctx, 256] fp16, ctx up to 4,096
  convState, recState [18, 1, 6144, 3], [18, 1, 16, 128, 128] fp16
output hidden [1, 128, 1024] fp16, every position

Per row of T ids, from zeroed states: ⌈T / 128⌉ calls, call k with ids[128k : 128k + 128] and position_ids 0..128k + 127. The last call is padded with <|endoftext|> (248044) and its padded rows are dropped. The rows of every call, cut to T, are the backbone’s final-norm last_hidden_state [T, 1024]. The bundle’s metadata.json says kind: decision-backbone and language.prefill_chunk: 128, and carries this order, the row layout, the delimiter ids, the render rules, the head formula and the response shape.

Host. Builds the rows, runs the calls, converts the hidden rows at the readout positions to float64, applies the head (head/head.safetensors, head/kev_head.json) and the per-question softmax in float64, rounds p to fp32 once, and writes the response. conversion/kev/decide.py is the Python reference; apps/Kev is the Swift one. With shared: true the state’s whole calls run once. A prepared state (prepare(state:), then decide(prepared:questionsJSON:)) runs the same calls in two steps, so questions that arrive later skip the state.

No engine is involved: the hosts drive the low-level runtime (AIModel + loadFunction, four zeroed states per row). No runtime patch.

Measured (Apple M4 Max, macOS 27.0 26A428, 2026-10-03 – 10-04)

The bar, fixed before any graph ran: the argmax equal to the oracle’s on every question whose oracle top-2 margin is above 0.02 (near-ties listed apart), max |Δp| ≤ 0.02 over every option of every question, the mean over rows of each row’s mean |Δp| ≤ 0.002, and every process re-running its opening row bit for bit. The Python gates load the AOT .aimodelc (coreai-build compile … --platform macOS --preferred-compute gpu --architecture h16c --expect-frequent-reshapes) with SpecializationOptions.default().

Before any graph: the merge and fp32 torch

The graph alone on the Mac GPU

The oracle’s ids in, the fp16 hidden rows read through the author’s fp32 head:

set questions argmax (margin > 0.02) near-ties agreeing max |Δp| mean of row means bar
fixture 434 420/420 13/14 0.0124 0.00096 PASS
held out 130 125/125 5/5 0.0058 0.00083 PASS

The near-tie that flips is a SemIf item with an oracle margin of 0.0032; it moves by 0.0026. The worst row is a SemIf item (0.0124). Every value is finite and every process’s re-run is bit-equal. Two red arms move p on the same graph: a state swapped for another record’s moves 4 of 5 argmaxes (max |Δp| 0.905); a grammatical “not” in five yes / no questions moves 2 of 5 (0.759). Transcript: gate-kev-0.8b-readout.json.

From the request: Python and Swift

Time per decision on the Mac

Swift, Release CLI, the AOT asset, in one machine-wide GPU lock window (2026-10-04 12:13–12:27 JST). Two processes per graph counted only when no other GPU job ran and the one-minute load average was at most 12; each made 10 decisions per item after one warm-up, and the table gives the median (p10–p90) over the 20. The earlier upload’s graph (below, under Bundle) ran in the same window:

request tokens calls ms p10–p90 earlier upload’s graph, ms
one question 94 1 29.9 29.6–30.1 100.1
one question 380 3 89.2 88.7–90.4 404.4
one question 1,518 12 357.0 355.4–359.3 1,600.1
one question 1,802 15 447.9 446.6–449.8 1,897.6
five questions on one 137-token state, each row from zero 240 10 296.6 294.6–298.5 845.7
the same five questions, the state run once (shared) 240 6 186.4 185.4–187.7 327.7
eight questions on that state, each row from zero 320 16 474.2 472.2–475.9 1,377.0
the same eight questions, shared 320 9 278.8 277.2–281.7 461.0
four questions on one 1,477-token state, each row from zero 1,622 49 1,455.9 1,453.7–1,460.4 6,366.1
the same four questions, shared 1,622 16 485.4 483.8–489.4 1,749.0
a follow-up question on a prepared 137-token state — — 30.3 — 34.1
a follow-up question on a prepared 1,477-token state — — 31.5 — 51.5

Preparing the two states took 34.1–34.7 ms and 335.9–339.6 ms (the two processes’ medians). In round 11’s lock window (09:43–10:45 JST) the .aimodel specialized by Swift, the shipped path, took 29.7 ms for the 94-token question and 186.5 ms for the five questions shared, against 30.0 and 189.1 ms for the AOT asset; the Python reference took 33.4 and 207.1 ms. Transcript: gate-kev-0.8b-timing-mac.json.

iPhone 18 Pro (iOS 27.0 24A437, Core AI arch h19p, 2026-10-04)

Through apps/KevGate, a headless gate app on the Swift host, Release, without the increased-memory-limit entitlement (3,529 MB available at launch). The phone was on USB power, the battery at 80 % and charging; every bench item started at thermal state nominal. Every p the app wrote was re-scored on the Mac from its bit patterns.

set questions argmax (margin > 0.02) near-ties agreeing max |Δp| mean of row means bar
fixture 434 420/420 13/14 0.0110 0.00097 PASS
held out 130 125/125 5/5 0.0059 0.00083 PASS

The shared prefix equals the direct run on all 20 multi-question records, and the re-run of the opening record is bit-equal. No hidden row equals the Mac’s (a different GPU): their p differ by at most 0.0026, with every argmax equal (434/434). The fixture pass peaked at a 389 MB footprint.

request tokens ms earlier upload’s graph, ms
one question 94 37.8 156.7
one question 380 110.6 627.9
one question 1,518 444.1 2,500.9
one question 1,802 559.1 2,977.8
five questions on one 137-token state, each row from zero / shared 240 361.9 / 223.2 1,331.4 / 507.4
eight questions on that state, each row from zero / shared 320 583.2 / 333.4 2,167.4 / 723.3
four questions on one 1,477-token state, each row from zero / shared 1,622 1,822.2 / 613.9 10,061.6 / 2,769.8
the five questions on a prepared state (the prepare: 38.8 / 211.6 ms) 240 183.4 295.5
the four questions on a prepared state (the prepare: 415.3 / 2,451.4 ms) 1,622 196.5 342.4

An earlier session on the same phone, with an earlier KevGate build, gave 36.5 ms for the 94-token question against 157.1 ms for the earlier upload’s graph. Transcript: gate-kev-0.8b-iphone.json.

Kev-4B on the phone: see the Kev-4B card.

Forms measured

Every form below was exported, compiled and gated the same way. The times come from different windows, so compare forms only within one window. Full records: gate-kev-0.8b-forms.json.

Mac (Swift Release CLI, the 94-token question, median of the counted processes):

form 94-token decision, ms window fixture max |Δp| note
Gated DeltaNet unrolled, S = 16 (the earlier upload) 100.1 round 15 0.0117  
unrolled, S = 32 139.0 round 11 0.0118  
unrolled, S = 64 — — 0.0100 not timed in a lock window (68.06 ms per call on a shared GPU)
in-graph chunk scan, S = 16 — — 0.0135 not timed in a lock window (10.90 ms per call on a shared GPU)
in-graph chunk scan, S = 32 — — — fails: non-near-tie argmax 112/127, max |Δp| 0.125, non-finite rows; breaks in fp32 torch too
Metal kernel, S = 16 / 32 / 64 63.6 / 42.7 / 41.5 round 11 0.0115 / 0.0134 / 0.0115  
Metal kernel, S = 128 (this release) 29.9 round 15 0.0124  
Metal kernel, query length 2..512, host calls of at most 512 ids 25.2 round 15 0.0112 not shipped: memory grows (below)
the same graph, host calls of at most 256 / 128 ids 25.1 / 25.0 round 15 0.0060 / 0.0060 (130 rows) not shipped
int8, every linear (unrolled S = 16) — — 0.0431 fails the bar (Precision)
any form on the Neural Engine — — — not run: the recurrence’s fp32 state is not an ANE element type (qwen3.5-static-ane.md)

iPhone 18 Pro (KevGate bench, the 94-token question):

form ms session fixture max |Δp|
unrolled, S = 16 (the earlier upload) 156.7 round 13, stage C 0.0102
Metal kernel, S = 16 / 32 / 64 92.4 / 58.5 / 49.3 round 13, stage A 0.0110 / 0.0133 / 0.0098
Metal kernel, S = 128 (this release) 37.8 round 13, stage C 0.0110
Metal kernel, query length 2..512, host calls of at most 512 ids 32.6 round 13, stage C gate stopped by memory (below)

Several short questions on one state. For these, the S = 64 graph takes less time than S = 128. In round 11’s Mac window, five questions on one 137-token state, shared, took 145.0 ms with S = 64 and 189.1 ms with S = 128; one 94-token question took 41.5 and 30.0 ms. On the iPhone (round 13, stage A) the same pairs took 180.3 against 224.4 ms and 49.3 against 36.5 ms. The S = 64 bundle is not in this release: it would need a new export and a new gate.

Why the dynamic-length graph does not ship. The graph whose query length is dynamic (2..512) passed the gate on the Mac and decided the 94-token question in 25.2 ms. A process that keeps changing its call length keeps growing, though. Cycling four call lengths (128, 256, 384, 512) three times added 467.1, 166.9 and 168.5 MB per lap in Swift and 460.7, 169.5 and 169.6 MB in Python. Alternating between two lengths did not grow after the opening calls; bringing in a third added 19.1 MB, then 1.3 MB. On the iPhone the gate run reached 3,497 MB with 43 MB available after 439 calls and stopped writing. A 2-second sleep gave back part of the memory and loading the function again gave back none; creating the AIModel again gave it back. Records: gate-kev-0.8b-forms.json; notes: knowledge/kev-port.md.

Precision

int8 was measured and is not shipped. It was measured on the unrolled S = 16 graph, the earlier upload; the Metal-kernel graphs were not measured in int8. int8 per block of 32 over every decoder linear (symmetric_with_clipping, weights only; the embedding table, the Gated DeltaNet conv1d and every norm fp16) fails the bar:

decoder (unrolled S = 16) main.mlirb bytes fixture max |Δp| / mean held out max |Δp| / mean bar
fp16 1,506,481,909 0.0117 / 0.00097 0.0057 / 0.00081 PASS
int8lin: every linear int8 1,040,063,350 0.0431 / 0.00326 0.0273 / 0.00294 FAIL
int8mix: int8 except layers 0–11 1,273,276,920 0.0089 / 0.00103 0.0043 / 0.00082 PASS

int8lin breaks the bar on 10 fixture rows (above 0.01: 51 rows; fp16: 1). An fp32 torch bisect with the exporter’s own int8 weights (conversion/kev/int8_bisect_torch.py, 60 rows) reproduces it (0.0420 against the graph’s 0.0431). Its rule, written before any result, asked for the worst row below 0.010 with at most half of the linear weights kept fp16; no set met it (the best: four kinds of linear kept fp16, 45.5 % of the weights, 0.0141). Layers 0–11 kept fp16, exactly half and outside the rule’s search, give 0.0046 on the bisect rows; that is int8mix. It keeps 85 % of fp16’s bytes, and it was chosen on the fixture, so it does not ship (gate-kev-0.8b-int8.json).

What the compiled asset holds. --expect-frequent-reshapes adds an fp16 copy of every linear weight to the compiled asset: for the unrolled graph the iPhone h19p asset is 2,505,674,625 bytes with it and 1,506,400,620 without, and this release’s Mac h16c asset is 2,502,047,194 bytes against a 1,505,385,733-byte main.mlirb. On a toy (an embedding and two linears) compiled with a fixed input length, the int8 version’s resources.bin is 12,582,936 bytes on macOS and iOS, with and without the flag, the bytes of the fp16 toy compiled without it, against an int8 main.mlirb of 10,623,080 bytes. An int8 .aimodel downloads smaller; the compiled asset is not smaller by the same amount (../kev-4b/gate-kev-4b-int8.json).

⬇️ Bundle

mlboydaisuke/Kev-0.8B-CoreAI, one folder under gpu-pipelined/:

file what bytes sha256
kev_0_8b_decode_fp16_metal_pf128.aimodel/main.mlirb the decoder, fp16 1,505,385,733 19d5a480…2c3f684e
metadata.json kind: decision-backbone, the readout contract, the gate 12,590 6e50e41a…47282227
tokenizer/tokenizer.json the base model’s, verbatim 12,807,196 fe000e3e…d50d2927
tokenizer/tokenizer_config.json the base model’s, verbatim 16,712 e611fbcc…c47885de
head/head.safetensors the pointer head: q / k weight [256, 1024] and bias, fp32 2,099,576 12038b02…00893a42
head/kev_head.json head size, scale, temperature, delimiter ids, provenance 2,046 73d51070…9d5c9d7c

The repository root carries the base model’s config.json, the adapter’s adapter_config.json and the Apache-2.0 LICENSE, verbatim. SHA256SUMS lists every file. Its Hub revision is 3792e4b (2026-10-04).

The earlier upload. Hub revision 7793533 held another graph of this port, kev_0_8b_decode_fp16_pf16: the Gated DeltaNet recurrence unrolled step by step, 16 tokens per call. It passed the same gate (fixture max |Δp| 0.0117, held out 0.0057). In the Mac window above it took 100.1 ms for the 94-token question against 29.9 ms for this release, and on the iPhone 156.7 ms against 37.8 ms.

No AOT asset ships: the Swift runtime specializes the .aimodel correctly on the Mac and on the phone (the JIT rows above). To compile one for the Mac anyway:

xcrun coreai-build compile kev_0_8b_decode_fp16_metal_pf128.aimodel --output aot --preferred-compute gpu \
    --platform macOS --architecture h16c --expect-frequent-reshapes

Use it

Swift, with the Kev package (macOS 27 / iOS 27; the system CoreAI framework, Accelerate and swift-transformers’ tokenizer), on a download of the repository:

import Kev

let bundle = URL(filePath: "Kev-0.8B-CoreAI/gpu-pipelined/kev_0_8b_decode_fp16_metal_pf128")
let kev = try await KevDecider(bundle: bundle)              // asset: nil = the .aimodel, specialized here (GPU, frequent reshapes)
let response = try await kev.decide(requestJSON: requestData, shared: true)   // shared: the state's whole calls run once
print(PythonFormat.dumps(response, asciiOnly: false))
// {"model": "kev_0_8b_decode_fp16_metal_pf128", "answers": {"team": {"type": "choice", "choice": "billing", "confidence": …,
//  "probabilities": {…}}, "urgent": {"type": "noul", "noul": …}}, "usage": {"input_tokens": …, "output_tokens": …}, "latency_ms": …}

// questions that arrive later, on the same state
let prepared = try await kev.prepare(state: try JSONParser.parse(stateData))       // the state's whole calls, once
let later = try await kev.decide(prepared: prepared, questionsJSON: questionsData)  // only the questions' rows run

The same from the Mac CLI (swift build -c release --package-path apps/Kev builds kev):

kev run --bundle Kev-0.8B-CoreAI/gpu-pipelined/kev_0_8b_decode_fp16_metal_pf128 --asset jit \
    --request req.json --shared --out resp.json

A request, in the SystemOne-compatible request shape:

{"model": "kev-0.8b",
 "state": "I was charged twice for my last order. Please refund one of the charges.",
 "questions": {
   "team": {"type": "choice", "instructions": "Which team should handle this?",
            "criteria": {"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}},
   "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}}}

Python: conversion/kev/decide.py run --model kev-0.8b --bundle <folder> --aimodelc <asset> --request req.json [--shared] --out resp.json is the gates’ own read-out with coreai.runtime. It loads an AOT .aimodelc (compile one with the command above); the Python gates never used the Python runtime’s JIT.

Reproduce

Environment: the zoo overlay venv (coreai-core 1.0.0b2, coreai-torch 0.4.1, torch 2.9.0, transformers 4.57.6), Xcode 27.0 RC. The author’s code (the oracle and the merge) runs in its own venv built from the author’s uv.lock at tag kev-1.0 (torch 2.8.0, transformers 5.17.0, peft 0.21.0). The steps, in order, with every flag, are in conversion/kev/README.md.

# the merged checkpoint, with the author's script (tag kev-1.0)
(cd $ZOO_WORK_ROOT/_kev/kev-src && python scripts/merge_lora_checkpoint.py \
    --lora jaredpalmer/kev-0.8b@788ddbdd65715bb03a56788c822f6c632c9a551d --out $ZOO_WORK_ROOT/_kev/merged/kev-0.8b-v1.0)
# the decoder (--aot adds the h16c .aimodelc the Python gates load)
python conversion/kev/export_decoder.py fp16 --gdn-scan metal --prefill-chunk 128 --aot
# the published bundle: the gated .aimodel bytes and their metadata, without exporting again
python conversion/kev/export_decoder.py fp16 --gdn-scan metal --prefill-chunk 128 --metadata-only \
    --from-aimodel <bundles>/kev_0_8b_decode_fp16_metal_pf128/kev_0_8b_decode_fp16_metal_pf128.aimodel \
    --bundle-dir <ship>/kev_0_8b_decode_fp16_metal_pf128

Recipe: recipe.toml. Port notes: knowledge/kev-port.md.

Other formats

On the Hub (2026-10-04), not run here:

No other Core AI conversion of Kev was listed on the Hub on 2026-10-04.

License

Apache-2.0: the adapter and the head (jaredpalmer/kev-0.8b) and the base (Qwen/Qwen3.5-0.8B-Base); the bundle inherits it, and the Hugging Face repository carries the base’s LICENSE. The author’s package runs only in the oracle and the merge, at gate time; it is not part of the bundle. The fixture file carries its own terms: its 144 SemIf authored144 items are MIT (github.com/TheoLeeCJ/SemIf at ca3ba65f), with their copyright and permission notice; the transfer-v4 records are references to the author’s file (line and row hashes), not its text; of the 20 records written for the port, 11 are published and 9 are withheld (an invented name in each was found in use on the web), keeping their ids and numbers.

Limits

From the author’s card: English only; text generation, chat, tool-call routing and fully automated consequential decisions about people are out of scope; the temperature was fitted on the author’s development rows (the card explains how to measure and refit it on one’s own data). This port adds: