Core AI model zoo

d1-omni-600M — Core AI

🤗 mlboydaisuke/d1-omni-600M-CoreAI · LFM Open License v1.0 · source LiquidAI/d1-omni-600M (revision 414f8d6)

A decision model from Liquid AI. Give it a state (text or JSON, with images or one voice clip) and named questions: noul (yes / no), choice (named options) or score (ordered levels). It returns a probability for every option of every question, read from one forward pass. It does not generate text.

The checkpoint has 587 M parameters: a bidirectional LFM2.5-Encoder-350M trunk with a decision head, a SigLIP2 vision tower and a 17-layer FastConformer audio tower. A request carries images or audio, not both. Liquid AI’s release blog calls it an early research release and reports speed for d1-3B only.

This port re-authors the three networks in plain PyTorch from model.safetensors and exports each one as a Core AI graph in fp16. The host builds the token rows, prepares the images and the audio, runs the graphs, and applies the publisher’s temperatures and softmax. Every stage is gated against the publisher’s own model in fp32 on the CPU.

What this repository holds

Three graphs. The decision graph runs at seven static lengths; a host picks the smallest one that holds a row. The vision graph takes one image crop per call. The audio graph runs at four clip lengths. Their MLIR debug locations were removed before shipping (the operations are unchanged).

folder graph .aimodel bytes h19p .aimodelc bytes
decide-fp16-L64 decision, 64 positions 761,631,060 761,747,067
decide-fp16-L128 decision, 128 positions 761,647,561 761,763,726
decide-fp16-L256 decision, 256 positions 761,680,330 761,796,331
decide-fp16-L512 decision, 512 positions 761,745,868 761,861,994
decide-fp16-L1024 decision, 1,024 positions 761,876,787 761,993,099
decide-fp16-L2048 decision, 2,048 positions 762,139,098 762,255,228
decide-fp16-L4096 decision, 4,096 positions (attention in two blocks of 2,048 keys) 762,676,555 — (not shipped: on the iPhone 18 Pro its first call exceeds the per-process memory limit; the .aimodel runs)
vision-fp16 one image crop → up to 256 prefix rows 188,218,374 188,303,983
audio-fp16-5s a clip of up to 5 s → up to 63 prefix rows 224,920,212 225,226,318
audio-fp16-10s up to 10 s → up to 125 rows 225,049,263 225,355,344
audio-fp16-20s up to 20 s → up to 250 rows 225,305,277 225,611,159
audio-fp16-30s up to 30 s → up to 375 rows 225,561,282 225,867,256

Each folder is in macos/ and ios/ (the same .aimodel; the device specializes it when it loads it) and, except decide-fp16-L4096, in ios-h19p/ (an ahead-of-time compile for the iPhone 18 Pro GPU, which other phones refuse). Each folder also holds metadata.json (its contract) and its host files: tokenizer/ for the decision graphs (the publisher’s files, unmodified), position_table.f32 for the vision graph, mel_filters_128x257_f32.bin for the audio graphs. The repository root holds the shared metadata.json, tokenizer/, host/ (the host files and the Swift host’s sources), reference/ (the public fixture and the publisher’s numbers on it), LICENSE and NOTICE.md.

Source: metadata.json staged, MANIFEST.json.

Use it

Swift: the D1Omni package is in host/, a copy of the zoo’s apps/D1Omni (macOS 27, the system CoreAI framework, Accelerate, ImageIO and swift-transformers 1.3.3). An app whose package folder holds a download of this repository depends on it by path. SwiftPM names a path package after its folder, so the package is host:

// swift-tools-version: 6.1
import PackageDescription

let package = Package(
    name: "Ask", platforms: [.macOS("27.0")],
    dependencies: [.package(path: "d1-omni-600M-CoreAI/host")],
    targets: [.executableTarget(name: "Ask", dependencies: [.product(name: "D1Omni", package: "host")])])

Sources/Ask/main.swift, run from the package folder:

import Foundation
import D1Omni

let dir = URL(filePath: "d1-omni-600M-CoreAI/macos")
let d1 = try await D1Omni(folders: D1Omni.folders(macos: dir), media: D1Omni.MediaFolders.find(macos: dir))
let questions: JSONValue = try JSONParser.parse(Data(#"""
    {"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
     "team": {"type": "choice", "instructions": "Which team should handle this?",
              "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults"}}}
    """#.utf8))
let text = try await d1.systemOne(state: .string("I was charged twice this month, please refund one of them."),
                                  questions: questions)
let heard = try await d1.systemOne(state: nil, questions: questions,                 // 16 kHz mono 16-bit WAV
                                   audio: URL(filePath: "d1-omni-600M-CoreAI/reference/audio/aud_01.wav"))
let room: JSONValue = try JSONParser.parse(Data(#"{"bed": {"type": "noul", "instructions": "Is there a bed in the room?"}}"#.utf8))
let seen = try await d1.systemOne(state: nil, questions: room,                       // PNG or JPEG files
                                  images: [URL(filePath: "d1-omni-600M-CoreAI/reference/images/img_01.png")])
print(PythonFormat.dumps(text, asciiOnly: false))

The same from the Mac command line. swift build -c release --package-path d1-omni-600M-CoreAI/host builds d1omni in d1-omni-600M-CoreAI/host/.build/release/:

d1omni ask --bundle-dir d1-omni-600M-CoreAI/macos --state "I was charged twice this month, please refund one of them." \
    --question '{"type": "noul", "instructions": "Is the customer asking for a refund?"}'
d1omni ask --bundle-dir d1-omni-600M-CoreAI/macos --image d1-omni-600M-CoreAI/reference/images/img_01.png \
    --question '{"type": "noul", "instructions": "Is there a bed in the room?"}'

Python, with the reference host conversion/d1_omni/host.py from a clone of the zoo (coreai-model-zoo/, beside the download) and the Core AI Python runtime (coreai-core), on the Mac. The Python runtime loads compiled assets, so compile the decision graphs before you run it:

for L in 64 128 256 512 1024 2048 4096; do
  xcrun coreai-build compile d1-omni-600M-CoreAI/macos/decide-fp16-L$L/d1_omni_decide_fp16_L$L.aimodel \
      --output aot --platform macOS --preferred-compute gpu --architecture h16c
done
import asyncio, json, sys
sys.path.insert(0, "coreai-model-zoo/conversion/d1_omni")
import host                                   # the publisher's prompt rules and readout, written out
import coreai.runtime as rt

async def ask(state, questions):
    tok = host.RawTokenizer("d1-omni-600M-CoreAI/tokenizer/tokenizer.json")
    host.check_token_ids(tok)
    gpu = rt.SpecializationOptions.from_preferred_compute_unit_kind(rt.ComputeUnitKind.gpu())
    rows = host.request_rows(tok, state, questions)                     # one row per question
    probabilities = []
    for row in rows:
        L = host.bucket_for(row.positions, host.ALL_BUCKETS)            # the smallest of 64 ... 4096
        model = await rt.AIModel.load(f"aot/d1_omni_decide_fp16_L{L}.h16c.aimodelc", gpu)
        inputs, markers = host.graph_inputs(row, L)
        out = await model.load_function("main")({k: rt.NDArray(v) for k, v in inputs.items()})
        scores = out["scores"].numpy().reshape(-1)
        probabilities.append(host.probabilities_from_logits(scores[markers], row.question, row.calibrate))
    return host.response(rows, probabilities)

print(json.dumps(asyncio.run(ask(
    "I was charged twice this month, please refund one of them.",
    {"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"}}))))

Contract

Rows. One row per question, the publisher’s prompt.py (host.py copies it):

[1 <|startoftext|>, 17 <|reserved_7|>] + state + [18 <|reserved_8|>] + instructions
  + for each option: [19 <|reserved_9|>, 16 <|mask|>] + " " + option + [20 <|reserved_10|>]
  + [21 <|reserved_11|>]

Every text piece is tokenized without special tokens, after <|name|> in caller text is rewritten to <¦name¦>. A JSON state is written with json.dumps(ensure_ascii=False). A state that does not fit is cut at its end. <|pad|> is 0 and <|im_end|> is 7; a host checks the nine ids against the tokenizer at load. Option text: choice → name: description; score → level i: description; noul → false: …, true: … from the criteria, or false: no, the statement does not hold / true: yes, the statement holds without them. Without criteria, an image request writes false: no / true: yes. An audio request always does, and writes choices as option_000: description, as those questions were trained.

Decision graph (main, static length L):

tensor shape dtype value
input_ids [1, L] int32 0 on the media prefix [0, P), the row’s ids on [P, P + n), 0 after
prefix_embeds [1, L, 1024] float32 the image or audio prefix on [0, P), 0.0 elsewhere
pad_mask [1, L] float32 1.0 on [0, P + n)
prefix_mask [1, L] float32 1.0 on [0, P)
keep_right [1, L] float32 0.0 at P − 1 when P > 0, 1.0 elsewhere
qtype_onehot [1, 3] float32 choice / score / noul
→ scores [1, L] float32 read at P + each <|mask|> position

L is the smallest of 64, 128, 256, 512, 1,024, 2,048, 4,096 that holds P + n. With an image the text is cut to 896 positions, with audio to 15,360; all positions together stay within 16,384.

Readout. The K marker scores of a text row are divided by the temperature of its type and option count, then softmaxed in fp32. Image and audio rows take the plain softmax. A noul is scored as [false, true] and reported as [yes, no].

temperature key T   temperature key T
choice:2 1.7465145587921143   noul:2 1.6663223505020142
choice:3-5 1.3998981714248657   score:3-5 1.7301132678985596
choice:6-10 1.1751071214675903   score:6-10 1.0
choice:11+ 1.372515082359314   choice / score / noul (any other K) 1.0

The answer is the publisher’s: noul → P(yes); choice → the argmax name, a confidence and every p; score → the expected level, a confidence, every p and the legend. usage.input_tokens counts every position the graphs read.

Vision graph (main, one crop): pixel_values [1, 1024, 768] (16-px patches row-major, (x − 127.5) / 127.5), pos_embed [1, 1024, 768] (the 16 × 16 position table, position_table.f32, resized to the crop’s patch grid, bilinear with antialias), patch_mask [1, 1024], unshuffle_index int32 [256, 4] → prefix [1, 256, 1024]. An image whose rounded area exceeds 2 × 512² is cut into the 512-px tiles of the closest grid (2 to 10 tiles) plus a thumbnail; a smaller one is one crop. A 384 × 384 image is 144 prefix rows.

Audio graph (main, one clip bucket of 5 / 10 / 20 / 30 s): mel [1, 128, F] and four time masks → prefix [1, T, 1024]. The clip is 16 kHz mono, cut at 30 s and padded to 0.5 s, not resampled. The mel: pre-emphasis 0.97; a centred 512-point STFT, hop 160, Hann(400) centred in 512; power; 128 Slaney mel bins (mel_filters_128x257_f32.bin); log(x + 2⁻²⁴); each mel row normalized over the clip’s frames, (x − mean) / (std + 1e-5). F = 501 / 1,001 / 2,001 / 3,001 columns; a clip takes the smallest bucket whose F holds its 1 + n // 160 frames.

Every value above is in metadata.json.

Measured

Time per decision. Apple M4 Max, macOS 27.0 (26A428), the GPU, the Core AI Python runtime on the h16c compiles of these graphs. One decision = the graph inputs, the graph call(s), the output copy, the marker read, the temperature and the softmax. Tokenizing, image decoding and the mel are not included. Each form ran 30 decisions after 5 warm-up rounds, interleaved with the other forms, inside a machine-wide measurement window (no other GPU job).

request positions graphs Mac M4 Max, median ms (p10–p90) iPhone 18 Pro, median ms, .aimodel / h19p compile
one question 47 decision L64 7.35 (7.20–7.52) 10.79 / 10.70
three questions, one call each 47 / 56 / 51 decision L64 × 3 21.97 (21.67–23.36) 32.52 / 32.07
one question on a 3.4k-token state 3,454 decision L4096 299.2 (302.53 and 295.87 in two windows) 758.2 / — (the compile exceeds the memory limit, see below)
one question on a 384 × 384 image 144 + 39 vision + decision L256 38.3 (38.50 and 38.12 in two windows) 57.67 / 57.72
one question on 9.6 s of audio 121 + 63 audio 10 s + decision L256 24.94 (24.76–25.06) 27.02 / 27.30
one question on 2.8 s of audio 35 + 63 audio 5 s + decision L128 17.72 (17.49–17.96) not measured
one question on 28 s of audio 350 + 63 audio 30 s + decision L512 43.9 (43.83 and 43.92 in two windows) not measured

The single-window rows ran in window d1d-r12 (2026-10-08 20:09:31–20:09:50 JST) on these stripped graphs. The two-window rows ran in windows d1d-r7-run1 and d1d-r7-run2 (12:39–12:48 JST) on the same graphs before their debug locations were removed: the operations and the compiled weights (resources.bin, by sha256) are the same, and these three were not timed again. Source: gate-d1-omni-600m-timing-mac.json (zoo) ← results/timing/ranking_r12.json, ranking.json.

iPhone 18 Pro (iPhone19,2, iOS 27.2.0 build 24B5099f, on USB power at 80 % charge, thermal state nominal on every decision). The Swift host ran in a gate app on the phone: 20 decisions per form after 3 warm-up decisions, each series under 20 s. The .aimodel column is the graph specialized on the phone at load (0.40–1.79 s cold); the h19p column is the ahead-of-time compile (0.29–1.65 s cold). On a subset of the fixture (78 rows: 60 text, 9 image, 9 audio) every row passed the gate above with the .aimodel graphs (max |Δp| 0.0059, mean 0.00090) and the h19p compiles gave the same bits on the 74 rows they ran. The L4096 compile loaded but its first call was killed at a footprint of 3,306 MB, the phone’s per-process limit, so it is not shipped: on the phone the 4,096-position graph runs as the .aimodel. The phone’s bits differ from the Mac’s (max |Δp| 0.0040 between them on the same rows; the host arrays are bit-equal). Source: gate-d1-omni-600m-iphone.json ← results/iphone_bench.json, iphone_parity_jit.json, iphone_parity_aot.json.

The Swift host (D1Omni, Release) gave the same times as the Python runtime with the .aimodel specialized at load (JIT) and with the AOT compile: one question 16.14 / 16.16 ms and three questions 48.46 / 48.59 ms (both at L256, before L64 / L128 existed), the 3.4k-token state 301.21 / 300.84 ms, the image 37.37 / 37.47 ms, 9.6 s of audio 24.48 / 24.78 ms. Its outputs were bit-equal between JIT and AOT on every timed call. Source: gate-d1-omni-600m-timing-mac.json ← ranking_r8.json, ranking_r9.json.

Gate. The oracle is the publisher’s model (AutoModel.from_pretrained(..., trust_remote_code=True), probabilities()), fp32 on the CPU. The bar for each row: the argmax equal to the oracle’s when the oracle’s top-2 margin is above 0.02 (near ties are reported apart), max |Δp| ≤ 0.02 over the options, and the mean of the rows’ max |Δp| ≤ 0.002. Every gate also runs a control that judges each row against another row’s oracle; it must fail, and it does.

graphs (Mac GPU, these stripped compiles) rows argmax near ties max |Δp| mean of row max
decision L64 64 64/64 — 0.0058 0.00095
decision L128 (278 + 20 shorter rows padded in) 298 290/290 8/8 0.0117 0.00109
decision L256 436 424/424 12/12 0.0067 0.00098
decision L512 (19 + 20 padded + 10 with a media prefix) 49 46/46 3/3 0.0044 0.00095
decision L1024 (20 padded + 10 with a media prefix) 30 29/29 1/1 0.0071 0.00102
decision L2048 (11 + 5 padded + 10 with a media prefix) 26 25/25 1/1 0.0071 0.00099
decision L4096 (4 + 5 padded + 11 with a media prefix) 20 20/20 — 0.0038 0.00064
image rows: vision → decision 46 46/46 — 0.0062 0.00111
audio rows: audio → decision (NumPy mel, the Swift host’s) 46 46/46 — 0.0036 0.00085
audio rows: audio → decision (the publisher’s torch mel) 46 46/46 — 0.0055 0.00067

The decision gates ran each row three times (five at L4096) and the scores never changed; the image and audio gates’ repeated calls did not change them either. Source: gate-d1-omni-600m-runtime-decide.json, -runtime-vision.json, -runtime-audio.json ← results/ship_runtime_*.json, small_runtime_audio_e2e_fp16.json.

Measured and not shipped. Decision graph at L256, one question; the same windows as above:

form gate (436 rows) one question, ms
fp16 (shipped) pass, max |Δp| 0.0067 16.81
fp32 pass, 1.1e-5 17.63
fp16 weights, fp32 compute pass, 0.0020 23.67
int8 weights (every decision linear, per block of 32) fail: argmax 429/436, max |Δp| 0.097 16.70
fp16 on the Neural Engine (52 regions) fail: argmax 409/436, max |Δp| 0.95 41.77
fp16 weights, fp32 compute, on the Neural Engine (1 region) fail: a score moves by up to 13.8 between calls 23.45

The int8 .aimodel is 468,960,581 bytes, but its compile’s resources.bin is 761,554,008 bytes, the size of the fp16 compile’s: only the download is smaller. The vision graph on the Neural Engine misses the bar (mean 0.0031); the audio graph on it passes (24 rows) at 44.45 ms per 9.6 s clip. On the iPhone 18 Pro the Neural Engine compiles of the audio and decision graphs failed to load; the .aimodel graphs with the Neural Engine preferred (placement not confirmed) gave 59.98 ms for the 9.6 s audio request (9 rows passed) and a failing decision graph (argmax 33/55). Source: gate-d1-omni-600m-forms.json, gate-d1-omni-600m-iphone.json.

Other formats

On the Hub on 2026-10-08 (21:17 JST), not run here:

No other Core AI conversion was listed.

Reference fixture

reference/records.json holds the 237 public records of the gate’s fixture and reference/oracle.json the publisher’s model’s numbers on their 360 questions. The gates also ran on measured-only records that are not published. Their sources:

records source licence
card_text, card_batch_00, card_batch_01 the publisher’s README examples LFM Open License v1.0
semif_* (144) TheoLeeCJ/SemIf authored144 at ca3ba65f MIT, Copyright (c) 2026 TheoLeeCJ (reference/LICENSE-SemIf)
tv4_000 … tv4_059 MMLU test questions (cais/mmlu at c30699e8), as lines 1–60 of jaredpalmer/kev’s transfer-v4 development file MIT, Copyright (c) 2020 Dan Hendrycks (reference/LICENSE-MMLU-MIT.txt)
own_*, long_3400 written for these ports as the conversion code (BSD-3-Clause)
img_01 own FLUX.2 klein 4B output Apache-2.0 model output
img_02, img_03 photographs from Wikimedia Commons, downscaled CC0 1.0
aud_01 … aud_15 speech synthesized with Kokoro-82M from scripts written for these ports Apache-2.0 model output

The zoo’s fixtures-d1-omni-600m.json carries the same records and numbers, with each row’s token ids, markers and bucket.

Reproduce

models/d1-omni-600m/recipe.toml in the zoo lists every step: the oracle in its own environment (transformers 5.19.0), the exports in the zoo’s environment (coreai-torch 0.4.1, coreai-core 1.0.0b2), the strip, the compiles (Xcode 27.0 RC, coreai-build 3600.83.1), the gates and the staging. The scripts are in conversion/d1_omni/; the lessons are in knowledge/d1-omni-port.md.

License

LFM Open License v1.0 (LICENSE, the publisher’s text, unmodified). The graphs here are converted and modified forms of the publisher’s model.safetensors; NOTICE.md lists the changes: the networks re-authored in PyTorch and exported in fp16, the debug locations removed, the decision graph cut into seven static lengths, the iPhone compiles, and the files written for this repository. Commercial use is licensed only to an entity below the licence’s threshold of 10 million US dollars in annual revenue (Section 5).

Limits