Core AI model zoo

Community Benchmarks

Community field data — measured by the Bench tab of CoreAIChat (TestFlight) on contributors’ own devices and submitted as bench-result issues (the public audit log). The app measures and builds the result blob; no number in this table was typed by a human. This is NOT a controlled-environment benchmark — background load and heat show up here as real-world variance.

Protocol pb-random-v1: fixed 128-token random prompt (seed 0) → 256 greedy decode tokens, S=1 prefill (COREAI_CHUNK_THRESHOLD=1), 1 cold + 3 warm runs on a freshly created engine. Cell = median across submissions of each submission’s median warm decode tok/s; n = accepted submissions. Cells with n < 3 are provisional (marked *). Blobs with Low Power Mode on or a serious/critical thermal state before the run are excluded from medians (counted below).

Add your device: TestFlight → Bench tab → Run → Submit on GitHub. Your device becomes a row here on the next aggregation (python3 scripts/aggregate_bench.py).

Generated by scripts/aggregate_bench.py — do not edit by hand. Last run: 2026-07-03 06:15 UTC.

Decode tok/s (median warm)

Device qwen3.5-0.8b lfm2.5-1.2b granite-4.0-h-1b
iPhone 17 Pro (A19 Pro, iPhone18,1) 68.4* (n=1)

Accepted submissions: 1 · excluded by environment filter: 0 · rejected (schema/protocol): 0

Contributors

@john-rocky