Community field data — measured by the Bench tab of CoreAIChat (TestFlight) on contributors’ own devices and submitted as bench-result issues (the public audit log). The app measures and builds the result blob; no number in this table was typed by a human. This is NOT a controlled-environment benchmark — background load and heat show up here as real-world variance.
Protocol pb-random-v1: fixed 128-token random prompt (seed 0) → 256 greedy
decode tokens, S=1 prefill (COREAI_CHUNK_THRESHOLD=1), 1 cold + 3 warm runs on a
freshly created engine. Cell = median across submissions of each submission’s
median warm decode tok/s; n = accepted submissions. Cells with n < 3 are
provisional (marked *). Blobs with Low Power Mode on or a serious/critical
thermal state before the run are excluded from medians (counted below).
Add your device: TestFlight → Bench tab → Run → Submit on GitHub. Your
device becomes a row here on the next aggregation
(python3 scripts/aggregate_bench.py).
Generated by scripts/aggregate_bench.py — do not edit by hand. Last run: 2026-07-03 06:15 UTC.
| Device | qwen3.5-0.8b | lfm2.5-1.2b | granite-4.0-h-1b |
|---|---|---|---|
iPhone 17 Pro (A19 Pro, iPhone18,1) |
68.4* (n=1) | — | — |
Accepted submissions: 1 · excluded by environment filter: 0 · rejected (schema/protocol): 0