{
  "artifacts": [
    {
      "file": "decider-2b-vision_int8.litertlm",
      "sha256": "5fb2e19aa2066d2e3955431366bc572edb7abbc6037db53d4eaf4a4a4d0e7bd9",
      "size_mb": 3024.179
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "export_decoder_g256.py --out out/decoder_g256_r4_fp32 --step-form relu_diff (prefill ladder 1024/256/64/16/4/1 + decode, cache 4096, externalized single-token embedder, 48 state buffers) -> quantize_r4.py dyn8 decoder / embedder (recipe wi8fc) -> bundle_r4.sh dyn8 (build_bundle_g256.py with the round-3 fp16 vision encoder + adapter from quantize_fp16_r3.py: fast_vlm, image 256x256, max_num_tokens 4096, identity jinja template, no start token, stop token 248044, the snapshot's tokenizer.json, prefer_activation_type fp32 on the decoder section; then add_executor_metadata.py for the 48 state buffers, 36 linear-attention + 12 K/V) (REPRODUCE.md rounds 3-4)",
    "quantization": "dynamic int8 CHANNELWISE on every decoder FULLY_CONNECTED (activations quantized on the fly on the CPU), token embedding int8 CHANNELWISE (recipe wi8fc); vision encoder and adapter fp16 float casting; fp32 activation preference declared on the decoder section (Hub card Files; REPRODUCE.md round 4)",
    "tool": "litert-torch decoder export (fork john-rocky/litert-torch @115a13607c730c81018bb9789138a3e5e5119e3d + the qwen35 hybrid patch 0a01e2ae…, with the derived-M-RoPE rotary of scripts/mrope_derived.py installed only inside the export process) + litert-torch / litert-converter vision encoder and merger export at 256x256 + ai-edge-quantizer weight forms + litert-lm-builder bundle (build_bundle_g256.py) + add_executor_metadata.py (repro = hf-to-litertlm decider2bv_work/, REPRODUCE.md rounds 2-4)",
    "tool_version": "decoder: litert-torch 0.9.2 (patched fork clone), torch 2.12.1, transformers 5.14.1, litert-converter 0.3.0, ai-edge-litert 2.1.6, ai-edge-quantizer 0.8.0 (venv-export lock); vision: litert-torch 0.9.3, litert-converter 0.4.0, torch 2.13.0, transformers 5.14.1 with transformers 5.17.0's position-embedding taps; bundle and runtime checks: litert-lm-builder / litert-lm 0.17.1, ai-edge-litert 2.2.0 (venv-readout lock); reference = the checkpoint's own decider/vision.py in fp32 on CPU, torch 2.14.0, transformers 5.17.0 (REPRODUCE.md section 1)"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": 24603,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-decider2bv-ask-app",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.1",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "harness: hfmodels-android samples/ask (github.com/john-rocky/hfmodels-android, branch decider-vision-demo at 419a52b, the app code is commit 60b397f; release APK sha256 dd9a4a40337e66218436d3022f0932e4489bec72da372d7eb981310ed4af9a5a), hfmodels SDK 0.1.3-SNAPSHOT over com.google.ai.edge.litertlm:litertlm-android 0.16.1. One model load per picture and ONE conversation per load: turn 1 = the 256x256 PNG + question 1, turns 2-5 = text only; GenerationOptions(maxOutputTokens = 1). answer_ms = wall from stream() to the first non-empty text piece, conversation_ms = wall of createConversation(), load_ms = wall of fromPretrained(); all measured by the app",
            "09-29 12:34:38.476 11836 11876 I hfmodels: imported decider-2b-vision_int8.litertlm: 3171081088 bytes, sha256 ok, from /storage/emulated/0/Android/data/io.github.johnrocky.hfmodels.samples.ask/files/decider-2b-vision_int8.litertlm [r2b_gpu_a, the first run after the push: the SDK hashed and imported the file; the later runs load the imported file]",
            "09-29 12:47:34.944 27664 27697 I hfmodels: resolve litert-community/decider-2b-vision-LiteRT: commit=6c024e94 via EXPLICIT_REVISION, descriptor=caller@explicit/hfmodels.json sha=a064e06f",
            "09-29 12:47:45.426 27664 27664 I hfmodels: ready litert-community/decider-2b-vision-LiteRT@6c024e94 variant=int8 profile=cpu language=CPU vision=CPU prepare_ms=10480",
            "09-29 12:47:40.033 27664 27697 I tflite  : Replacing 24603 out of 25764 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2245 partitions for subgraph 0 (prefill_1024).",
            "09-29 12:47:42.701 27664 27697 I tflite  : Replacing 17709 out of 18852 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 1 (prefill_256).",
            "09-29 12:47:43.103 27664 27697 I tflite  : Replacing 15693 out of 16836 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 2 (prefill_64).",
            "09-29 12:47:43.523 27664 27697 I tflite  : Replacing 15783 out of 16926 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 3 (prefill_16).",
            "09-29 12:47:43.945 27664 27697 I tflite  : Replacing 15783 out of 16926 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 4 (prefill_4).",
            "09-29 12:47:44.086 27664 27697 I tflite  : Replacing 2071 out of 2129 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 49 partitions for subgraph 5 (prefill_1).",
            "09-29 12:47:44.092 27664 27697 I tflite  : Replacing 2125 out of 2183 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 50 partitions for subgraph 6 (decode).",
            "09-29 12:47:45.423 27664 27697 I tflite  : Replacing 1 out of 4 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2 partitions for subgraph 0 (main). [4-node graph named main]",
            "09-29 12:47:48.530 27664 28114 I tflite  : Replacing 1804 out of 1804 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0 (main). [vision encoder]",
            "09-29 12:47:50.780 27664 28114 I tflite  : Replacing 22 out of 28 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main). [vision adapter]",
            "09-29 12:47:45.439 27664 27664 I ask     : READY profile=cpu load_ms=10497 variant=int8 commit=6c024e94 downloaded=false",
            "09-29 12:47:54.159 27664 27664 I ask     : TURN 1 id=tallest letter=E ok=true answer_ms=2113.2",
            "09-29 12:47:57.180 27664 27664 I ask     : TURN 2 id=shortest letter=D ok=true answer_ms=1015.1",
            "09-29 12:47:59.893 27664 27664 I ask     : TURN 3 id=count letter=D ok=true answer_ms=709.5",
            "09-29 12:48:02.557 27664 27664 I ask     : TURN 4 id=taller letter=A ok=true answer_ms=659.6",
            "09-29 12:48:05.123 27664 27664 I ask     : TURN 5 id=leftmost letter=A ok=true answer_ms=561.5",
            "r2b_cpu_c (ask-result-1790653685.json, chart_02, 5 bars, profile cpu): load_ms 10497, conversation_ms 2790.2, answer_ms 2113.2 / 1015.1 / 709.5 / 659.6 / 561.5, letters EDDAA, equal to the Mac fresh-engine conversation 5/5, to the reference readout 5/5, to the chart's data 5/5; thermal status 0 -> 0; POLLS=135 PEAK_VmHWM_KB=3742560 MIN_MEMAVAILABLE_KB=5971500; MEMAVAILABLE_KB_BEFORE=6106532",
            "memory: the app process VmHWM and the system MemAvailable polled on the phone every 0.2 s (pong_guard.sh; host memory only, the OpenCL heap is outside VmHWM); peak_mem_mb = the median over the runs of PEAK_VmHWM_KB x 1024 / 10^6, metrics.min_memavailable_mb = the lowest MIN_MEMAVAILABLE_KB of the runs",
            "no reboot, no framework restart in any run: r2b_cpu_c /proc/uptime 556041.1 -> 556073.53, system_server pid 4856 -> 4856",
            "conditions: r2b_cpu_c: before the launch 'MemAvailable 6607192 kB, SKIN \tTemperature{mValue=37.9, mType=3, mName=SKIN, mStatus=0}'; after 'thermal after: Thermal Status: 0'. The app ran in cpuset /top-app. thermal status in the run lines is PowerManager.currentThermalStatus read by the app before and after the five questions, that is after the model load",
            "answers: the letters are compared with the same five turns sent through the LiteRT-LM 0.16.1 Python Conversation API on a Mac (CPU, a new Engine per chart; conv5_cpu.json), with a per-question readout of the reference implementation in the model repo (fp16-int8vocab file, Mac CPU, each question alone; ref.json), and with the data the chart was drawn from; the PNG the phone sent is pixel-identical to the chart the Mac used in every run (png_max_abs_diff 0)",
            "files: hfmodels-android samples/ask/results/2026-09-29-s26/<run>.json and <run>.parity.json (branch decider-vision-demo); litertlm-convert/decider2bv_work/demo/round2/results/s26/<run>.guard / .log / .logcat.txt / .delegation.txt / .poll / .parity.log, r2b_window.txt; harness run_s26_ask.sh, pong_guard.sh, parity_ask.py; Mac side demo/round0b/conv5_charts.py, conv5_cpu.json, ref.json"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "answers_equal_chart_data": 5,
            "answers_equal_mac_fresh_engine": 5,
            "answers_equal_reference_readout": 5,
            "answers_total": 5,
            "conversation_create_s_median": 2.79,
            "first_answer_ms_max": 2113.2,
            "first_answer_ms_median": 2113.2,
            "first_answer_ms_min": 2113.2,
            "followup_answer_ms_max": 1015.1,
            "followup_answer_ms_median": 684.55,
            "followup_answer_ms_min": 561.5,
            "load_s_max": 10.497,
            "load_s_median": 10.497,
            "load_s_min": 10.497,
            "max_output_tokens": 1,
            "min_memavailable_mb": 6114.82,
            "n_runs": 1,
            "peak_vmhwm_mb_max": 3832.38,
            "peak_vmhwm_mb_min": 3832.38,
            "thermal_status_max": 0,
            "turns_per_conversation": 5
          },
          "output_match": null,
          "peak_mem_mb": 3832.38,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "Android app (hfmodels-android samples/ask, release build); one conversation, image + 5 questions",
          "total_ops": 25764,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.1/2026-09-29/decider-2b-vision-int8__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": 24603,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-decider2bv-gate",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 24603 out of 25764 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2245 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 2125 out of 2183 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 50 partitions for subgraph 6 (decode).",
            "VERBOSE: Replacing 1804 out of 1804 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0 (main). [vision encoder, EncoderBackend: CPU]",
            "Time to first token: 1.22 s",
            "Prefill Turn 1: Processed 150 tokens in 1.097457344s duration. Prefill Speed: 136.68 tokens/sec.",
            "Init Total: 17732.80 ms; Init Conversation: 1993.70 ms",
            "Decode Turn 1: Processed 3 tokens in 370.012343ms duration. Decode Speed: 8.11 tokens/sec. (3 output tokens under --max_output_tokens=3: informational, not recorded as decode_tokens_per_s)",
            "stdout 'C) stay': first generated token = upstream fp32 vocabulary top-1 'C' (id 34); prefill 150 = the upstream input length; 0 'Validation error' lines",
            "probe r5_L1_dyn8_cpu_game_pong_atari_level.probe.txt: PEAK_VmHWM_KB=4433672 (process VmHWM polled every 0.2 s, host memory only) -> peak_mem_mb = 4433672 x 1024 / 10^6; MEMAVAILABLE_KB_BEFORE=6866560, MIN_MEMAVAILABLE_KB=5509628",
            "no reboot, no framework restart: /proc/uptime 519312.67 -> 519335.01, system_server pid 4856 -> 4856",
            "gate, 6/6 published rows (one fresh process each; first token = upstream vocabulary top-1 and prefill = upstream token count on every row, 0 'Validation error'): game_pong_atari_level 'C) stay' prefill 150/150; game_breakout_atari_center 'B' prefill 146/146; game_breakout_atari_noball 'D) launch' prefill 146/146; v2_color_purple 'A) A' prefill 105/105; synth_teal 'B' prefill 102/102; synth_center_circle 'B' prefill 110/110. Rows 2-6 ran on the runtime cache the first row built, with cpufreq policies below cpuinfo_max (their speeds are not recorded); their peak VmHWM 4707900-4734996 kB",
            "conditions before the run: thermal status 0, SKIN 37.9 C, every cpufreq policy at cpuinfo_max, runtime process in cpuset '/', rested 180.7 s; cold (bundle freshly pushed and touched, no runtime cache)",
            "device fingerprint samsung/m1qjpnx/m1q:16/BP4A.251205.006/S942QOPS1AZF2_SJP1AZF2:user/release-keys",
            "files: litertlm-convert/decider2bv_work/logs/s26_r5/r5_L1_dyn8_cpu_game_pong_atari_level.err / .out / .probe.txt / .poll; litertlm-convert/decider2bv_work/results/s26_r5_rows/r5_L1_dyn8_cpu_game_pong_atari_level.json"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 3,
            "gate_rows_passed": 6,
            "gate_rows_run": 6,
            "init_conversation_s": 1.9937,
            "init_s": 17.7328,
            "max_num_tokens": 4096,
            "prefill_tokens": 150,
            "threads": 4
          },
          "output_match": null,
          "peak_mem_mb": 4540.08,
          "prefill_tokens_per_s": 136.68,
          "provenance": "measured",
          "runs": true,
          "total_ops": 25764,
          "ttft_ms": 1220.0
        },
        "source": "data/device_runs/0.16.0/2026-09-29/decider-2b-vision-int8__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": 25764,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-decider2bv-ask-app",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.1",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "harness: hfmodels-android samples/ask (github.com/john-rocky/hfmodels-android, branch decider-vision-demo at 419a52b, the app code is commit 60b397f; release APK sha256 dd9a4a40337e66218436d3022f0932e4489bec72da372d7eb981310ed4af9a5a), hfmodels SDK 0.1.3-SNAPSHOT over com.google.ai.edge.litertlm:litertlm-android 0.16.1. One model load per picture and ONE conversation per load: turn 1 = the 256x256 PNG + question 1, turns 2-5 = text only; GenerationOptions(maxOutputTokens = 1). answer_ms = wall from stream() to the first non-empty text piece, conversation_ms = wall of createConversation(), load_ms = wall of fromPretrained(); all measured by the app",
            "09-29 12:34:38.476 11836 11876 I hfmodels: imported decider-2b-vision_int8.litertlm: 3171081088 bytes, sha256 ok, from /storage/emulated/0/Android/data/io.github.johnrocky.hfmodels.samples.ask/files/decider-2b-vision_int8.litertlm [r2b_gpu_a, the first run after the push: the SDK hashed and imported the file; the later runs load the imported file]",
            "09-29 12:43:11.768 22042 22086 I hfmodels: resolve litert-community/decider-2b-vision-LiteRT: commit=6c024e94 via EXPLICIT_REVISION, descriptor=caller@explicit/hfmodels.json sha=a064e06f",
            "09-29 12:43:51.239 22042 22042 I hfmodels: ready litert-community/decider-2b-vision-LiteRT@6c024e94 variant=int8 profile=gpu language=GPU vision=GPU prepare_ms=39469",
            "09-29 12:43:16.687 22042 22086 I tflite  : Replacing 25764 out of 25764 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "09-29 12:43:29.149 22042 22086 I tflite  : Replacing 18852 out of 18852 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_256).",
            "09-29 12:43:33.492 22042 22086 I tflite  : Replacing 16836 out of 16836 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_64).",
            "09-29 12:43:37.325 22042 22086 I tflite  : Replacing 16926 out of 16926 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_16).",
            "09-29 12:43:41.755 22042 22086 I tflite  : Replacing 16926 out of 16926 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 4 (prefill_4).",
            "09-29 12:43:46.183 22042 22086 I tflite  : Replacing 2129 out of 2129 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 5 (prefill_1).",
            "09-29 12:43:46.338 22042 22086 I tflite  : Replacing 2183 out of 2183 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 6 (decode).",
            "09-29 12:43:51.234 22042 22086 I tflite  : Replacing 1 out of 4 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2 partitions for subgraph 0 (main). [4-node graph named main]",
            "09-29 12:44:01.331 22042 23656 I tflite  : Replacing 1804 out of 1804 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main). [vision encoder]",
            "09-29 12:44:06.329 22042 23656 I tflite  : Replacing 22 out of 28 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main). [vision adapter]",
            "09-29 12:43:51.243 22042 22042 I ask     : READY profile=gpu load_ms=39477 variant=int8 commit=6c024e94 downloaded=false",
            "09-29 12:44:08.091 22042 22042 I ask     : TURN 1 id=tallest letter=C ok=true answer_ms=903.2",
            "09-29 12:44:10.343 22042 22042 I ask     : TURN 2 id=shortest letter=B ok=true answer_ms=247.7",
            "09-29 12:44:12.599 22042 22042 I ask     : TURN 3 id=count letter=C ok=true answer_ms=251.2",
            "09-29 12:44:14.940 22042 22042 I ask     : TURN 4 id=taller letter=B ok=true answer_ms=336.2",
            "09-29 12:44:17.186 22042 22042 I ask     : TURN 5 id=leftmost letter=A ok=true answer_ms=241.1",
            "r2b_gpu_a (ask-result-1790652938.json, chart_00, 3 bars, profile gpu): load_ms 46589, conversation_ms 5321.2, answer_ms 990.5 / 289.9 / 288.9 / 409.1 / 268.6, letters BCBAA, equal to the Mac fresh-engine conversation 5/5, to the reference readout 5/5, to the chart's data 5/5; thermal status 2 -> 2; POLLS=294 PEAK_VmHWM_KB=3472204 MIN_MEMAVAILABLE_KB=887208; MEMAVAILABLE_KB_BEFORE=5548292",
            "r3_take1 (ask-result-1790653191.json, chart_02, 5 bars, profile gpu): load_ms 38706, conversation_ms 5111.7, answer_ms 935.4 / 256.3 / 256.5 / 331 / 245.8, letters BDDAA, equal to the Mac fresh-engine conversation 4/5, to the reference readout 4/5, to the chart's data 4/5; thermal status 2 -> 2; POLLS=329 PEAK_VmHWM_KB=4094560 MIN_MEMAVAILABLE_KB=835632; MEMAVAILABLE_KB_BEFORE=6020320",
            "r3_take2 (ask-result-1790653457.json, chart_01, 4 bars, profile gpu): load_ms 39477, conversation_ms 5110.8, answer_ms 903.2 / 247.7 / 251.2 / 336.2 / 241.1, letters CBCBA, equal to the Mac fresh-engine conversation 5/5, to the reference readout 5/5, to the chart's data 5/5; thermal status 2 -> 2; POLLS=331 PEAK_VmHWM_KB=3347124 MIN_MEMAVAILABLE_KB=1338432; MEMAVAILABLE_KB_BEFORE=5992928",
            "r2b_gpu_c (ask-result-1790653942.json, chart_02, 5 bars, profile gpu): load_ms 38831, conversation_ms 4826.4, answer_ms 1143.5 / 317.4 / 305 / 434.8 / 291.5, letters BDDAA, equal to the Mac fresh-engine conversation 4/5, to the reference readout 4/5, to the chart's data 4/5; thermal status 1 -> 1; POLLS=259 PEAK_VmHWM_KB=4271352 MIN_MEMAVAILABLE_KB=928404; MEMAVAILABLE_KB_BEFORE=6335388",
            "memory: the app process VmHWM and the system MemAvailable polled on the phone every 0.2 s (pong_guard.sh; host memory only, the OpenCL heap is outside VmHWM); peak_mem_mb = the median over the runs of PEAK_VmHWM_KB x 1024 / 10^6, metrics.min_memavailable_mb = the lowest MIN_MEMAVAILABLE_KB of the runs",
            "no reboot, no framework restart in any run: r2b_gpu_a /proc/uptime 555258.57 -> 555326.97, system_server pid 4856 -> 4856; r3_take1 /proc/uptime 555513.11 -> 555589.85, system_server pid 4856 -> 4856; r3_take2 /proc/uptime 555777.88 -> 555855.51, system_server pid 4856 -> 4856; r2b_gpu_c /proc/uptime 556270.67 -> 556331.0, system_server pid 4856 -> 4856",
            "conditions: r2b_gpu_a: before the launch 'MemAvailable 6289332 kB, SKIN \tTemperature{mValue=37.9, mType=3, mName=SKIN, mStatus=0}'; after 'thermal after: Thermal Status: 2'; r3_take1: before the launch 'MemAvailable 6874980 kB, SKIN \tTemperature{mValue=37.9, mType=3, mName=SKIN, mStatus=0}'; after 'thermal after: Thermal Status: 2' (airplane mode on, screen recording on); r3_take2: before the launch 'MemAvailable 6673060 kB, SKIN \tTemperature{mValue=37.9, mType=3, mName=SKIN, mStatus=0}'; after 'thermal after: Thermal Status: 2' (airplane mode on, screen recording on); r2b_gpu_c: before the launch 'MemAvailable 6934320 kB, SKIN \tTemperature{mValue=37.9, mType=3, mName=SKIN, mStatus=0}'; after 'thermal after: Thermal Status: 1'. The app ran in cpuset /top-app. thermal status in the run lines is PowerManager.currentThermalStatus read by the app before and after the five questions, that is after the model load",
            "answers: the letters are compared with the same five turns sent through the LiteRT-LM 0.16.1 Python Conversation API on a Mac (CPU, a new Engine per chart; conv5_cpu.json), with a per-question readout of the reference implementation in the model repo (fp16-int8vocab file, Mac CPU, each question alone; ref.json), and with the data the chart was drawn from; the PNG the phone sent is pixel-identical to the chart the Mac used in every run (png_max_abs_diff 0)",
            "chart_02, 'Which bar is the tallest?': both GPU runs answered B (the purple bar, 152) where the CPU run on the same phone, the Mac conversation and the reference readout (p_top 0.9959) answer E (the blue bar, 180, the tallest in the data); the other 18 GPU answers equal the Mac run. Cause not established",
            "files: hfmodels-android samples/ask/results/2026-09-29-s26/<run>.json and <run>.parity.json (branch decider-vision-demo); litertlm-convert/decider2bv_work/demo/round2/results/s26/<run>.guard / .log / .logcat.txt / .delegation.txt / .poll / .parity.log, r2b_window.txt; harness run_s26_ask.sh, pong_guard.sh, parity_ask.py; Mac side demo/round0b/conv5_charts.py, conv5_cpu.json, ref.json"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "answers_equal_chart_data": 18,
            "answers_equal_mac_fresh_engine": 18,
            "answers_equal_reference_readout": 18,
            "answers_total": 20,
            "conversation_create_s_median": 5.111,
            "first_answer_ms_max": 1143.5,
            "first_answer_ms_median": 962.95,
            "first_answer_ms_min": 903.2,
            "followup_answer_ms_max": 434.8,
            "followup_answer_ms_median": 289.4,
            "followup_answer_ms_min": 241.1,
            "load_s_max": 46.589,
            "load_s_median": 39.154,
            "load_s_min": 38.706,
            "max_output_tokens": 1,
            "min_memavailable_mb": 855.69,
            "n_runs": 4,
            "peak_vmhwm_mb_max": 4373.86,
            "peak_vmhwm_mb_min": 3427.45,
            "thermal_status_max": 2,
            "turns_per_conversation": 5
          },
          "output_match": null,
          "peak_mem_mb": 3874.18,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "Android app (hfmodels-android samples/ask, release build); one conversation, image + 5 questions",
          "total_ops": 25764,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.1/2026-09-29/decider-2b-vision-int8__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": 25764,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-decider2bv-gate",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 25764 out of 25764 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 18852 out of 18852 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_256).",
            "VERBOSE: Replacing 16836 out of 16836 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_64).",
            "VERBOSE: Replacing 16926 out of 16926 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_16).",
            "VERBOSE: Replacing 16926 out of 16926 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 4 (prefill_4).",
            "VERBOSE: Replacing 2129 out of 2129 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 5 (prefill_1).",
            "VERBOSE: Replacing 2183 out of 2183 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 6 (decode).",
            "VERBOSE: Replacing 1804 out of 1804 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main). [vision encoder, EncoderBackend: GPU]",
            "VERBOSE: Replacing 22 out of 28 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main). [vision adapter, AdapterBackend: CPU]",
            "VERBOSE: Replacing 1 out of 4 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2 partitions for subgraph 0 (main). [4-node graph named main, CPU]",
            "Time to first token: 0.56 s",
            "Prefill Turn 1: Processed 150 tokens in 440.188854ms duration. Prefill Speed: 340.76 tokens/sec.",
            "Init Total: 64366.57 ms; Init Conversation: 6953.45 ms",
            "Decode Turn 1: Processed 3 tokens in 350.185312ms duration. Decode Speed: 8.57 tokens/sec. (3 output tokens under --max_output_tokens=3: informational, not recorded as decode_tokens_per_s)",
            "stdout 'INFO: Loaded OpenCL library with dlopen. | C)': first generated token = upstream fp32 vocabulary top-1 'C' (id 34); prefill 150 = the upstream input length; 0 'Validation error' lines",
            "probe r5_L3_dyn8_gpu_game_pong_atari_level.probe.txt: PEAK_VmHWM_KB=4220064 (process VmHWM polled every 0.2 s, host memory only) -> peak_mem_mb = 4220064 x 1024 / 10^6; MEMAVAILABLE_KB_BEFORE=7218776, MIN_MEMAVAILABLE_KB=751960 (the GPU heap is outside VmHWM)",
            "no reboot, no framework restart: /proc/uptime 519874.27 -> 519951.46, system_server pid 4856 -> 4856",
            "gate, 6/6 published rows (one fresh process each; first token = upstream vocabulary top-1 and prefill = upstream token count on every row, 0 'Validation error'): game_pong_atari_level 'C)' prefill 150/150; game_breakout_atari_center 'B' prefill 146/146; game_breakout_atari_noball 'D' prefill 146/146; v2_color_purple 'A) red' prefill 105/105; synth_teal 'B) blue' prefill 102/102; synth_center_circle 'B' prefill 110/110. Rows 2-6 ran on the runtime cache the first row built, with cpufreq policies below cpuinfo_max (their speeds are not recorded); their peak VmHWM 3749592-4258076 kB",
            "conditions before the run: thermal status 0, SKIN 37.9 C, every cpufreq policy at cpuinfo_max, runtime process in cpuset '/', rested 180.6 s; cold (bundle freshly pushed and touched, no runtime cache)",
            "post-run logcat read of the hold window: 8 system / app process starts at 02:45:50-51, during this row's conversation init (about 8 s before its prefill); the L1 speed row window had none",
            "device fingerprint samsung/m1qjpnx/m1q:16/BP4A.251205.006/S942QOPS1AZF2_SJP1AZF2:user/release-keys",
            "files: litertlm-convert/decider2bv_work/logs/s26_r5/r5_L3_dyn8_gpu_game_pong_atari_level.err / .out / .probe.txt / .poll; litertlm-convert/decider2bv_work/results/s26_r5_rows/r5_L3_dyn8_gpu_game_pong_atari_level.json"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 3,
            "gate_rows_passed": 6,
            "gate_rows_run": 6,
            "init_conversation_s": 6.95345,
            "init_s": 64.36657,
            "max_num_tokens": 4096,
            "min_memavailable_mb": 770.01,
            "prefill_tokens": 150
          },
          "output_match": null,
          "peak_mem_mb": 4321.35,
          "prefill_tokens_per_s": 340.76,
          "provenance": "measured",
          "runs": true,
          "total_ops": 25764,
          "ttft_ms": 560.0
        },
        "source": "data/device_runs/0.16.0/2026-09-29/decider-2b-vision-int8__galaxy-s26.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "0.17.1",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "litertlm-convert/decider2bv_work/results/runtime_r4.json variants.dyn8.legs.cpu.summary: {\"engine_created\": 12, \"engine_errors\": [], \"n_attempted\": 12, \"n_done\": 12, \"n_first_equals_oracle_top1\": 12, \"n_first_equals_variant_graph_top1\": 11, \"n_pass_i\": 12, \"n_rows\": 12, \"partial_delegation_lines\": 0, \"shape_mismatch_lines\": 0, \"unsupported_op_lines\": 0, \"validation_error_lines\": 0, \"webgpu_delegate_init_lines\": 0}",
            "rows (first streamed token / oracle top-1, prefill / upstream count): game_pong_atari_up 'A'/'A' 150/150; game_pong_atari_level 'C'/'C' 150/150; game_breakout_atari_noball 'D'/'D' 146/146; game_breakout_atari_center 'B'/'B' 146/146; color_red 'A'/'A' 105/105; v2_color_purple 'A'/'A' 105/105; v2_color_split_red_blue 'A'/'A' 105/105; synth_unanswerable 'B'/'B' 116/116; synth_teal 'B'/'B' 102/102; synth_center_circle 'B'/'B' 110/110; synth_count_dots 'B'/'B' 119/119; synth_half_half 'A'/'A' 105/105",
            "11/12 rows equal this file's own CPU graph top-1 (litertlm-convert/decider2bv_work/results/readout_r4_dyn8.json)",
            "files: litertlm-convert/decider2bv_work/results/runtime_r4_rows/dyn8/cpu/*.json, litertlm-convert/decider2bv_work/logs/runtime_r4/dyn8/cpu_*.log"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "gate_rows_passed": 12,
            "gate_rows_run": 12,
            "max_num_tokens": 4096,
            "max_output_tokens": 3,
            "threads": 8
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.17.1/2026-09-29/decider-2b-vision-int8__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-29",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "0.17.1",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "litertlm-convert/decider2bv_work/results/runtime_r4.json variants.dyn8.legs.gpu.summary: {\"engine_created\": 12, \"engine_errors\": [], \"n_attempted\": 12, \"n_done\": 12, \"n_first_equals_oracle_top1\": 12, \"n_first_equals_variant_graph_top1\": 11, \"n_pass_i\": 12, \"n_rows\": 12, \"partial_delegation_lines\": 0, \"shape_mismatch_lines\": 0, \"unsupported_op_lines\": 0, \"validation_error_lines\": 0, \"webgpu_delegate_init_lines\": 96}",
            "rows (first streamed token / oracle top-1, prefill / upstream count): game_pong_atari_up 'A'/'A' 150/150; game_pong_atari_level 'C'/'C' 150/150; game_breakout_atari_noball 'D'/'D' 146/146; game_breakout_atari_center 'B'/'B' 146/146; color_red 'A'/'A' 105/105; v2_color_purple 'A'/'A' 105/105; v2_color_split_red_blue 'A'/'A' 105/105; synth_unanswerable 'B'/'B' 116/116; synth_teal 'B'/'B' 102/102; synth_center_circle 'B'/'B' 110/110; synth_count_dots 'B'/'B' 119/119; synth_half_half 'A'/'A' 105/105",
            "I0000 00:00:1790611633.785552 83523725 environment.cc:526] Selected adapter: Apple M4 Max, arch=metal-3, vendor=apple, backend=Metal, adapterType=Integrated GPU (litertlm-convert/decider2bv_work/logs/runtime_r4/dyn8/gpu_game_pong_atari_up.log:19)",
            "litertlm-convert/decider2bv_work/logs/runtime_r4/dyn8/gpu_game_pong_atari_up.log: 8 'Initializing WebGPU-based API' delegate kernels (7 decoder signatures + the vision encoder), no 'not supported by GPU delegate' block",
            "11/12 rows equal this file's own CPU graph top-1 (litertlm-convert/decider2bv_work/results/readout_r4_dyn8.json)",
            "files: litertlm-convert/decider2bv_work/results/runtime_r4_rows/dyn8/gpu/*.json, litertlm-convert/decider2bv_work/logs/runtime_r4/dyn8/gpu_*.log"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "gate_rows_passed": 12,
            "gate_rows_run": 12,
            "max_num_tokens": 4096,
            "max_output_tokens": 3
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.17.1/2026-09-29/decider-2b-vision-int8__mac-studio-m4-max.json"
      }
    ]
  },
  "model": {
    "family": "decider",
    "id": "decider-2b-vision-int8",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/decider-2b-vision-LiteRT",
    "task": "image-text-to-text"
  },
  "pitfalls": [
    "Use the GPU when the probabilities matter: on the CPU this file quantizes activations on the fly and its probabilities move against upstream fp32 on the published image questions (51/53 answers equal, max |dp| 0.429, p95 0.164), and they also depend on how the prompt is split into prefill chunks (max |dp| 0.233 with one padded chunk); on the Mac GPU (Metal CompiledModel, one padded chunk per answer slot) it computes the int8 weights in float: 52/53, max |dp| 0.032, p95 0.011. Probabilities on the phone were not measured, because the runtime returns text only (Hub card Files note; FINDINGS section 5).",
    "Galaxy S26 (SM-S942Q, MemTotal 11,389,756 kB, LiteRT-LM v0.16.0 litert_lm_advanced_main, a different runtime version from the Mac rows): the first token was upstream's answer letter on 6/6 published rows on the CPU and on the GPU, with prefill counts equal to upstream's and no 'Validation error' lines. Speed row (game_pong_atari_level, 150 prompt tokens, greedy, cold, n = 1): CPU 4 threads TTFT 1.22 s, prefill 136.68 tok/s, engine creation 17.7 s (+2.0 s conversation set-up), peak VmHWM 4.54 GB; GPU (OpenCL) 0.56 s, 340.76 tok/s, 64.4 s (+7.0 s), 4.32 GB (Hub card Performance).",
    "On the S26 GPU all 7 decoder signatures were fully delegated to OpenCL (one partition each); the vision encoder ran on the GPU and the adapter on the CPU. VmHWM does not include GPU memory: during the int8 GPU run the phone's MemAvailable fell to 0.77 GB, so the int8 GPU run is close to the limit of a phone of this RAM class; the other five int8 CPU rows peaked at 4.82-4.85 GB VmHWM (Hub card Performance).",
    "Mac (LiteRT-LM 0.17.1): the int8 file was not timed; its first token was upstream's letter on 12/12 single-question image rows on the CPU and on the GPU (11/12 equal to this file's own CPU graph), and synth_center_circle gave the reference answer B on both runtime legs while the int8 CPU graph gives A under both prefill feedings; the runtime's own chunk plan is not observed (Hub card Performance; FINDINGS sections 5 and 8).",
    "The model returns option probabilities, not text: upstream reads the letter logits A-J at one answer slot per question (a ' (' token after ':' with 'Answer' within the 5 tokens before it), cuts them to the option count and softmaxes at T = 1. Through LiteRT-LM only the answer letter is available: the greedy first token of one question per request, one image first, and it is the top token over the whole vocabulary, not only over the option letters (on one two-question published row the top token for a two-option question is 'C'). LiteRT-LM's text-scoring call cannot stand in for the probabilities (its scores disagree with its own greedy decode on 0.17.1, LiteRT-LM #3561). Probabilities, text-only requests and several questions per request go through the bundled reference readout, reference/decider_litert.py (ai-edge-litert 2.2.0, numpy, Pillow, tokenizers; no torch, transformers or litert-lm) (Hub card Use it, Limitations).",
    "Fixed 256x256 input (64 image tokens) where upstream picks a resolution per image: for 224x224, 256x240 and 256x256 originals upstream's own processing gives the same pixels (13 published requests, 21 slots, |dp| 0); every other size gets fewer image tokens than upstream would use (70 for the 160x210 frames, 144 to 475 for the larger synthetic images), and upstream fp32 itself then matches its original-image answer on 30/32 published slots, max |dp| 0.373, p95 0.318, up to 0.58 on photographs and documents kept out of the repo. Resize to 256x256 with PIL bicubic before sending; the runtime's own image resize was never exercised (Hub card The price of the fixed 256x256 input, Limitations; FINDINGS section 2).",
    "Positions: the checkpoint uses 3-channel M-RoPE (image tokens on an 8x8 grid) and LiteRT-LM passes one 1-D position, so the decoder derives the three channels inside its graph with element-wise ops; plain 1-D positions flipped 2 of the 53 published image answers and moved one probability by 0.198 in upstream fp32. Text-only requests must start at position 65, where the derived positions keep upstream's relative positions; the runtime numbers them from 0 (the 9 text-only questions moved by up to 0.0497 with one flip in the fp32 graph), so text-only requests go through the reference readout (Hub card What matches upstream 4-5; FINDINGS sections 1 and 7).",
    "Rebuilders: do not export the derived rotary's step as clamp(x, 0, 1). The converter lowers it to RELU_0_TO_1 (9 per signature), which the WebGPU delegate of LiteRT-LM 0.17.1 on the Mac refuses ('RELU_0_TO_1: Not supported op', 25,675 of 25,927 prefill_1024 ops on the GPU, 'Hint fully delegated to single delegate is set, but the graph is not fully delegated', engine creation fails, twice). The shipped decoders write it as relu(x) - relu(x - 1) (equal on integer inputs; fp32 letter logits bit-equal) and create the full-GPU engine on the Mac and, for this file, on the S26 OpenCL delegate (Hub card Performance; FINDINGS section 4).",
    "XNNPACK cache and GPU caches: LiteRT-LM's CPU cache folder holds 1,896,902,288 B for this decoder plus 1,312,687,576 B for the vision encoder and adapter; with the GPU backend the decoder program cache in a cache folder grew by 270,663,680 B at every engine creation in the lane's Mac runs, and a warm folder did not shorten GPU engine creation (Hub card Files; FINDINGS section 6 and Corrections).",
    "Bundle metadata shared by all three files: identity template (no role markers), no start token, stop token 248044, one 256x256 image, a 4096-token cache (rows longer than 4096 tokens do not fit; prompts above 2048 tokens were not exercised, and the derived rotary computes positions in float), ExecutorMetadata for the 48 state buffers and an fp32 activation preference for the decoder (no runtime log line states the precision the GPU executor used). One image per request, English only; the text behaviour is upstream's (decider-2b v5 text weights) (Hub card Files, Limitations; FINDINGS section 8 and Open questions).",
    "License Apache-2.0 as declared by the upstream checkpoint (source revision 863e290863655f1d6b69324d77d09ac972d21609); changes: text decoder, vision encoder and merger converted to LiteRT flatbuffers, M-RoPE derived in the decoder graph from a 1-D position, image fixed at 256x256, weights cast as listed, chat template replaced with an identity template, tokenizer repackaged unchanged, ExecutorMetadata and an fp32 activation preference added; the multi-token-prediction head is not included; community conversion, not an official Mapika or Qwen release (Hub card License and changes)."
  ],
  "schema_version": "1.2"
}
