{
  "artifacts": [
    {
      "file": "gliner25_decide_s512_wfp16.tflite",
      "sha256": "bbe150fa6f7bf29cf4b12a6eff630b3a8fb59f41b41ea14a68b1b693fea6fc71",
      "size_mb": 773.791
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-26/gliner2.5-decide__gliner25_decide_s512_wfp16.json"
  },
  "conversion": {
    "command": "cd conversion && python -B export_shape.py --seq 512 && python -B quantize_wfp16.py --seq 512  (conversion/README.md 'Order': the fp32 export first, then the weight-only float16 cast of that gated export; exports/gliner25_decide_s512_wfp16.tflite)",
    "quantization": "float16 weight storage (ai-edge-quantizer 0.8.0 FLOAT_CASTING, weight-only 16-bit): the 146 FULLY_CONNECTED weight tensors stored as float16 with a DEQUANTIZE to float32; activations and every other constant stay float32; 1,780 operators vs 1,634 in the fp32 reference graph; no INT8 file is shipped (card 'Files', 'Limits'; conversion/README.md)",
    "tool": "litert-torch",
    "tool_version": "0.9.3 (ai-edge-litert 2.1.6; torch 2.12.1; transformers 4.57.6; gliner2 2.0.0; ai-edge-quantizer 0.8.0 for the float16 weight storage)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "SQUARED_DIFFERENCE"
    ],
    "coverage_ops_pct": 89.0,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 1731,
      "unmatched": 49
    },
    "partitions": 196
  },
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 1780,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliner25decide-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliner25-decide-round2-v1, the GLiNER2.5-Small runner rebuilt for three inputs; vendor libLiteRt.so + libLiteRtClGlAccelerator.so sha256-pinned in android/provenance.json), one process per job, fixture inputs pushed as .f32 files (inputs_embeds, attention_mask, label_routing), sha256-checked on the device. Round-3 inputs: rows of the float16 embedding table upcast to float32 on the host = the shipping path (results/device/inputs_s512/inputs.json table.kind 'fp16', word_embeddings_fp16.bin sha256 f90f5d83...). Sources: ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/gpu_gate_r3.json (job g: 'not a gate: S26 CompiledModel CPU (card CPU column)'), raw outputs ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r3_g_s512_wfp16_cpu/NNN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r3_g_s512_wfp16_cpu/runner.stderr.log, supervisor recount ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/SUPERVISOR.md.",
            "logcat/stderr: 'Replacing 1780 out of 1780 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions'; compile PASS in 692.4 ms; is_fully_accelerated=true; hardware_accelerators 'CPU only (bitmask 1)'",
            "conditions (round 3, card CPU column, not a gate): cool start (cool after 121.5 s wait; rule kgsl <= 50 C and thermal_status 0); screen kept awake via 'svc power stayon usb' (restored to false afterwards), USB powered; before: kgsl 41.5 C, max clock 1300 MHz, thermal_pwrlevel 0, thermal_status 0, battery 37.5 C 86 %, wakefulness Awake; after: kgsl 50.8 C, max clock 500 MHz, thermal_pwrlevel 10, thermal_status 2, battery 41.2 C; runner 2026-09-26T13:14:40+09:00 -> 2026-09-26T13:22:27+09:00.",
            "parity: 42/42 decisions equal the official gliner2 2.0.0 fp32 classify_text result (42 requests = 21 model-card examples + 21 fast-decisions dev rows that fit the window); max |dlogit| vs official 3.311e-03 (ticket_route_019), max |dprob| 2.256e-04; NaN/Inf in any repetition: 0.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/table_gate_raw/s512_graph-wfp16_table-fp16), recomputed 2026-09-26 from the saved .f32 outputs: 42/42 fixtures pass over all 32 slots (0 violating elements; valid label slots: 42/42, 0 violations); max |d| 1.192e-05, max rel 2.239e-04.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 2 warm-ups then 5 timed runs per fixture, 42 fixtures = 210 timed runs; median 1665.98 / min 863.07 / max 1788.94 ms; per-fixture medians 879.92 ms (first) -> 1785.20 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 1665.9826295,
          "loads": true,
          "max_abs_diff": 1.1920928955078125e-05,
          "max_rel_diff": 0.0002239291206933558,
          "metrics": {
            "battery_temp_c_end": 41.2,
            "battery_temp_c_start": 37.5,
            "cool_start_wait_s": 121.5,
            "decisions_equal_official": 42,
            "fixtures": 42,
            "gpu_max_clock_mhz_end": 500,
            "iterations": 210,
            "kgsl_temp_c_end": 50.8,
            "kgsl_temp_c_start": 41.5,
            "latency_max_ms": 1788.941093,
            "latency_min_ms": 863.065,
            "load_compile_ms": 692.383593,
            "max_abs_dlogit_vs_official": 0.0033109188079833984,
            "max_abs_dprob_vs_official": 0.00022557377815246582,
            "per_fixture_median_first_ms": 879.916354,
            "per_fixture_median_last_ms": 1785.198124,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 42,
            "shared_rule_violating_elements": 0,
            "thermal_status_end": 2,
            "thermal_status_start": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK (native runner, default threads)",
          "total_ops": 1780,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliner2.5-decide__gliner25_decide_s512_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 1780,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliner25decide-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Android sample app gate (litert-community/GLiNER2.5-Decide-LiteRT android/, package com.gliner25decide, debug build APK sha256 0c5f08b04413..., com.google.ai.edge.litert:litert:2.2.0): one Environment, Accelerator.CPU, CpuOptions 4 threads (XNNPACK); the app builds the inputs itself (SentencePiece tokenizer, gliner2 schema, memory-mapped float16 table) and the on-device token ids, [L] positions and padding equal the captured Python batches on all 126 (request, window) pairs (126/126). Sources: ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/app_gate_r5.json (gate.CPU.windows.s512), ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r5_gate_cpu/pulled/files/gate/gate_CPU.json (per-pair logits), logcat ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r5_gate_cpu/logcat.txt.",
            "logcat (three compiles in one process, one per window): '09-26 14:45:50.856  9723  9775 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0 (main).' | '09-26 14:45:52.430  9723  9775 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0 (main).' | '09-26 14:45:54.399  9723  9775 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 0 (main).'.",
            "conditions: screen behind the secure keyguard, activity shown over it (setShowWhenLocked, debug diagnostics only), verified top-resumed / cpuset top-app; cool start (cool after 192.3 s wait); before: kgsl 41.1 C, max clock 1300 MHz, thermal_status 0, battery 37.2 C 83 %; after the whole gate (all three windows, 2026-09-26T14:45:50+09:00 + 332 s): kgsl 51.9 C, max clock 500 MHz, thermal_pwrlevel 10, thermal_status 2, battery 43.4 C. This window's numbers are one third of that run; the per-window thermal state is not recorded separately.",
            "parity: 42/42 pairs' decisions equal the official gliner2 2.0.0 fp32 result, 42/42 all logits finite; max |dlogit| vs official 3.311e-03; CPU logits bit-identical to the native-runner round-3 job (max |d| 0.0). output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/table_gate_raw/s512_graph-wfp16_table-fp16), recomputed 2026-09-26 from the per-pair logits: 42/42 pairs pass over all 32 slots (0 violating elements); max |d| 1.192e-05, max rel 2.239e-04.",
            "latency: graph phase = First input write through synchronized output readback; 1 warm-up + 3 timed runs per pair, 42 pairs = 126 timed runs; median 1242.72 / min 711.42 / max 1405.62 ms over all timed runs; pair medians: median 1243.16, first three 735.08, 982.06, 1065.04 -> last three 1244.27, 1390.12, 1402.48 ms; tokenize+embed pair-median 22.73 ms, decode 0.051 ms (host phases, outside latency_p50_ms)."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 1242.723984,
          "loads": true,
          "max_abs_diff": 1.1920928955078125e-05,
          "max_rel_diff": 0.0002239291206933558,
          "metrics": {
            "cpu_threads": 4,
            "decisions_equal_official": 42,
            "decode_pair_median_ms": 0.0506515,
            "gpu_max_clock_mhz_end_whole_gate": 500,
            "inputs_identical_to_python_pairs": 126,
            "iterations": 126,
            "kgsl_temp_c_end_whole_gate": 51.9,
            "kgsl_temp_c_start": 41.1,
            "latency_max_ms": 1405.6192179999998,
            "latency_min_ms": 711.421094,
            "max_abs_dlogit_vs_native_runner": 0.0,
            "max_abs_dlogit_vs_official": 0.0033109188079833984,
            "pair_median_first_ms": 735.0803129999999,
            "pair_median_last_ms": 1402.4848440000003,
            "pair_median_ms": 1243.1611715,
            "pairs": 42,
            "shared_rule_pass_fixtures": 42,
            "shared_rule_violating_elements": 0,
            "thermal_status_end_whole_gate": 2,
            "thermal_status_start": 0,
            "timed_runs_per_pair": 3,
            "tokenize_embed_pair_median_ms": 22.7265885,
            "warmup_runs_per_pair": 1
          },
          "output_match": true,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK 4 threads; Android sample app (debug build)",
          "total_ops": 1780,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliner2.5-decide__gliner25_decide_s512_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 1780,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliner25decide-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliner25-decide-round2-v1, the GLiNER2.5-Small runner rebuilt for three inputs; vendor libLiteRt.so + libLiteRtClGlAccelerator.so sha256-pinned in android/provenance.json), one process per job, fixture inputs pushed as .f32 files (inputs_embeds, attention_mask, label_routing), sha256-checked on the device. Round-3 inputs: rows of the float16 embedding table upcast to float32 on the host = the shipping path (results/device/inputs_s512/inputs.json table.kind 'fp16', word_embeddings_fp16.bin sha256 f90f5d83...). Sources: ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/gpu_gate_r3.json (job c: 'G3 S26 GPU precision FP32, wfp16 storage, fp16 table inputs'), raw outputs ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r3_c_s512_wfp16_fp32/NNN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r3_c_s512_wfp16_fp32/runner.stderr.log, supervisor recount ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/SUPERVISOR.md.",
            "logcat/stderr: 'Replacing 1780 out of 1780 node(s) with delegate (LITERT_CL) node, yielding 1 partitions'; compile PASS in 2766.0 ms; is_fully_accelerated=true; hardware_accelerators 'GPU only (bitmask 2)'; gpu_options_toml 'precision = 2'",
            "conditions (round 3, gate G3): cool start (cool after 151.8 s wait; rule kgsl <= 50 C and thermal_status 0); screen kept awake via 'svc power stayon usb' (restored to false afterwards), USB powered; before: kgsl 41.8 C, max clock 1300 MHz, thermal_pwrlevel 0, thermal_status 0, battery 37.6 C 92 %, wakefulness Awake; after: kgsl 58.5 C, max clock 646 MHz, thermal_pwrlevel 8, thermal_status 2, battery 44.7 C; runner 2026-09-26T12:48:20+09:00 -> 2026-09-26T12:52:00+09:00.",
            "parity: 42/42 decisions equal the official gliner2 2.0.0 fp32 classify_text result (42 requests = 21 model-card examples + 21 fast-decisions dev rows that fit the window); max |dlogit| vs official 3.301e-03 (ticket_route_019), max |dprob| 2.258e-04; NaN/Inf in any repetition: 0.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/table_gate_raw/s512_graph-wfp16_table-fp16), recomputed 2026-09-26 from the saved .f32 outputs: 42/42 fixtures pass over all 32 slots (0 violating elements; valid label slots: 42/42, 0 violations); max |d| 2.575e-05, max rel 7.895e-05.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 2 warm-ups then 5 timed runs per fixture, 42 fixtures = 210 timed runs; median 733.12 / min 573.49 / max 936.55 ms; per-fixture medians 575.23 ms (first) -> 912.12 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 733.1154160000001,
          "loads": true,
          "max_abs_diff": 2.574920654296875e-05,
          "max_rel_diff": 7.895220187492669e-05,
          "metrics": {
            "battery_temp_c_end": 44.7,
            "battery_temp_c_start": 37.6,
            "cool_start_wait_s": 151.8,
            "decisions_equal_official": 42,
            "fixtures": 42,
            "gpu_max_clock_mhz_end": 646,
            "iterations": 210,
            "kgsl_temp_c_end": 58.5,
            "kgsl_temp_c_start": 41.8,
            "latency_max_ms": 936.549166,
            "latency_min_ms": 573.492812,
            "load_compile_ms": 2765.954582,
            "max_abs_dlogit_vs_official": 0.003300905227661133,
            "max_abs_dprob_vs_official": 0.00022584199905395508,
            "per_fixture_median_first_ms": 575.227343,
            "per_fixture_median_last_ms": 912.124947,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 42,
            "shared_rule_violating_elements": 0,
            "thermal_status_end": 2,
            "thermal_status_start": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions precision=FP32",
          "total_ops": 1780,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliner2.5-decide__gliner25_decide_s512_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 1780,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliner25decide-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Android sample app gate (litert-community/GLiNER2.5-Decide-LiteRT android/, package com.gliner25decide, debug build APK sha256 0c5f08b04413..., com.google.ai.edge.litert:litert:2.2.0): one Environment, Accelerator.GPU with GpuOptions(precision = FP32) (explicit); the app builds the inputs itself (SentencePiece tokenizer, gliner2 schema, memory-mapped float16 table) and the on-device token ids, [L] positions and padding equal the captured Python batches on all 126 (request, window) pairs (126/126). Sources: ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/app_gate_r5.json (gate.GPU.windows.s512), ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r5_gate_gpu/pulled/files/gate/gate_GPU.json (per-pair logits), logcat ~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/device/r5_gate_gpu/logcat.txt.",
            "logcat (three compiles in one process, one per window): '09-26 14:38:51.610  6980  7036 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).' | '09-26 14:38:54.126  6980  7036 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).' | '09-26 14:38:58.145  6980  7036 I tflite  : Replacing 1780 out of 1780 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).'.",
            "conditions: screen behind the secure keyguard, activity shown over it (setShowWhenLocked, debug diagnostics only), verified top-resumed / cpuset top-app; cool start (cool after 0.1 s wait); before: kgsl 41.1 C, max clock 1300 MHz, thermal_status 0, battery 32.3 C 85 %; after the whole gate (all three windows, 2026-09-26T14:38:50+09:00 + 212 s): kgsl 49.6 C, max clock 1200 MHz, thermal_pwrlevel 1, thermal_status 2, battery 43.1 C. This window's numbers are one third of that run; the per-window thermal state is not recorded separately.",
            "parity: 42/42 pairs' decisions equal the official gliner2 2.0.0 fp32 result, 42/42 all logits finite; max |dlogit| vs official 3.301e-03; GPU logits bit-identical to the native-runner round-3 job (max |d| 0.0). output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliner25-decide-litert/results/table_gate_raw/s512_graph-wfp16_table-fp16), recomputed 2026-09-26 from the per-pair logits: 42/42 pairs pass over all 32 slots (0 violating elements); max |d| 2.575e-05, max rel 7.895e-05.",
            "latency: graph phase = First input write through synchronized output readback; 1 warm-up + 3 timed runs per pair, 42 pairs = 126 timed runs; median 639.24 / min 576.02 / max 1047.52 ms over all timed runs; pair medians: median 638.77, first three 579.43, 578.33, 582.49 -> last three 672.03, 704.54, 695.11 ms; tokenize+embed pair-median 43.09 ms, decode 0.072 ms (host phases, outside latency_p50_ms)."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 639.23664,
          "loads": true,
          "max_abs_diff": 2.574920654296875e-05,
          "max_rel_diff": 7.895220187492669e-05,
          "metrics": {
            "decisions_equal_official": 42,
            "decode_pair_median_ms": 0.07210949999999999,
            "gpu_max_clock_mhz_end_whole_gate": 1200,
            "inputs_identical_to_python_pairs": 126,
            "iterations": 126,
            "kgsl_temp_c_end_whole_gate": 49.6,
            "kgsl_temp_c_start": 41.1,
            "latency_max_ms": 1047.515729,
            "latency_min_ms": 576.0178639999999,
            "max_abs_dlogit_vs_native_runner": 0.0,
            "max_abs_dlogit_vs_official": 0.003300905227661133,
            "pair_median_first_ms": 579.433229,
            "pair_median_last_ms": 695.1102599999999,
            "pair_median_ms": 638.7681769999999,
            "pairs": 42,
            "shared_rule_pass_fixtures": 42,
            "shared_rule_violating_elements": 0,
            "thermal_status_end_whole_gate": 2,
            "thermal_status_start": 0,
            "timed_runs_per_pair": 3,
            "tokenize_embed_pair_median_ms": 43.090417,
            "warmup_runs_per_pair": 1
          },
          "output_match": true,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions precision=FP32; Android sample app (debug build)",
          "total_ops": 1780,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliner2.5-decide__gliner25_decide_s512_wfp16__galaxy-s26.json"
      }
    ]
  },
  "model": {
    "family": "gliner2.5-decide",
    "id": "gliner2.5-decide__gliner25_decide_s512_wfp16",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/GLiNER2.5-Decide-LiteRT",
    "task": "text-classification"
  },
  "pitfalls": [
    "Weight storage and GPU computation precision are separate settings: at the runtime's default GPU precision the s128 graph (float32 weights) compiled and returned finite logits on the Galaxy S26, but only 14 of 42 decisions matched the official result; with GpuOptions(precision = FP32) all 42 matched. Use the explicit FP32 option for every file (card 'GPU precision').",
    "The graph is the DeBERTa-v3-large encoder plus the classification head only: the host looks up the word embeddings before the graph (host_assets/word_embeddings_fp16.bin, [128011,1024] float16, upcast to float32 on lookup; the graph applies the embedding LayerNorm itself, so feed the raw rows) and turns the logits into decisions after it (softmax or sigmoid per task, thresholds). HOST_CONTRACT.md specifies the encoded sequence, the [L] marker positions and the decision rules (card 'Files'; HOST_CONTRACT.md).",
    "Assign the inputs by shape, not by name: the converter names them args_0..args_2 (inputs_embeds [1,N,1024], attention_mask [1,N] with 1.0 for encoded tokens, label_routing [1,32,N] one-hot at the j-th [L] marker); output_0 [1,1,1,32] holds one logit per label slot in request order, and slots past the label count hold a constant that is ignored (HOST_CONTRACT.md 'Graph signature').",
    "Fixed windows N = 128 / 256 / 512: the task schemas plus the text must fit in N encoded tokens with at most 32 labels in total; longer requests are rejected, never truncated, and gliner2's long-text chunking (classify_text_long) is not ported. Use the smallest window that holds the request (card 'Files', 'Limits').",
    "Back-to-back requests heat the phone: on the Galaxy S26 the GPU clock limit steps down from 1,300 MHz (s256 902 MHz, s512 646 MHz at the end of the job) and each s256 / s512 request gets slower during the job — the opening request is the cool value, the median mixes it with the throttled end (card 'Latency').",
    "GPU validation covers the Galaxy S26 (Adreno) only; NPU execution was not evaluated. No INT8 file is shipped: on GLiNER2.5 Small, the same family of graphs, dynamic-range INT8 FULLY_CONNECTED weights did not compile on the LiteRT 2.2.0 GPU, and no INT8 variant of this model was built or tested (card 'Limits').",
    "English only, as the source model; one text per call, no batching; few-shot examples and gliner2's extraction tasks (entities, relations, structures) are not part of the graphs (card 'Limits')."
  ],
  "schema_version": "1.2"
}
