{
  "artifacts": [
    {
      "file": "gliformer_large_ner_s128_wfp16.tflite",
      "sha256": "5c1f4ce26f953573d10c64f9d1d1559c669b52d649bafbfc462f24dddbcf9375",
      "size_mb": 674.243
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16.json"
  },
  "conversion": {
    "command": "not published as a script (card 'Provenance, conversion and license'): litert-torch 0.9.3 fixed-shape fp32 export of the NER path of knowledgator/gliformer-large-v1 (rev d0a4e53d) re-expressed for the GPU delegate without changing its math — host-side token lookup, one-hot routing as matmul, float masks, attention at rank 4, the exact DeBERTa logarithmic relative-position buckets as projected tables, the word BiLSTM unrolled for the window, one packed output tensor; then ai-edge-quantizer 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights (the wfp16 file); the layout and page embeddings of the backbone are skipped exactly as the upstream text path skips them",
    "quantization": "float16 weight storage (ai-edge-quantizer 0.8.0 FLOAT_CASTING on the FULLY_CONNECTED weights only): 153 float16 weight tensors read through a DEQUANTIZE to float32, every activation and every other constant float32; 3,996 operators in the fp32 reference (gliformer_large_ner_s128_fp32.tflite, 1,361,304,680 B, published as the exact reference) — 4,149 operators in this file; dynamic-range INT8 was not attempted — it does not compile on the LiteRT 2.2.0 GPU delegate for this graph family (card 'Provenance, conversion and license')",
    "tool": "litert-torch",
    "tool_version": "0.9.3 (torch 2.12.1; ai-edge-litert 2.1.6; ai-edge-quantizer 0.8.0 for the float16 weight storage; gliformer 0.1.2, gliner 0.2.29, transformers 5.16.1 — requirements-lock.txt / upstream.json)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "SQUARED_DIFFERENCE",
      "LOGISTIC",
      "TANH"
    ],
    "coverage_ops_pct": 83.6,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 3622,
      "unmatched": 527
    },
    "partitions": 681
  },
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 4146,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round4-six-inputs-memory-default-probe-v2; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/011_hash_gliformer_large_ner_s128_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round4_s128_wfp16_cpu.json (gate s128_wfp16_cpu, 'CPU only'), raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/s128_wfp16_cpu/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/025_run_s128_wfp16_cpu.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/027_own_pid_logcat_s128_wfp16_cpu.stdout.log.",
            "logcat/stderr: 'Created TensorFlow Lite XNNPACK delegate for CPU.' | 'Replacing 4146 out of 4149 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main).'; compile status 0 in 1786.9 ms; hardware_accelerators 'CPU only (bitmask 1)'; native default CPU options (thread count not set by the runner); the three non-XNNPACK nodes are not named by the log.",
            "conditions (round 4, the four s128 jobs ran back to back in one adb session, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before the s128 shape 29.9 C 94 % (2026-09-26T02:03:22.264820+00:00), after this job 33.6 C 94 % (2026-09-26T02:05:15.255268+00:00); runner wall 24.8 s; compile 1786.9 ms (status 0); process peak RSS (VmHWM) 2,532,122,624 B, RSS after the last readback 2,496,192,512 B. This job ran third of the four s128 jobs (after the fp32 GPU job).",
            "parity: 10/10 fixtures' entity span sets (label, start, end) equal the official gliformer 0.1.2 fp32 CPU result (checkpoint rev d0a4e53d, same tokenizer, threshold 0.5); all outputs finite: True; max score diff 2.401e-04 (gate tolerance 0.01); max |dlogit| vs official 7.509e-03.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/litert_cpu_outputs_round2_s128_wfp16.npz; the device pack android/fixtures/s128/*.f32 is byte-identical to the fixtures/graph_inputs_NN.npz rows of the float32 table that run consumed, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 10/10 fixtures pass over all 48 x 15 slots (0 violating elements of 7,200); max |d| 4.387e-05, max rel 1.816e-04.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 5 timed runs after 2 warm-ups per fixture, 10 fixtures = 50 timed runs; median 322.55 / min 311.34 / max 347.11 ms; per-fixture medians 316.08 ms (first) -> 323.48 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": 322.54960900000003,
          "loads": true,
          "max_abs_diff": 4.38690185546875e-05,
          "max_rel_diff": 0.00018157862359657884,
          "metrics": {
            "battery_temp_c_end": 33.6,
            "battery_temp_c_start": 29.9,
            "exact_span_sets_vs_official": 10,
            "fixtures": 10,
            "iterations": 50,
            "latency_max_ms": 347.108906,
            "latency_min_ms": 311.338854,
            "load_compile_ms": 1786.856041,
            "loaded_rss_bytes": 2496192512,
            "max_abs_dlogit_vs_official": 0.0075092315673828125,
            "max_score_diff_vs_official": 0.00024008750915527344,
            "peak_rss_bytes": 2532122624,
            "per_fixture_median_first_ms": 316.077864,
            "per_fixture_median_last_ms": 323.480729,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 10,
            "shared_rule_violating_elements": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": 2532.12,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK (native runner, default threads)",
          "total_ops": 4149,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 4146,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Android sample app gate (litert-community/GLiFormer-Large-NER-LiteRT android/, package com.gliformer, debug build APK sha256 3bf2de72f19f... (installed 2026-09-26 21:29 JST, ~/code/codex-conversions/2026-09-26/gliformer-android/results/round5/install.json), com.google.ai.edge.litert:litert:2.2.0, round 5): one Environment, Accelerator.CPU with CpuOptions(numThreads = 4) (GliformerExtractor.kt), one window (128) resident; the app builds the inputs itself (whitespace word splitter, SentencePiece Unigram tokenizer, memory-mapped float16 table upcast to float32, routing construction) and its token ids / first-subtoken / parent / entity positions / masks equal the captured Python inputs on all 80/80 tokenizer checks (400 raw graph-tensor comparisons). Sources: ~/code/codex-conversions/2026-09-26/gliformer-android/results/gate_s128_cpu_r5.json, raw logits ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s128_cpu_r5/raw_logits.tar (per-row sha256 in the JSON), logcat ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s128_cpu_r5/logcat_own_pid.log.",
            "logcat: '09-26 21:31:44.487 18942 18999 I tflite  : Created TensorFlow Lite XNNPACK delegate for CPU.' | '09-26 21:31:44.488 18942 18999 I tflite  : Replacing 4146 out of 4149 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main).'.",
            "conditions: screen unlocked and awake (mDreamingLockscreen false, mWakefulness Awake), 'svc power stayon usb' on, USB powered; app battery 37 C 80 % before -> 38.2 C after the gate (2026-09-26T12:31:43.393480Z -> 2026-09-26T12:32:51.579070Z); load 1981.5 ms, 12 full-pipeline warm-up passes 1488.6 ms before the timed rows; memory after load VmRSS 2,663,985,152 B / VmHWM 2,713,726,976 B / total PSS 2,585,922,560 B, after the gate VmRSS 2,731,327,488 B / VmHWM 2,756,358,144 B. No GPU clock / thermal_status sample in this gate.",
            "parity: 60/60 inputs' entity span sets equal the official gliformer 0.1.2 fp32 result and 60/60 equal the Python LiteRT 2.1.6 CPU fp16-table reference (all 5 repetitions); all logits finite; max score diff vs the matched-fp16 Python reference 1.907e-06 (gate tolerance 1e-05), vs official 6.462e-04 (tolerance 0.001); max |dlogit| vs the Python reference over the valid (word, label) slots 4.578e-05.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Python LiteRT (ai-edge-litert 2.1.6 CompiledModel, CPU, 4 threads, fp16 host table) reference logits of the same graph on the same request (~/code/codex-conversions/2026-09-26/gliformer-android/fixtures/references_fp16.json + fixtures/logits_fp16/corpus_NN_s128.bin, the app's matched-fp16 acceptance reference; per-file sha256 checked), recomputed 2026-09-26 from the saved raw outputs: 60/60 fixtures pass over all 48 x 15 slots (0 violating elements of 43,200); max |d| 4.578e-05, max rel 5.270e-04.",
            "latency: graph phase = first input write through synchronized output readback; 5 timed samples per row after the warm-up, 60 rows = 300 samples; median 189.90 / min 119.75 / max 211.45 ms over all samples; row medians: median 188.95, first three 127.60, 122.39, 122.29 -> last three 206.53, 205.36, 205.37 ms; tokenize+lookup median 5.44 ms, decode 0.203 ms (host phases, outside latency_p50_ms). This gate ran right after the GPU gate on the same warm phone."
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": 189.904922,
          "loads": true,
          "max_abs_diff": 4.57763671875e-05,
          "max_rel_diff": 0.0005269837565720081,
          "metrics": {
            "battery_temp_c_end": 38.2,
            "battery_temp_c_start": 37,
            "cpu_threads": 4,
            "decode_median_ms": 0.20278649999999998,
            "iterations": 300,
            "latency_max_ms": 211.448073,
            "latency_min_ms": 119.750937,
            "load_ms": 1981.520364,
            "max_score_diff_vs_official": 0.0006461739540100098,
            "max_score_diff_vs_python_reference": 1.9073486328125e-06,
            "max_valid_logit_diff_vs_python_reference": 4.57763671875e-05,
            "oracle_span_sets_identical": 60,
            "pipeline_total_median_ms": 195.67973949999998,
            "pss_after_load_bytes": 2585922560,
            "python_span_sets_identical": 60,
            "repetitions_per_row": 5,
            "row_median_first_ms": 127.595208,
            "row_median_last_ms": 205.3675,
            "row_median_ms": 188.9513545,
            "rows": 60,
            "shared_rule_pass_rows": 60,
            "shared_rule_violating_elements": 0,
            "tokenize_lookup_median_ms": 5.4434895,
            "tokenizer_checks_identical": 80,
            "vmhwm_after_gate_bytes": 2756358144,
            "vmhwm_after_load_bytes": 2713726976,
            "vmrss_after_gate_bytes": 2731327488,
            "vmrss_after_load_bytes": 2663985152,
            "warmup_passes": 12,
            "warmup_total_ms": 1488.60677
          },
          "output_match": true,
          "peak_mem_mb": 2756.36,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK 4 threads; Android sample app (debug build)",
          "total_ops": 4149,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 4149,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round4-six-inputs-memory-default-probe-v2; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/011_hash_gliformer_large_ner_s128_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round4_s128_wfp16_default.json (gate s128_wfp16_default, status PROBE: 'GPU DEFAULT (GpuOptions omitted)', not a shipping configuration), raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/s128_wfp16_default/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/030_run_s128_wfp16_default.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/032_own_pid_logcat_s128_wfp16_default.stdout.log.",
            "logcat/stderr: 'Replacing 4149 out of 4149 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).'; compile status 0 in 4692.3 ms; is_fully_accelerated=true; hardware_accelerators 'GPU only (bitmask 2)'; gpu_options_toml 'omitted: runtime DEFAULT precision'.",
            "conditions (round 4, the four s128 jobs ran back to back in one adb session, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before the s128 shape 29.9 C 94 % (2026-09-26T02:03:22.264820+00:00), after this job 33.9 C 94 % (2026-09-26T02:05:21.447521+00:00); runner wall 5.5 s; compile 4692.3 ms (status 0); process peak RSS (VmHWM) 3,592,392,704 B, RSS after the last readback 1,738,186,752 B. This probe ran last of the four s128 jobs (after the CPU job), 0 warm-ups and 1 timed run per fixture.",
            "parity: 0/10 fixtures' entity span sets equal the official gliformer 0.1.2 fp32 result; the decoder returned 0 entities where the official result has 37 (0 matched spans); all outputs finite: True; max |dlogit| vs official 2.680e+01; packed-logit norm ratio vs official 0.0035-0.0060 (|logit| max 0.0718 vs the reference's 24.67: finite but collapsed logits, a correctness failure, not a compile failure). Cause not established (no tensor-level bisect; no op named by any log).",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/litert_cpu_outputs_round2_s128_wfp16.npz; the device pack android/fixtures/s128/*.f32 is byte-identical to the fixtures/graph_inputs_NN.npz rows of the float32 table that run consumed, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 0/10 fixtures pass over all 48 x 15 slots (2355 violating elements of 7,200; worst element |d| 2.423e+01 at |ref| 2.430e+01); max |d| 2.680e+01, max rel 1.344e+00. Against this device's own CPU XNNPACK run: 0/10 pass, 2355 violating elements, max |d| 2.680e+01.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 0 warm-ups and 1 timed run per fixture, 10 fixtures = 10 timed runs; median 47.00 / min 46.28 / max 51.48 ms (the first call of every fixture, so not comparable with the warmed FP32 row)."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 46.997577500000006,
          "loads": true,
          "max_abs_diff": 26.80441665649414,
          "max_rel_diff": 1.3435370922088623,
          "metrics": {
            "battery_temp_c_end": 33.9,
            "battery_temp_c_start": 29.9,
            "decoded_entities": 0,
            "exact_span_sets_vs_official": 0,
            "fixtures": 10,
            "iterations": 10,
            "latency_max_ms": 51.482291,
            "latency_min_ms": 46.282448,
            "load_compile_ms": 4692.345207,
            "loaded_rss_bytes": 1738186752,
            "max_abs_dlogit_vs_official": 26.803625106811523,
            "norm_ratio_vs_official_max": 0.006013061620896899,
            "norm_ratio_vs_official_min": 0.0034750215214196358,
            "oracle_entities": 37,
            "peak_rss_bytes": 3592392704,
            "per_fixture_median_first_ms": 51.482291,
            "per_fixture_median_last_ms": 48.322604,
            "repetitions_per_fixture": 1,
            "shared_rule_pass_fixtures": 0,
            "shared_rule_violating_elements": 2355,
            "shared_rule_vs_device_cpu_pass_fixtures": 0,
            "shared_rule_vs_device_cpu_violating_elements": 2355,
            "warmups_per_fixture": 0
          },
          "output_match": false,
          "peak_mem_mb": 3592.39,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions omitted (runtime default GPU precision)",
          "total_ops": 4149,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 4149,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round4-six-inputs-memory-default-probe-v2; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/011_hash_gliformer_large_ner_s128_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round4_s128_wfp16.json (gate s128_wfp16, mode fp32), raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/s128_wfp16/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/013_run_s128_wfp16.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round4/s128/015_own_pid_logcat_s128_wfp16.stdout.log; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s128_wfp16.json.",
            "logcat/stderr: 'Replacing 4149 out of 4149 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).'; compile status 0 in 3416.5 ms; is_fully_accelerated=true; hardware_accelerators 'GPU only (bitmask 2)'; gpu_options_toml 'precision = 2'.",
            "conditions (round 4, the four s128 jobs ran back to back in one adb session, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before the s128 shape 29.9 C 94 % (2026-09-26T02:03:22.264820+00:00), after this job 29.9 C 94 % (2026-09-26T02:03:55.115165+00:00); runner wall 9.6 s; compile 3416.5 ms (status 0); process peak RSS (VmHWM) 4,542,996,480 B, RSS after the last readback 2,591,805,440 B.",
            "parity: 10/10 fixtures' entity span sets (label, start, end) equal the official gliformer 0.1.2 fp32 CPU result (checkpoint rev d0a4e53d, same tokenizer, threshold 0.5); all outputs finite: True; max score diff 2.406e-04 (gate tolerance 0.01); max |dlogit| vs official 7.504e-03.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/litert_cpu_outputs_round2_s128_wfp16.npz; the device pack android/fixtures/s128/*.f32 is byte-identical to the fixtures/graph_inputs_NN.npz rows of the float32 table that run consumed, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 10/10 fixtures pass over all 48 x 15 slots (0 violating elements of 7,200); max |d| 6.104e-05, max rel 1.359e-04. The same rule against this device's own CPU XNNPACK run of the same graph and inputs (the cpu_xnnpack row): 10/10 pass, 0 violating elements, max |d| 7.439e-05, max rel 2.451e-04.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 5 timed runs after 2 warm-ups per fixture, 10 fixtures = 50 timed runs; median 82.27 / min 80.84 / max 86.15 ms; per-fixture medians 82.40 ms (first) -> 82.87 ms (last) in run order.",
            "cold start (round 9, a separate fresh native process, build gliformer-round9-cold-process-start-v1, same graph and device, GpuOptions precision=FP32, 0 warm-ups; ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/cold_start_round9.json, public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s128_cold_start.json): process start -> first result 5022.7 ms (main entry -> first result 4976.2 ms; compile 4707.9 ms; first inference readback 90.53 ms), second call 80.24 ms; peak RSS 4,556,410,880 B, loaded 2,602,950,656 B; existing runtime/driver disk caches were not cleared (process cold start, not a cache-purged boot); the first call's entity span set equals the official result (5/5)."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 82.27380199999999,
          "loads": true,
          "max_abs_diff": 6.103515625e-05,
          "max_rel_diff": 0.0001359053421765566,
          "metrics": {
            "battery_temp_c_end": 29.9,
            "battery_temp_c_start": 29.9,
            "cold_start_compile_ms": 4707.870935,
            "cold_start_first_inference_ms": 90.525365,
            "cold_start_loaded_rss_bytes": 2602950656,
            "cold_start_peak_rss_bytes": 4556410880,
            "cold_start_process_start_to_first_result_ms": 5022.70679997,
            "cold_start_second_call_ms": 80.239375,
            "exact_span_sets_vs_official": 10,
            "fixtures": 10,
            "iterations": 50,
            "latency_max_ms": 86.147917,
            "latency_min_ms": 80.840416,
            "load_compile_ms": 3416.543801,
            "loaded_rss_bytes": 2591805440,
            "max_abs_dlogit_vs_official": 0.007504463195800781,
            "max_score_diff_vs_official": 0.00024056434631347656,
            "peak_rss_bytes": 4542996480,
            "per_fixture_median_first_ms": 82.398542,
            "per_fixture_median_last_ms": 82.869115,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 10,
            "shared_rule_violating_elements": 0,
            "shared_rule_vs_device_cpu_max_abs_diff": 7.43865966796875e-05,
            "shared_rule_vs_device_cpu_pass_fixtures": 10,
            "shared_rule_vs_device_cpu_violating_elements": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": 4543.0,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions precision=FP32",
          "total_ops": 4149,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 4149,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Android sample app gate (litert-community/GLiFormer-Large-NER-LiteRT android/, package com.gliformer, debug build APK sha256 3bf2de72f19f... (installed 2026-09-26 21:29 JST, ~/code/codex-conversions/2026-09-26/gliformer-android/results/round5/install.json), com.google.ai.edge.litert:litert:2.2.0, round 5): one Environment, Accelerator.GPU with GpuOptions(precision = FP32) (explicit), one window (128) resident; the app builds the inputs itself (whitespace word splitter, SentencePiece Unigram tokenizer, memory-mapped float16 table upcast to float32, routing construction) and its token ids / first-subtoken / parent / entity positions / masks equal the captured Python inputs on all 80/80 tokenizer checks (400 raw graph-tensor comparisons). Sources: ~/code/codex-conversions/2026-09-26/gliformer-android/results/gate_s128_gpu_r5.json, raw logits ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s128_gpu_r5/raw_logits.tar (per-row sha256 in the JSON), logcat ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s128_gpu_r5/logcat_own_pid.log.",
            "logcat: '09-26 21:29:50.648 18266 18323 I tflite  : Replacing 4149 out of 4149 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).'.",
            "conditions: screen unlocked and awake (mDreamingLockscreen false, mWakefulness Awake), 'svc power stayon usb' on, USB powered; app battery 33.2 C 80 % before -> 37 C after the gate (2026-09-26T12:29:49.588977Z -> 2026-09-26T12:30:39.350418Z); load 3157.9 ms, 12 full-pipeline warm-up passes 1087.8 ms before the timed rows; memory after load VmRSS 2,463,789,056 B / VmHWM 4,389,380,096 B / total PSS 4,202,232,832 B, after the gate VmRSS 2,472,722,432 B / VmHWM 4,389,380,096 B. No GPU clock / thermal_status sample in this gate.",
            "parity: 60/60 inputs' entity span sets equal the official gliformer 0.1.2 fp32 result and 60/60 equal the Python LiteRT 2.1.6 CPU fp16-table reference (all 5 repetitions); all logits finite; max score diff vs the matched-fp16 Python reference 3.874e-06 (gate tolerance 1e-05), vs official 6.465e-04 (tolerance 0.001); max |dlogit| vs the Python reference over the valid (word, label) slots 8.059e-05.",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Python LiteRT (ai-edge-litert 2.1.6 CompiledModel, CPU, 4 threads, fp16 host table) reference logits of the same graph on the same request (~/code/codex-conversions/2026-09-26/gliformer-android/fixtures/references_fp16.json + fixtures/logits_fp16/corpus_NN_s128.bin, the app's matched-fp16 acceptance reference; per-file sha256 checked), recomputed 2026-09-26 from the saved raw outputs: 59/60 fixtures pass over all 48 x 15 slots (1 violating elements of 43,200; worst element |d| 1.641e-05 at |ref| 1.423e-02); max |d| 8.059e-05, max rel 1.153e-03. (the one violating element sits at 1.15e-03 relative, above the 1e-3 relative tolerance by 15 %; every span set is exact).",
            "latency: graph phase = first input write through synchronized output readback (encoder_readback in the app's timing); 5 timed samples per row after the warm-up, 60 rows = 300 samples; median 116.22 / min 80.73 / max 171.52 ms over all samples; row medians: median 115.58, first three 81.50, 80.96, 81.16 -> last three 147.21, 148.71, 147.13 ms (the phone warms over the gate); tokenize+lookup median 13.59 ms, decode 0.573 ms (host phases, outside latency_p50_ms).",
            "first tap after launch (benchmark build APK sha256 79fb0bef8cba..., non-debuggable, a separate app process pid 21047, screen interactive, GPU FP32, fp16 table, window 128, the same device; ~/code/codex-conversions/2026-09-26/gliformer-android/results/first_tap_r5.json): Activity.onCreate -> Ready 4775.4 ms (load 3404.7 + 12 warm-up passes 1368.0 ms; host command -> observed Ready upper bound 4963.4 ms); tap -> result 168.06 ms (tap 1) ... 154.81 ms (tap 5), tap -> rendered 190.47 / 167.15 ms; graph readback per tap 145.87, 153.07, 143.15, 145.80, 143.07 ms; battery 40.7 -> 40.5 C."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 116.21625,
          "loads": true,
          "max_abs_diff": 8.058547973632812e-05,
          "max_rel_diff": 0.0011530901538208127,
          "metrics": {
            "battery_temp_c_end": 37,
            "battery_temp_c_start": 33.2,
            "decode_median_ms": 0.5731245,
            "first_tap_1_graph_readback_ms": 145.8675,
            "first_tap_1_to_render_ms": 190.465781,
            "first_tap_1_to_result_ms": 168.060573,
            "first_tap_5_graph_readback_ms": 143.066354,
            "first_tap_5_to_render_ms": 167.149635,
            "first_tap_5_to_result_ms": 154.813749,
            "first_tap_launch_to_ready_host_upper_bound_ms": 4963.370207929984,
            "first_tap_launch_to_ready_ms": 4775.414842,
            "first_tap_startup_load_ms": 3404.723957,
            "first_tap_startup_warmup_ms": 1367.992447,
            "iterations": 300,
            "latency_max_ms": 171.518385,
            "latency_min_ms": 80.734896,
            "load_ms": 3157.873696,
            "max_score_diff_vs_official": 0.0006464719772338867,
            "max_score_diff_vs_python_reference": 3.874301910400391e-06,
            "max_valid_logit_diff_vs_python_reference": 8.058547973632812e-05,
            "oracle_span_sets_identical": 60,
            "pipeline_total_median_ms": 127.5003645,
            "pss_after_load_bytes": 4202232832,
            "python_span_sets_identical": 60,
            "repetitions_per_row": 5,
            "row_median_first_ms": 81.49625,
            "row_median_last_ms": 147.125781,
            "row_median_ms": 115.5812495,
            "rows": 60,
            "shared_rule_pass_rows": 59,
            "shared_rule_violating_elements": 1,
            "tokenize_lookup_median_ms": 13.588568,
            "tokenizer_checks_identical": 80,
            "vmhwm_after_gate_bytes": 4389380096,
            "vmhwm_after_load_bytes": 4389380096,
            "vmrss_after_gate_bytes": 2472722432,
            "vmrss_after_load_bytes": 2463789056,
            "warmup_passes": 12,
            "warmup_total_ms": 1087.803958
          },
          "output_match": false,
          "peak_mem_mb": 4389.38,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions precision=FP32; Android sample app (debug build)",
          "total_ops": 4149,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s128_wfp16__galaxy-s26.json"
      }
    ]
  },
  "model": {
    "family": "gliformer-large-ner",
    "id": "gliformer-large-ner__gliformer_large_ner_s128_wfp16",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/GLiFormer-Large-NER-LiteRT",
    "task": "token-classification"
  },
  "pitfalls": [
    "Weight storage and GPU computation precision are separate settings: with GpuOptions omitted (the runtime's default GPU precision) the s128 wfp16 graph compiled on the Galaxy S26 and returned finite logits, but no entities were extracted — the Kotlin block's comment reads 'default precision returns no entities' and HOST_CONTRACT.md 'Numerical and memory limits' says 'Explicit GPU FP32 computation is required; default GPU precision failed entity agreement'. Set GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32) on every GPU graph (card 'Kotlin'; HOST_CONTRACT.md).",
    "The graphs cover the encoder and the NER head only: the host looks up the word embeddings before the graph (host_assets/word_embeddings_fp32.bin, headerless little-endian float32 [128008,1024], 524,320,768 B, the default; or word_embeddings_fp16.bin, 262,160,384 B, upcast to float32 on lookup — gated at s128 only, longer-window fp16-table accuracy was not gated) and decodes the [1,1,T,15] start/end/inside logits with the upstream pairing decoder after it (threshold 0.5, flat spans). Build the encoded sequence with the pinned gliformer 0.1.2 processor ([SCHEMA] parent token, the five [ENTITY] label pairs, the separator, the text words; right-pad ids with 0 to N): a tokenizer call over a hand-concatenated prompt is not equivalent (HOST_CONTRACT.md 'Prompt, embeddings and routing').",
    "Bind the input buffers by name: the native subgraph (CompiledModel) buffer order is inputs_embeds [1,128,1024], attention_mask [1,128], text_routing [1,48,128], parent_routing [1,1,128], label_routing [1,5,128], text_mask [1,48], but the flatbuffer signature map lists the names alphabetically (attention_mask, inputs_embeds, label_routing, parent_routing, text_mask, text_routing) — resolve the names against the loaded signature rather than sorting them; every tensor is float32 (attention_mask 1.0 for real tokens, routing one-hot at the first sub-token / [SCHEMA] / [ENTITY] positions, text_mask 1.0 for real words); output_0 [1,1,48,15] reshapes to [1,48,5,3] with start, end, inside logits on the last axis — not mutually exclusive BIO classes (HOST_CONTRACT.md 'Capacities and execution', 's128'; graph_contract_s128.json).",
    "Fixed windows and a fixed class axis: N = 128 (T = 48 text words) / 256 / 512 encoded tokens including the complete label prompt, exactly five labels in a fixed order (the measured order is person, organization, location, product, date; other five-label sets are accepted by the API but were not numerically gated), batch 1. Inputs above the s512 limits raise an error naming both capacities; nothing is truncated or chunked implicitly — chunk_by_sentences(text, max_words) is a caller-side helper whose entity offsets are local to each chunk (card 'Host contract'; HOST_CONTRACT.md 'One extraction call').",
    "s256 and s512 ship as an encoder graph plus a head graph because the single s256 graph does not compile on the LiteRT 2.2.0 GPU: the checkpoint's word-level BiLSTM is unrolled for T steps and unrolled heads of 6,286 operators or more crash the GPU compiler (the head crashes at 570 MB RSS — a runtime defect reported with reproducers, not a memory limit), while the DeBERTa encoder alone (1,773 operators) compiles at every window. Create two CompiledModels — the encoder with Accelerator.GPU and FP32 precision, the head with Accelerator.CPU — and pass the encoder's [1,1,N,1024] output buffer to the head as its first input; the numbers are exact either way, only the placement differs (card 'What you get', 'Kotlin').",
    "Memory is the cost of this model (Galaxy S26, native CompiledModel processes without the token table, tokenizer or decoder): the 707 MB s128 graph is 2,591,805,440 B resident after loading and peaks at 4,542,996,480 B while the GPU delegate compiles it; the s256 encoder on the GPU with the fp32 head on the CPU in one process is 4,571,271,168 B resident; the s512 head alone is 4,676,505,600 B resident on the CPU — so treat s256 as the practical top window in an app, chunk longer documents by sentence, and plan for flagship-class phones; the token table adds 262 MB (fp16) or 524 MB (fp32) when resident (card 'What you get', memory table).",
    "Cold start: the first call after process start took 5,023 ms (compile 4,708 ms) for the s128 wfp16 graph in a fresh native process and 80 ms afterwards; the Android sample warms the whole pipeline 12 times before Ready and reported launch -> Ready 4,775 ms and a first Extract tap of 168 ms (fifth 155 ms) in a non-debuggable build with the screen on (card 'What you get', 'Kotlin', 'Measured quality and performance').",
    "Returned start / end are Unicode code-point offsets into the original Python string (end exclusive; text[start:end] is the entity text), not UTF-8 bytes and not Kotlin UTF-16 indices — translate them when supplementary characters occur (HOST_CONTRACT.md 'Output decoding').",
    "Validation scope: GPU execution was validated on the Galaxy S26 (SM-S942Q, SM8850 Adreno, Android 16, LiteRT 2.2.0) only — other Android GPU families, lower-memory phones and NPUs were not evaluated; English only, one text per call, the NER path only (classification, relation, structuring and embedding heads are not converted); the quality figures are agreement with the official gliformer 0.1.2 fp32 implementation (micro-F1 1.000, 70/70 identical span sets on desktop CPU; 10/10, 15/15, 20/20 on the phone), not a human-labelled accuracy estimate; latency is a single-device sample at battery 30-41 C with the phone idle — the s512 encoder number was taken at 41 C (card 'Measured quality and performance', 'Provenance, conversion and license')."
  ],
  "schema_version": "1.2"
}
