{
  "artifacts": [
    {
      "file": "gliformer_large_ner_s256_head_wfp16.tflite",
      "sha256": "de4c62cacdf14ae8fd412d50cf958ee5b424c3062cca70fe13e290d85e9c1a30",
      "size_mb": 50.99
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": 2425.31,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": true
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-26/gliformer-large-ner__gliformer_large_ner_s256_head_wfp16.json"
  },
  "conversion": {
    "command": "not published as a script (card 'Provenance, conversion and license'): litert-torch 0.9.3 fixed-shape fp32 export of the NER path of knowledgator/gliformer-large-v1 (rev d0a4e53d) re-expressed for the GPU delegate without changing its math — host-side token lookup, one-hot routing as matmul, float masks, attention at rank 4, the exact DeBERTa logarithmic relative-position buckets as projected tables, the word BiLSTM unrolled for the window, one packed output tensor; then ai-edge-quantizer 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights (the wfp16 file); the layout and page embeddings of the backbone are skipped exactly as the upstream text path skips them",
    "quantization": "float16 weight storage (ai-edge-quantizer 0.8.0 FLOAT_CASTING on the FULLY_CONNECTED weights only): 9 float16 weight tensors read through a DEQUANTIZE to float32, every activation and every other constant float32; 12,559 operators in the fp32 reference (gliformer_large_ner_s256_head_fp32.tflite, 103,920,220 B, published as the exact reference) — 12,568 operators in this file; dynamic-range INT8 was not attempted — it does not compile on the LiteRT 2.2.0 GPU delegate for this graph family (card 'Provenance, conversion and license')",
    "tool": "litert-torch",
    "tool_version": "0.9.3 (torch 2.12.1; ai-edge-litert 2.1.6; ai-edge-quantizer 0.8.0 for the float16 weight storage; gliformer 0.1.2, gliner 0.2.29, transformers 5.16.1 — requirements-lock.txt / upstream.json)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "LOGISTIC",
      "TANH"
    ],
    "coverage_ops_pct": 79.6,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 10010,
      "unmatched": 2558
    },
    "partitions": 2568
  },
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 12565,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round5-variable-input-probes-loaded-rss-v3; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s256_cpu/007_hash_gliformer_large_ner_s256_head_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s256_split.json (head_cpu) + ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/device_round9_head_s256_cpu.json, raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s256_cpu/output/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s256_cpu/017_run_graph.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s256_cpu/019_own_pid_logcat.stdout.log; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s256_split.json. Inputs: the round-9 GPU encoder readbacks of this window plus the captured routing tensors.",
            "logcat/stderr: 'Created TensorFlow Lite XNNPACK delegate for CPU.' | 'Replacing 12565 out of 12568 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main).'; compile status 0 in 2059.1 ms; hardware_accelerators 'CPU only (bitmask 1)'; native default CPU options (thread count not set by the runner); the three non-XNNPACK nodes are not named by the log.",
            "conditions (round 9/10 stage job 'head_s256_cpu', one native process for this graph only, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before 36.7 C 86 % -> after 36.7 C 86 %; runner wall 21.5 s; compile 2059.1 ms (status 0); compile peak RSS (VmHWM) 2,396,839,936 B, process peak 2,414,317,568 B, loaded RSS after the last readback 2,413,006,848 B (this graph alone; the two graphs of a split window were measured in separate processes, so these are stage values, not the application's residency).",
            "pair (round 9 s256 split gate, ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s256_split.json; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s256_split.json): the pulled GPU encoder hidden states of this stage fed the wfp16 head on the CPU in a separate process; 15/15 fixtures' entity span sets equal the official gliformer 0.1.2 fp32 result, max score diff 2.414e-04, max |dlogit| vs official 1.484e-02, all finite; sum of the two stage medians 342.59 ms (not a same-process end-to-end measurement).",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/fixtures/round9/device_s256_head/cpu_outputs.npz, written by scripts/prepare_split_device_round9.py from the very fixtures/round9/device_s{seq}_head/NN.npz the device consumed, whose encoder_hidden arrays equal the round-9 GPU encoder readbacks byte for byte, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 15/15 fixtures pass over all 256 x 15 slots (0 violating elements of 57,600); max |d| 1.526e-05, max rel 2.530e-05.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 5 timed runs after 2 warm-ups per fixture, 15 fixtures = 75 timed runs; median 167.03 / min 161.92 / max 193.32 ms; per-fixture medians 162.97 ms (first) -> 181.16 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": 167.031874,
          "loads": true,
          "max_abs_diff": 1.52587890625e-05,
          "max_rel_diff": 2.530466554162558e-05,
          "metrics": {
            "battery_temp_c_end": 36.7,
            "battery_temp_c_start": 36.7,
            "compile_peak_rss_bytes": 2396839936,
            "fixtures": 15,
            "iterations": 75,
            "latency_max_ms": 193.319427,
            "latency_min_ms": 161.921354,
            "load_compile_ms": 2059.099791,
            "loaded_rss_bytes": 2413006848,
            "pair_exact_span_sets_vs_official": 15,
            "pair_max_abs_dlogit_vs_official": 0.014841079711914062,
            "pair_max_score_diff_vs_official": 0.00024139881134033203,
            "pair_sum_of_stage_medians_ms": 342.588644,
            "peak_rss_bytes": 2414317568,
            "per_fixture_median_first_ms": 162.967135,
            "per_fixture_median_last_ms": 181.163177,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 15,
            "shared_rule_violating_elements": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": 2414.32,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK (native runner, default threads)",
          "total_ops": 12568,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s256_head_wfp16__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 12565,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Android sample app gate (litert-community/GLiFormer-Large-NER-LiteRT android/, package com.gliformer, debug build APK sha256 3bf2de72f19f... (installed 2026-09-26 21:29 JST, ~/code/codex-conversions/2026-09-26/gliformer-android/results/round5/install.json), com.google.ai.edge.litert:litert:2.2.0, round 5): one Environment, the s256 split: the s256 wfp16 encoder with Accelerator.GPU and GpuOptions(precision = FP32) and this head with Accelerator.CPU and CpuOptions(numThreads = 4) (GliformerExtractor.kt) resident together, window 256 forced for every input; the app builds the inputs itself (whitespace word splitter, SentencePiece Unigram tokenizer, memory-mapped float16 table upcast to float32, routing construction) and its token ids / first-subtoken / parent / entity positions / masks equal the captured Python inputs on all 80/80 tokenizer checks (400 raw graph-tensor comparisons). Sources: ~/code/codex-conversions/2026-09-26/gliformer-android/results/gate_s256_split_r5.json, raw logits ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s256_split_r5/raw_logits.tar (per-row sha256 in the JSON), logcat ~/code/codex-conversions/2026-09-26/gliformer-android/logs/round5/gate_s256_split_r5/logcat_own_pid.log.",
            "logcat: '09-26 21:33:23.720 19628 19681 I tflite  : Replacing 1773 out of 1773 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).' | '09-26 21:33:26.254 19628 19681 I tflite  : Created TensorFlow Lite XNNPACK delegate for CPU.' | '09-26 21:33:26.256 19628 19681 I tflite  : Replacing 12565 out of 12568 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main).'.",
            "conditions: screen unlocked and awake (mDreamingLockscreen false, mWakefulness Awake), 'svc power stayon usb' on, USB powered; app battery 38.2 C 80 % before -> 40.8 C after the gate (2026-09-26T12:33:22.629435Z -> 2026-09-26T12:36:26.850167Z); load 4185.2 ms, 12 full-pipeline warm-up passes 3993.7 ms before the timed rows; memory after load VmRSS 4,267,446,272 B / VmHWM 4,317,204,480 B / total PSS 5,569,874,944 B, after the gate VmRSS 4,390,236,160 B / VmHWM 4,414,648,320 B. No GPU clock / thermal_status sample in this gate.",
            "parity: 65/65 inputs' entity span sets equal the official gliformer 0.1.2 fp32 result and 65/65 equal the Python LiteRT 2.1.6 CPU fp16-table reference (all 5 repetitions); all logits finite; max score diff vs the matched-fp16 Python reference 4.232e-06 (gate tolerance 1e-05), vs official 6.465e-04 (tolerance 0.001); max |dlogit| vs the Python reference over the valid (word, label) slots 7.415e-05.",
            "output_match left null: this head consumed the app's own GPU encoder hidden states, and the Python reference's head consumed the Mac CPU encoder's, so the head alone has no CPU reference on identical inputs. Pair-level shared element rule against the Python LiteRT (ai-edge-litert 2.1.6 CompiledModel, CPU, 4 threads, fp16 host table) reference logits of the same graph on the same request (~/code/codex-conversions/2026-09-26/gliformer-android/fixtures/references_fp16.json + fixtures/logits_fp16/corpus_NN_s256.bin, the app's matched-fp16 acceptance reference; per-file sha256 checked): 64/65 inputs pass over all 256 x 15 slots (1 violating element of 249,600; worst |d| 1.643e-05 at |ref| 1.423e-02), max |d| 7.415e-05, max rel 1.154e-03; every span set exact.",
            "latency: head phase = this graph's input write through synchronized output readback inside the app's graph phase; 5 timed samples per row, 65 rows = 325 samples; median 128.38 / min 114.31 / max 142.39 ms; row medians: median 128.26, first three 117.24, 118.48, 123.43 -> last three 134.69, 135.09, 135.79 ms; the pair's graph phase median 525.67 ms, encoder phase median 390.95 ms (the s256 encoder row). Both graphs resident: memory after load VmRSS 4,267,446,272 B."
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": 128.378021,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "battery_temp_c_end": 40.8,
            "battery_temp_c_start": 38.2,
            "cpu_threads": 4,
            "decode_median_ms": 0.495573,
            "encoder_median_ms": 390.950938,
            "head_iterations": 325,
            "head_latency_max_ms": 142.390625,
            "head_latency_min_ms": 114.312448,
            "head_row_median_first_ms": 117.243958,
            "head_row_median_last_ms": 135.786719,
            "head_row_median_ms": 128.257708,
            "load_ms": 4185.194165,
            "max_score_diff_vs_official": 0.0006464719772338867,
            "max_score_diff_vs_python_reference": 4.231929779052734e-06,
            "max_valid_logit_diff_vs_python_reference": 7.414817810058594e-05,
            "oracle_span_sets_identical": 65,
            "pair_graph_median_ms": 525.669115,
            "pair_graph_row_median_ms": 525.669115,
            "pair_shared_rule_max_abs_diff": 7.414817810058594e-05,
            "pair_shared_rule_max_rel_diff": 0.0011543332366272807,
            "pair_shared_rule_pass_rows": 64,
            "pair_shared_rule_violating_elements": 1,
            "pipeline_total_median_ms": 540.003125,
            "pss_after_load_bytes": 5569874944,
            "python_span_sets_identical": 65,
            "repetitions_per_row": 5,
            "rows": 65,
            "tokenize_lookup_median_ms": 12.128593,
            "tokenizer_checks_identical": 80,
            "vmhwm_after_gate_bytes": 4414648320,
            "vmhwm_after_load_bytes": 4317204480,
            "vmrss_after_gate_bytes": 4390236160,
            "vmrss_after_load_bytes": 4267446272,
            "warmup_passes": 12,
            "warmup_total_ms": 3993.714999
          },
          "output_match": null,
          "peak_mem_mb": 4414.65,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK 4 threads; Android sample app (debug build)",
          "total_ops": 12568,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s256_head_wfp16__galaxy-s26.json"
      }
    ]
  },
  "model": {
    "family": "gliformer-large-ner",
    "id": "gliformer-large-ner__gliformer_large_ner_s256_head_wfp16",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/GLiFormer-Large-NER-LiteRT",
    "task": "token-classification"
  },
  "pitfalls": [
    "The head graphs run on the CPU on Android (Accelerator.CPU): unrolled BiLSTM heads of this size crash the LiteRT 2.2.0 GPU compiler (see the split pitfall), so no GPU precision setting applies to this file; the encoder that feeds it needs GpuOptions(precision = FP32) — HOST_CONTRACT.md 'Numerical and memory limits': 'Explicit GPU FP32 computation is required; default GPU precision failed entity agreement' (card 'Kotlin'; HOST_CONTRACT.md).",
    "The graphs cover the encoder and the NER head only: the host looks up the word embeddings before the graph (host_assets/word_embeddings_fp32.bin, headerless little-endian float32 [128008,1024], 524,320,768 B, the default; or word_embeddings_fp16.bin, 262,160,384 B, upcast to float32 on lookup — gated at s128 only, longer-window fp16-table accuracy was not gated) and decodes the [1,1,T,15] start/end/inside logits with the upstream pairing decoder after it (threshold 0.5, flat spans). Build the encoded sequence with the pinned gliformer 0.1.2 processor ([SCHEMA] parent token, the five [ENTITY] label pairs, the separator, the text words; right-pad ids with 0 to N): a tokenizer call over a hand-concatenated prompt is not equivalent (HOST_CONTRACT.md 'Prompt, embeddings and routing').",
    "Bind the input buffers by name: the head's native buffer order is encoder_hidden [1,1,256,1024], text_routing [1,256,256], parent_routing [1,1,256], label_routing [1,5,256], text_mask [1,256] (the flatbuffer signature map sorts the names: encoder_hidden, label_routing, parent_routing, text_mask, text_routing); the head performs the text, parent and label routing matmuls internally and includes one bidirectional LSTM layer, anchor fusion and the span scorer; output_0 [1,1,256,15] reshapes to [1,256,5,3] with start, end, inside logits on the last axis (HOST_CONTRACT.md 's256'; graph_contract_s256.json).",
    "Fixed windows and a fixed class axis: N = 128 (T = 48 text words) / 256 / 512 encoded tokens including the complete label prompt, exactly five labels in a fixed order (the measured order is person, organization, location, product, date; other five-label sets are accepted by the API but were not numerically gated), batch 1. Inputs above the s512 limits raise an error naming both capacities; nothing is truncated or chunked implicitly — chunk_by_sentences(text, max_words) is a caller-side helper whose entity offsets are local to each chunk (card 'Host contract'; HOST_CONTRACT.md 'One extraction call').",
    "s256 and s512 ship as an encoder graph plus a head graph because the single s256 graph does not compile on the LiteRT 2.2.0 GPU: the checkpoint's word-level BiLSTM is unrolled for T steps and unrolled heads of 6,286 operators or more crash the GPU compiler (the head crashes at 570 MB RSS — a runtime defect reported with reproducers, not a memory limit), while the DeBERTa encoder alone (1,773 operators) compiles at every window. Create two CompiledModels — the encoder with Accelerator.GPU and FP32 precision, the head with Accelerator.CPU — and pass the encoder's [1,1,N,1024] output buffer to the head as its first input; the numbers are exact either way, only the placement differs (card 'What you get', 'Kotlin').",
    "Memory is the cost of this model (Galaxy S26, native CompiledModel processes without the token table, tokenizer or decoder): the 707 MB s128 graph is 2,591,805,440 B resident after loading and peaks at 4,542,996,480 B while the GPU delegate compiles it; the s256 encoder on the GPU with the fp32 head on the CPU in one process is 4,571,271,168 B resident; the s512 head alone is 4,676,505,600 B resident on the CPU — so treat s256 as the practical top window in an app, chunk longer documents by sentence, and plan for flagship-class phones; the token table adds 262 MB (fp16) or 524 MB (fp32) when resident (card 'What you get', memory table).",
    "Cold start: the first call after process start took 5,023 ms (compile 4,708 ms) for the s128 wfp16 graph in a fresh native process and 80 ms afterwards; the Android sample warms the whole pipeline 12 times before Ready and reported launch -> Ready 4,775 ms and a first Extract tap of 168 ms (fifth 155 ms) in a non-debuggable build with the screen on (card 'What you get', 'Kotlin', 'Measured quality and performance').",
    "Returned start / end are Unicode code-point offsets into the original Python string (end exclusive; text[start:end] is the entity text), not UTF-8 bytes and not Kotlin UTF-16 indices — translate them when supplementary characters occur (HOST_CONTRACT.md 'Output decoding').",
    "Validation scope: GPU execution was validated on the Galaxy S26 (SM-S942Q, SM8850 Adreno, Android 16, LiteRT 2.2.0) only — other Android GPU families, lower-memory phones and NPUs were not evaluated; English only, one text per call, the NER path only (classification, relation, structuring and embedding heads are not converted); the quality figures are agreement with the official gliformer 0.1.2 fp32 implementation (micro-F1 1.000, 70/70 identical span sets on desktop CPU; 10/10, 15/15, 20/20 on the phone), not a human-labelled accuracy estimate; latency is a single-device sample at battery 30-41 C with the phone idle — the s512 encoder number was taken at 41 C (card 'Measured quality and performance', 'Provenance, conversion and license').",
    "An fp32 head does not buy memory: on the Galaxy S26 the fp32 s256 head loaded to 2,464,145,408 B against the wfp16 head's 2,413,006,848 B and ran 8 % faster (153.3 vs 167.0 ms median), so the fp16-weight heads stay the default; HostRuntime(head_storage='fp32') selects the references (card 'Measured quality and performance')."
  ],
  "schema_version": "1.2"
}
