{
  "artifacts": [
    {
      "file": "gliformer_large_ner_s512_head_wfp16.tflite",
      "sha256": "7488194439006fae6433ffd4be090ca580905e179711cfb67356ac4c059183a3",
      "size_mb": 53.874
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-26/gliformer-large-ner__gliformer_large_ner_s512_head_wfp16.json"
  },
  "conversion": {
    "command": "not published as a script (card 'Provenance, conversion and license'): litert-torch 0.9.3 fixed-shape fp32 export of the NER path of knowledgator/gliformer-large-v1 (rev d0a4e53d) re-expressed for the GPU delegate without changing its math — host-side token lookup, one-hot routing as matmul, float masks, attention at rank 4, the exact DeBERTa logarithmic relative-position buckets as projected tables, the word BiLSTM unrolled for the window, one packed output tensor; then ai-edge-quantizer 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights (the wfp16 file); the layout and page embeddings of the backbone are skipped exactly as the upstream text path skips them",
    "quantization": "float16 weight storage (ai-edge-quantizer 0.8.0 FLOAT_CASTING on the FULLY_CONNECTED weights only): 9 float16 weight tensors read through a DEQUANTIZE to float32, every activation and every other constant float32; 25,103 operators in the fp32 reference (gliformer_large_ner_s512_head_fp32.tflite, 107,042,396 B, published as the exact reference) — 25,112 operators in this file; dynamic-range INT8 was not attempted — it does not compile on the LiteRT 2.2.0 GPU delegate for this graph family (card 'Provenance, conversion and license')",
    "tool": "litert-torch",
    "tool_version": "0.9.3 (torch 2.12.1; ai-edge-litert 2.1.6; ai-edge-quantizer 0.8.0 for the float16 weight storage; gliformer 0.1.2, gliner 0.2.29, transformers 5.16.1 — requirements-lock.txt / upstream.json)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "LOGISTIC",
      "TANH"
    ],
    "coverage_ops_pct": 79.6,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 19994,
      "unmatched": 5118
    },
    "partitions": 5128
  },
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu_xnnpack",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 25109,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round5-variable-input-probes-loaded-rss-v3; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s512_cpu/007_hash_gliformer_large_ner_s512_head_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s512_split.json (head_cpu) + ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/device_round9_head_s512_cpu.json, raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s512_cpu/output/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s512_cpu/017_run_graph.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_head_s512_cpu/019_own_pid_logcat.stdout.log; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s512_split.json. Inputs: the round-9 GPU encoder readbacks of this window plus the captured routing tensors.",
            "logcat/stderr: 'Created TensorFlow Lite XNNPACK delegate for CPU.' | 'Replacing 25109 out of 25112 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 3 partitions for subgraph 0 (main).'; compile status 0 in 3729.8 ms; hardware_accelerators 'CPU only (bitmask 1)'; native default CPU options (thread count not set by the runner); the three non-XNNPACK nodes are not named by the log.",
            "conditions (round 9/10 stage job 'head_s512_cpu', one native process for this graph only, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before 38.7 C 85 % -> after 40.0 C 85 %; runner wall 58.4 s; compile 3729.8 ms (status 0); compile peak RSS (VmHWM) 4,593,754,112 B, process peak 4,679,655,424 B, loaded RSS after the last readback 4,676,505,600 B (this graph alone; the two graphs of a split window were measured in separate processes, so these are stage values, not the application's residency).",
            "pair (round 9 s512 split gate, ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s512_split.json; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s512_split.json): the pulled GPU encoder hidden states of this stage fed the wfp16 head on the CPU in a separate process; 20/20 fixtures' entity span sets equal the official gliformer 0.1.2 fp32 result, max score diff 2.414e-04, max |dlogit| vs official 1.484e-02, all finite; sum of the two stage medians 1280.35 ms (not a same-process end-to-end measurement).",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/fixtures/round9/device_s512_head/cpu_outputs.npz, written by scripts/prepare_split_device_round9.py from the very fixtures/round9/device_s{seq}_head/NN.npz the device consumed, whose encoder_hidden arrays equal the round-9 GPU encoder readbacks byte for byte, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 20/20 fixtures pass over all 512 x 15 slots (0 violating elements of 153,600); max |d| 1.717e-05, max rel 3.633e-05.",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 5 timed runs after 2 warm-ups per fixture, 20 fixtures = 100 timed runs; median 348.86 / min 315.18 / max 404.47 ms; per-fixture medians 316.25 ms (first) -> 400.06 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": 348.8584375,
          "loads": true,
          "max_abs_diff": 1.71661376953125e-05,
          "max_rel_diff": 3.6327957786852494e-05,
          "metrics": {
            "battery_temp_c_end": 40.0,
            "battery_temp_c_start": 38.7,
            "compile_peak_rss_bytes": 4593754112,
            "fixtures": 20,
            "iterations": 100,
            "latency_max_ms": 404.467447,
            "latency_min_ms": 315.182604,
            "load_compile_ms": 3729.792655,
            "loaded_rss_bytes": 4676505600,
            "pair_exact_span_sets_vs_official": 20,
            "pair_max_abs_dlogit_vs_official": 0.014841079711914062,
            "pair_max_score_diff_vs_official": 0.00024139881134033203,
            "pair_sum_of_stage_medians_ms": 1280.352161,
            "peak_rss_bytes": 4679655424,
            "per_fixture_median_first_ms": 316.254531,
            "per_fixture_median_last_ms": 400.059791,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 20,
            "shared_rule_violating_elements": 0,
            "warmups_per_fixture": 2
          },
          "output_match": true,
          "peak_mem_mb": 4679.66,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "CPU XNNPACK (native runner, default threads)",
          "total_ops": 25112,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s512_head_wfp16__galaxy-s26.json"
      }
    ]
  },
  "model": {
    "family": "gliformer-large-ner",
    "id": "gliformer-large-ner__gliformer_large_ner_s512_head_wfp16",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/GLiFormer-Large-NER-LiteRT",
    "task": "token-classification"
  },
  "pitfalls": [
    "The head graphs run on the CPU on Android (Accelerator.CPU): unrolled BiLSTM heads of this size crash the LiteRT 2.2.0 GPU compiler (see the split pitfall), so no GPU precision setting applies to this file; the encoder that feeds it needs GpuOptions(precision = FP32) — HOST_CONTRACT.md 'Numerical and memory limits': 'Explicit GPU FP32 computation is required; default GPU precision failed entity agreement' (card 'Kotlin'; HOST_CONTRACT.md).",
    "The graphs cover the encoder and the NER head only: the host looks up the word embeddings before the graph (host_assets/word_embeddings_fp32.bin, headerless little-endian float32 [128008,1024], 524,320,768 B, the default; or word_embeddings_fp16.bin, 262,160,384 B, upcast to float32 on lookup — gated at s128 only, longer-window fp16-table accuracy was not gated) and decodes the [1,1,T,15] start/end/inside logits with the upstream pairing decoder after it (threshold 0.5, flat spans). Build the encoded sequence with the pinned gliformer 0.1.2 processor ([SCHEMA] parent token, the five [ENTITY] label pairs, the separator, the text words; right-pad ids with 0 to N): a tokenizer call over a hand-concatenated prompt is not equivalent (HOST_CONTRACT.md 'Prompt, embeddings and routing').",
    "Bind the input buffers by name: the head's native buffer order is encoder_hidden [1,1,512,1024], text_routing [1,512,512], parent_routing [1,1,512], label_routing [1,5,512], text_mask [1,512] (the flatbuffer signature map sorts the names: encoder_hidden, label_routing, parent_routing, text_mask, text_routing); the head performs the text, parent and label routing matmuls internally and includes one bidirectional LSTM layer, anchor fusion and the span scorer; output_0 [1,1,512,15] reshapes to [1,512,5,3] with start, end, inside logits on the last axis (HOST_CONTRACT.md 's512'; graph_contract_s512.json).",
    "Fixed windows and a fixed class axis: N = 128 (T = 48 text words) / 256 / 512 encoded tokens including the complete label prompt, exactly five labels in a fixed order (the measured order is person, organization, location, product, date; other five-label sets are accepted by the API but were not numerically gated), batch 1. Inputs above the s512 limits raise an error naming both capacities; nothing is truncated or chunked implicitly — chunk_by_sentences(text, max_words) is a caller-side helper whose entity offsets are local to each chunk (card 'Host contract'; HOST_CONTRACT.md 'One extraction call').",
    "s256 and s512 ship as an encoder graph plus a head graph because the single s256 graph does not compile on the LiteRT 2.2.0 GPU: the checkpoint's word-level BiLSTM is unrolled for T steps and unrolled heads of 6,286 operators or more crash the GPU compiler (the head crashes at 570 MB RSS — a runtime defect reported with reproducers, not a memory limit), while the DeBERTa encoder alone (1,773 operators) compiles at every window. Create two CompiledModels — the encoder with Accelerator.GPU and FP32 precision, the head with Accelerator.CPU — and pass the encoder's [1,1,N,1024] output buffer to the head as its first input; the numbers are exact either way, only the placement differs (card 'What you get', 'Kotlin').",
    "Memory is the cost of this model (Galaxy S26, native CompiledModel processes without the token table, tokenizer or decoder): the 707 MB s128 graph is 2,591,805,440 B resident after loading and peaks at 4,542,996,480 B while the GPU delegate compiles it; the s256 encoder on the GPU with the fp32 head on the CPU in one process is 4,571,271,168 B resident; the s512 head alone is 4,676,505,600 B resident on the CPU — so treat s256 as the practical top window in an app, chunk longer documents by sentence, and plan for flagship-class phones; the token table adds 262 MB (fp16) or 524 MB (fp32) when resident (card 'What you get', memory table).",
    "Cold start: the first call after process start took 5,023 ms (compile 4,708 ms) for the s128 wfp16 graph in a fresh native process and 80 ms afterwards; the Android sample warms the whole pipeline 12 times before Ready and reported launch -> Ready 4,775 ms and a first Extract tap of 168 ms (fifth 155 ms) in a non-debuggable build with the screen on (card 'What you get', 'Kotlin', 'Measured quality and performance').",
    "Returned start / end are Unicode code-point offsets into the original Python string (end exclusive; text[start:end] is the entity text), not UTF-8 bytes and not Kotlin UTF-16 indices — translate them when supplementary characters occur (HOST_CONTRACT.md 'Output decoding').",
    "Validation scope: GPU execution was validated on the Galaxy S26 (SM-S942Q, SM8850 Adreno, Android 16, LiteRT 2.2.0) only — other Android GPU families, lower-memory phones and NPUs were not evaluated; English only, one text per call, the NER path only (classification, relation, structuring and embedding heads are not converted); the quality figures are agreement with the official gliformer 0.1.2 fp32 implementation (micro-F1 1.000, 70/70 identical span sets on desktop CPU; 10/10, 15/15, 20/20 on the phone), not a human-labelled accuracy estimate; latency is a single-device sample at battery 30-41 C with the phone idle — the s512 encoder number was taken at 41 C (card 'Measured quality and performance', 'Provenance, conversion and license').",
    "The s512 head is the largest resident graph of the package: 4,676,505,600 B on the CPU alone (4,727,980,032 B for the fp32 head, 8 % faster at 319.5 vs 348.9 ms median), and s512 simultaneous residency with its encoder was not measured — the card's advice is to treat s256 as the practical top window in an app and chunk longer documents by sentence (card 'What you get', 'Measured quality and performance'; HOST_CONTRACT.md 'Numerical and memory limits')."
  ],
  "schema_version": "1.2"
}
