{
  "artifacts": [
    {
      "file": "gliformer_large_ner_s512_encoder_wfp16.tflite",
      "sha256": "714c1ec99657484aed7aa9128b25b96bf922b0569be5b8acdb108870a0606da8",
      "size_mb": 769.783
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-26",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-26/gliformer-large-ner__gliformer_large_ner_s512_encoder_wfp16.json"
  },
  "conversion": {
    "command": "not published as a script (card 'Provenance, conversion and license'): litert-torch 0.9.3 fixed-shape fp32 export of the NER path of knowledgator/gliformer-large-v1 (rev d0a4e53d) re-expressed for the GPU delegate without changing its math — host-side token lookup, one-hot routing as matmul, float masks, attention at rank 4, the exact DeBERTa logarithmic relative-position buckets as projected tables, the word BiLSTM unrolled for the window, one packed output tensor; then ai-edge-quantizer 0.8.0 float16 FLOAT_CASTING on the FULLY_CONNECTED weights (the wfp16 file); the layout and page embeddings of the backbone are skipped exactly as the upstream text path skips them",
    "quantization": "float16 weight storage (ai-edge-quantizer 0.8.0 FLOAT_CASTING on the FULLY_CONNECTED weights only): 144 float16 weight tensors read through a DEQUANTIZE to float32, every activation and every other constant float32; 1,629 operators in the fp32 reference (gliformer_large_ner_s512_encoder_fp32.tflite, 1,411,108,128 B, published as the exact reference) — 1,773 operators in this file; dynamic-range INT8 was not attempted — it does not compile on the LiteRT 2.2.0 GPU delegate for this graph family (card 'Provenance, conversion and license')",
    "tool": "litert-torch",
    "tool_version": "0.9.3 (torch 2.12.1; ai-edge-litert 2.1.6; ai-edge-quantizer 0.8.0 for the float16 weight storage; gliformer 0.1.2, gliner 0.2.29, transformers 5.16.1 — requirements-lock.txt / upstream.json)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "SQUARED_DIFFERENCE"
    ],
    "coverage_ops_pct": 89.1,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 1724,
      "unmatched": 49
    },
    "partitions": 194
  },
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu_mldrift",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": 1773,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-gliformer-gate",
            "os_build": "Android 16 (SDK 36, build BP4A.251205.006)",
            "runtime": "litert",
            "runtime_version": "2.2.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "Native LiteRT 2.2.0 CompiledModel C-API runner (gpu_runner, build gliformer-round5-variable-input-probes-loaded-rss-v3; vendor libLiteRt.so sha256 97355a36cb8a... + libLiteRtClGlAccelerator.so sha256 7c63d606a48e... pinned in android/round4/provenance.json (rounds 9 / 10 rebuilt the runner source against the same libraries, android/round9/provenance.json, android/round10/provenance.json)), one process per job, fixture inputs pushed as .f32 files and the model sha256-checked on the device (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_encoder_s512_gpu/007_hash_gliformer_large_ner_s512_encoder_wfp16.stdout.log). Sources: ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s512_split.json (encoder_gpu) + ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/device_round9_encoder_s512_gpu.json, raw outputs ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_encoder_s512_gpu/output/NN.f32, runner stderr ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_encoder_s512_gpu/017_run_graph.stderr.log, own-pid logcat ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/device_encoder_s512_gpu/019_own_pid_logcat.stdout.log; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s512_split.json.",
            "logcat/stderr: 'Replacing 1773 out of 1773 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).'; compile status 0 in 4136.0 ms; fully_accelerated=true; hardware_accelerators 'GPU only (bitmask 2)'; gpu_options_toml 'precision = 2'.",
            "conditions (round 9/10 stage job 'encoder_s512_gpu', one native process for this graph only, USB powered, no cool-start rule and no GPU clock / thermal_status sample in this lane's runner): battery before 36.5 C 86 % -> after 41.1 C 85 %; runner wall 120.5 s; compile 4136.0 ms (status 0); compile peak RSS (VmHWM) 4,035,690,496 B, process peak 4,035,690,496 B, loaded RSS after the last readback 2,325,684,224 B (this graph alone; the two graphs of a split window were measured in separate processes, so these are stage values, not the application's residency).",
            "pair (round 9 s512 split gate, ~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/gpu_round9_s512_split.json; public copy litert-community/GLiFormer-Large-NER-LiteRT/validation/s26_s512_split.json): the pulled GPU encoder hidden states of this stage fed the wfp16 head on the CPU in a separate process; 20/20 fixtures' entity span sets equal the official gliformer 0.1.2 fp32 result, max score diff 2.414e-04, max |dlogit| vs official 1.484e-02, all finite; sum of the two stage medians 1280.35 ms (not a same-process end-to-end measurement).",
            "output_match = shared element rule |out - ref| <= max(1e-5, 1e-3 * |ref|) against the Mac ai-edge-litert 2.1.6 CompiledModel CPU run of the same graph on the same inputs (~/code/codex-conversions/2026-09-26/gliformer-large-v1/results/round9/encoder_s512_wfp16_cpu_outputs.npz from fixtures/round9/encoder_s512/NN.npz; the device inputs fixtures/round9/device_s512_encoder/NN.npz are array-equal to those, checked 2026-09-26), recomputed 2026-09-26 from the saved raw outputs: 17/20 fixtures pass over all 512 x 1024 hidden values (30 violating elements of 10,485,760; worst element |d| 1.625e-05 at |ref| 2.341e-03); max |d| 1.061e-04, max rel 3.021e+01. All 30 violating elements sit at valid (unmasked) token positions (30 valid / 0 padded) and at |ref| below 0.032 (an fp32 GPU vs CPU difference of about 1e-4 at near-zero references, over the 1e-5 absolute floor); the head's outputs computed from these hidden states pass the same rule (the head rows).",
            "latency scope: all input lock/write/unlock + CompiledModel run + output lock/read/unlock; no file I/O; 5 timed runs after 2 warm-ups per fixture, 20 fixtures = 100 timed runs; median 931.49 / min 568.98 / max 1271.44 ms; per-fixture medians 571.62 ms (first) -> 930.17 ms (last) in run order."
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": 931.4937235,
          "loads": true,
          "max_abs_diff": 0.00010609626770019531,
          "max_rel_diff": 30.214284896850586,
          "metrics": {
            "battery_temp_c_end": 41.1,
            "battery_temp_c_start": 36.5,
            "compile_peak_rss_bytes": 4035690496,
            "fixtures": 20,
            "iterations": 100,
            "latency_max_ms": 1271.435156,
            "latency_min_ms": 568.975365,
            "load_compile_ms": 4135.989166,
            "loaded_rss_bytes": 2325684224,
            "pair_exact_span_sets_vs_official": 20,
            "pair_max_abs_dlogit_vs_official": 0.014841079711914062,
            "pair_max_score_diff_vs_official": 0.00024139881134033203,
            "pair_sum_of_stage_medians_ms": 1280.352161,
            "peak_rss_bytes": 4035690496,
            "per_fixture_median_first_ms": 571.618854,
            "per_fixture_median_last_ms": 930.171667,
            "repetitions_per_fixture": 5,
            "shared_rule_pass_fixtures": 17,
            "shared_rule_violating_elements": 30,
            "shared_rule_violating_elements_valid_positions": 30,
            "warmups_per_fixture": 2
          },
          "output_match": false,
          "peak_mem_mb": 4035.69,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "signature": "GpuOptions precision=FP32",
          "total_ops": 1773,
          "ttft_ms": null
        },
        "source": "data/device_runs/2.2.0/2026-09-26/gliformer-large-ner__gliformer_large_ner_s512_encoder_wfp16__galaxy-s26.json"
      }
    ]
  },
  "model": {
    "family": "gliformer-large-ner",
    "id": "gliformer-large-ner__gliformer_large_ner_s512_encoder_wfp16",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/GLiFormer-Large-NER-LiteRT",
    "task": "token-classification"
  },
  "pitfalls": [
    "Weight storage and GPU computation precision are separate settings: with GpuOptions omitted (the runtime's default GPU precision) the s128 wfp16 graph compiled on the Galaxy S26 and returned finite logits, but no entities were extracted — the Kotlin block's comment reads 'default precision returns no entities' and HOST_CONTRACT.md 'Numerical and memory limits' says 'Explicit GPU FP32 computation is required; default GPU precision failed entity agreement'. Set GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32) on every GPU graph (card 'Kotlin'; HOST_CONTRACT.md).",
    "The graphs cover the encoder and the NER head only: the host looks up the word embeddings before the graph (host_assets/word_embeddings_fp32.bin, headerless little-endian float32 [128008,1024], 524,320,768 B, the default; or word_embeddings_fp16.bin, 262,160,384 B, upcast to float32 on lookup — gated at s128 only, longer-window fp16-table accuracy was not gated) and decodes the [1,1,T,15] start/end/inside logits with the upstream pairing decoder after it (threshold 0.5, flat spans). Build the encoded sequence with the pinned gliformer 0.1.2 processor ([SCHEMA] parent token, the five [ENTITY] label pairs, the separator, the text words; right-pad ids with 0 to N): a tokenizer call over a hand-concatenated prompt is not equivalent (HOST_CONTRACT.md 'Prompt, embeddings and routing').",
    "Bind the input buffers by name: the encoder's native buffer order is inputs_embeds [1,512,1024], attention_mask [1,512] (the flatbuffer signature map lists attention_mask first); attention_mask is 1.0 for real encoded tokens and 0.0 for padding; the sole output output_0 [1,1,512,1024] is the hidden state after the embedding LayerNorm and all 24 DeBERTa layers (absolute position embeddings disabled, exact relative-position bucket indices baked in) and passes unchanged to the head graph's encoder_hidden — the host does not gather (HOST_CONTRACT.md 's512', 'Prompt, embeddings and routing'; graph_contract_s512.json).",
    "Fixed windows and a fixed class axis: N = 128 (T = 48 text words) / 256 / 512 encoded tokens including the complete label prompt, exactly five labels in a fixed order (the measured order is person, organization, location, product, date; other five-label sets are accepted by the API but were not numerically gated), batch 1. Inputs above the s512 limits raise an error naming both capacities; nothing is truncated or chunked implicitly — chunk_by_sentences(text, max_words) is a caller-side helper whose entity offsets are local to each chunk (card 'Host contract'; HOST_CONTRACT.md 'One extraction call').",
    "s256 and s512 ship as an encoder graph plus a head graph because the single s256 graph does not compile on the LiteRT 2.2.0 GPU: the checkpoint's word-level BiLSTM is unrolled for T steps and unrolled heads of 6,286 operators or more crash the GPU compiler (the head crashes at 570 MB RSS — a runtime defect reported with reproducers, not a memory limit), while the DeBERTa encoder alone (1,773 operators) compiles at every window. Create two CompiledModels — the encoder with Accelerator.GPU and FP32 precision, the head with Accelerator.CPU — and pass the encoder's [1,1,N,1024] output buffer to the head as its first input; the numbers are exact either way, only the placement differs (card 'What you get', 'Kotlin').",
    "Memory is the cost of this model (Galaxy S26, native CompiledModel processes without the token table, tokenizer or decoder): the 707 MB s128 graph is 2,591,805,440 B resident after loading and peaks at 4,542,996,480 B while the GPU delegate compiles it; the s256 encoder on the GPU with the fp32 head on the CPU in one process is 4,571,271,168 B resident; the s512 head alone is 4,676,505,600 B resident on the CPU — so treat s256 as the practical top window in an app, chunk longer documents by sentence, and plan for flagship-class phones; the token table adds 262 MB (fp16) or 524 MB (fp32) when resident (card 'What you get', memory table).",
    "Cold start: the first call after process start took 5,023 ms (compile 4,708 ms) for the s128 wfp16 graph in a fresh native process and 80 ms afterwards; the Android sample warms the whole pipeline 12 times before Ready and reported launch -> Ready 4,775 ms and a first Extract tap of 168 ms (fifth 155 ms) in a non-debuggable build with the screen on (card 'What you get', 'Kotlin', 'Measured quality and performance').",
    "Returned start / end are Unicode code-point offsets into the original Python string (end exclusive; text[start:end] is the entity text), not UTF-8 bytes and not Kotlin UTF-16 indices — translate them when supplementary characters occur (HOST_CONTRACT.md 'Output decoding').",
    "Validation scope: GPU execution was validated on the Galaxy S26 (SM-S942Q, SM8850 Adreno, Android 16, LiteRT 2.2.0) only — other Android GPU families, lower-memory phones and NPUs were not evaluated; English only, one text per call, the NER path only (classification, relation, structuring and embedding heads are not converted); the quality figures are agreement with the official gliformer 0.1.2 fp32 implementation (micro-F1 1.000, 70/70 identical span sets on desktop CPU; 10/10, 15/15, 20/20 on the phone), not a human-labelled accuracy estimate; latency is a single-device sample at battery 30-41 C with the phone idle — the s512 encoder number was taken at 41 C (card 'Measured quality and performance', 'Provenance, conversion and license')."
  ],
  "schema_version": "1.2"
}
