{
  "artifacts": [
    {
      "file": "OvisOCR2_int8.litertlm",
      "sha256": "6482ff193826b70093a9b86f559669b1c12c13ab170f6dfc0f4eb6037a88887d",
      "size_mb": 1241.934
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "MODEL=ATH-MaaS/OvisOCR2 IMG=512 convert_qwen35_vision.py out/ovisocr2-vision; decoder export on the patched worktree with PREFILL=1024,256,64,16,4,1 CACHE=4096 (export_qwen35vl_decoder.py per the 0.8B rail); bundle -> out/ovisocr2-bundle/OvisOCR2_int8_final.litertlm with ExecutorMetadata retrofitted (FINDINGS.md); step-by-step = hf-to-litertlm qwen35vl_work/ (HF card)",
    "quantization": "decoder int8 (dynamic on linears + embedding; convs and the delta rule stay float, fp32 activations declared) + fp16 vision encoder / int8 vision adapter, static 512x512, six-signature prefill ladder",
    "tool": "litert-torch + qwen35 hybrid-cache patch (decoder, patched worktree), convert_qwen35_vision.py (vision), bundle + ExecutorMetadata retrofit; public repro = hf-to-litertlm qwen35vl_work/ REPRODUCE section with the env-override scripts (HF card; FINDINGS.md)",
    "tool_version": "litert-torch 0.9.3 + litert-converter 0.4.0 (~/venvs/lt093ctl — FINDINGS.md); decoder from the qwen35_work patched worktree per the inherited 0.8B rail (qwen35vl_work/FINDINGS.md: 115a136 + qwen35_hybrid_litert_torch.patch)"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": null,
          "delegated_ops": 24542,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": "generation gate FAIL: exit 0 but no answer text (journal generated=false); output tail …'France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of'",
          "evidence": [
            "VERBOSE: Replacing 24542 out of 25703 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2245 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 17648 out of 18791 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 1 (prefill_256).",
            "VERBOSE: Replacing 15632 out of 16775 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 2 (prefill_64).",
            "VERBOSE: Replacing 15722 out of 16865 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 3 (prefill_16).",
            "community_accel_work/s7_manifest_backfill/logs/OvisOCR2-LiteRT__OvisOCR2_int8.gatecpu.log: generation gate verdict FAIL (journal rows.jsonl, fail_class=None, exit=0, generated=False, timeout=False); the same journal's --benchmark numbers for this leg are withheld because the gate output is not a real answer",
            "protocol: each bench run cold (caches deleted between runs), device cooled <42 C before every run; runtime litert_lm_advanced_main v0.16.0 (S4 kit)",
            "gate process peak 1784 MB (journal gate_peak_mb)"
          ],
          "failure_class": "degenerate_output",
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": 25703,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.0/2026-09-05/ovisocr2-int8__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": 35.18,
          "delegated_ops": 25703,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 25703 out of 25703 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 18791 out of 18791 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_256).",
            "VERBOSE: Replacing 16775 out of 16775 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_64).",
            "VERBOSE: Replacing 16865 out of 16865 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_16).",
            "results block: prefill=537.1 decode=35.18 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 195.0,
            "init_s": 81.31138,
            "prefill_tokens": 207.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 537.1,
          "provenance": "measured",
          "runs": true,
          "total_ops": 25703,
          "ttft_ms": 410.0
        },
        "source": "data/device_runs/0.16.0/2026-09-05/ovisocr2-int8__galaxy-s26.json"
      },
      {
        "device": "iphone-17-pro",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-01",
          "decode_tokens_per_s": 18.05,
          "delegated_ops": null,
          "env": {
            "device": "iPhone 17 Pro",
            "machine_label": "iphone-17-pro-thermal-serious",
            "os_build": "iOS 27.0",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "A19-Pro",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "device.modelIdentifier=iPhone18,1",
            "throughput taken from task 'quality' (127 generated tokens, 2026-09-01T07:08:30Z); peak_mem_mb = max memoryPeakDuringDecodeMB over 1 task file(s)",
            "[quality 2026-09-01T07:08:30Z] prefill=197.18224683186213 decode=18.048044107432524 ttft_ms=1022.0 peak_mb=914.988395690918 stop=stop output='1. What is 17 + 25?\\\\n\\\\n2. What is the capital of Japan?\\ …[trace truncated]"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "quality.decode_tokens_per_s": 18.048044,
            "quality.generated_tokens": 127.0,
            "quality.load_s": 9.858049,
            "quality.peak_mem_mb": 914.988396,
            "quality.prefill_tokens_per_s": 197.182247,
            "quality.prompt_tokens": 138.0,
            "quality.ttft_ms": 1022.0
          },
          "output_match": null,
          "peak_mem_mb": 914.988,
          "prefill_tokens_per_s": 197.18,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 1022.0
        },
        "source": "data/device_runs/0.16.0/2026-09-01/ovisocr2-int8__iphone-17-pro.json"
      },
      {
        "device": "iphone-17-pro",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-01",
          "decode_tokens_per_s": 48.5,
          "delegated_ops": null,
          "env": {
            "device": "iPhone 17 Pro",
            "machine_label": "iphone-17-pro-cold-unplugged",
            "os_build": "iOS 27.0",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "A19-Pro",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "device.modelIdentifier=iPhone18,1",
            "throughput taken from task 'quality' (106 generated tokens, 2026-09-01T07:06:58Z); peak_mem_mb = max memoryPeakDuringDecodeMB over 3 task file(s)",
            "[quality 2026-09-01T07:06:58Z] prefill=602.9798345847167 decode=48.49926331998226 ttft_ms=1513.0 peak_mb=3828.3493041992188 stop=stop output='1. What is 17 + 25?\\\\n\\\\n2. What is the capital of Japan?\\ …[trace truncated]",
            "[vision-has-text 2026-09-01T07:07:33Z] prefill=0.0 decode=64.6893528256163 ttft_ms=1467.0 peak_mb=3832.676902770996 stop=length output='## SAFETY PRO'",
            "[vision-no-text 2026-09-01T07:08:09Z] prefill=0.0 decode=62.3199497791695 ttft_ms=1763.0 peak_mb=3982.6925506591797 stop=length output='<img src=\"images'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "quality.decode_tokens_per_s": 48.499263,
            "quality.generated_tokens": 106.0,
            "quality.load_s": 31.166061,
            "quality.peak_mem_mb": 3828.349304,
            "quality.prefill_tokens_per_s": 602.979835,
            "quality.prompt_tokens": 138.0,
            "quality.ttft_ms": 1513.0,
            "vision_has_text.decode_tokens_per_s": 64.689353,
            "vision_has_text.generated_tokens": 4.0,
            "vision_has_text.load_s": 30.373604,
            "vision_has_text.peak_mem_mb": 3832.676903,
            "vision_has_text.prefill_tokens_per_s": 0.0,
            "vision_has_text.prompt_tokens": 0.0,
            "vision_has_text.ttft_ms": 1467.0,
            "vision_no_text.decode_tokens_per_s": 62.31995,
            "vision_no_text.generated_tokens": 4.0,
            "vision_no_text.load_s": 32.210293,
            "vision_no_text.peak_mem_mb": 3982.692551,
            "vision_no_text.prefill_tokens_per_s": 0.0,
            "vision_no_text.prompt_tokens": 0.0,
            "vision_no_text.ttft_ms": 1763.0
          },
          "output_match": null,
          "peak_mem_mb": 3982.693,
          "prefill_tokens_per_s": 602.98,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 1513.0
        },
        "source": "data/device_runs/0.16.0/2026-09-01/ovisocr2-int8__iphone-17-pro.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-01",
          "decode_tokens_per_s": 48.93,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": null,
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "results block: prefill=659.46 decode=48.93 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 256.0,
            "prefill_tokens": 256.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 659.46,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 429.0
        },
        "source": "data/device_runs/0.16.0/2026-09-01/ovisocr2-int8__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-01",
          "decode_tokens_per_s": 140.9,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": null,
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "results block: prefill=2174.87 decode=140.9 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 256.0,
            "prefill_tokens": 256.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 2174.87,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 133.7
        },
        "source": "data/device_runs/0.16.0/2026-09-01/ovisocr2-int8__mac-studio-m4-max.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": null,
          "delegated_ops": 24542,
          "env": {
            "device": "Pixel 8a",
            "machine_label": "pixel-8a-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Tensor G3",
            "vendor_sdk": null
          },
          "error": "generation gate FAIL: exit 0 but no answer text (journal generated=false); output tail …'France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of France? The capital of'",
          "evidence": [
            "VERBOSE: Replacing 24542 out of 25703 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2245 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 17648 out of 18791 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 1 (prefill_256).",
            "VERBOSE: Replacing 15632 out of 16775 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 2 (prefill_64).",
            "VERBOSE: Replacing 15722 out of 16865 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 2209 partitions for subgraph 3 (prefill_16).",
            "community_accel_work/s7_manifest_backfill/logs_p8a/OvisOCR2-LiteRT__OvisOCR2_int8.gatecpu.log: generation gate verdict FAIL (journal rows_p8a.jsonl, fail_class=None, exit=0, generated=False, timeout=False); the same journal's --benchmark numbers for this leg are withheld because the gate output is not a real answer",
            "protocol: each bench run cold (caches deleted between runs), device cooled <42 C before every run; runtime litert_lm_advanced_main v0.16.0 (S4 kit)",
            "gate process peak 1723 MB (journal gate_peak_mb)"
          ],
          "failure_class": "degenerate_output",
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": 25703,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.0/2026-09-05/ovisocr2-int8__pixel-8a.json"
      }
    ]
  },
  "model": {
    "family": "qwen3_5",
    "id": "ovisocr2-int8",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/mlboydaisuke/OvisOCR2-LiteRT",
    "task": "image-text-to-text"
  },
  "pitfalls": [
    "The OCR post-training erodes general instruction following, so the chat-style gates are not valid gates for this checkpoint: the 8-question text gate reads Mac GPU 4/8 (1 degenerate) / CPU 3/8, but every miss is the model transcribing or echoing the prompt ('17 + 25? / Answer briefly.' echoed back, a 'quick brown fox' non sequitur, numbered-list loops) — not the wrong-token face of quant collapse. The fp32 original on the stock HF stack fails the same questions the same way, question by question (echoes '17 + 25', loops 'Tokyo', loops a numbered list on the French question, and leaks literal think scaffolds into its own output), so the conversion is exonerated per-question, not just by score (FINDINGS.md Mac gates + HF fp32 anchors; HF card 'Scope').",
    "BANANA hermetic pad sweep 12-51 reads 1/40 and is INVALID for this checkpoint: every 'failure' is the same uniform echo of the filler at every fill count (coherent, not garbage) because the instrument assumes instruction-following ('reply BANANA'), and the fp32 original echoes the filler too. Pad-guard assurance rests instead on the decoder being structurally identical to the Qwen3.5-0.8B VL decoder that passed 40/40 (pad handling is architecture, not weights) and on the OCR E2E runs exercising ~110-token multi-chunk prefill with byte-identical CPU/GPU output (FINDINGS.md).",
    "Transcription-mode echo on the quality task: the on-device 'quality' legs (iPhone 17 Pro Metal and CPU) produce transcription-mode echo of the prompt — checkpoint behavior, same as Mac — so those rows measure echo output, not answers. This is a document parser, not a chat model: feed it document pages with the upstream reference transcription prompt (verbatim on the card); on-device image legs are grounded (a poster's title transcribed; a photo without text gets the picture-region img tag, exactly its document convention) (FINDINGS.md iPhone gate; HF card 'Scope' + Usage).",
    "No start token; stop tokens are <|endoftext|> (248044) and <|im_end|> (248046). The upstream repo has no generation_config.json, so both ids were taken from config.json + the tokenizer and verified against the tokenizer at bundle build. Consequence for anyone comparing against stock transformers: HF's eos is only <|endoftext|>, so fp32 generation runs through the turn end and rambles (literal </think>, hallucinated bbox tags, formula repeats) — add 248046 to eos_token_id on the HF side before comparing; the bundle stops cleanly (HF card conversion notes + 'What it does'; FINDINGS.md HF fp32 anchors).",
    "Requires litert-lm >= 0.15 (HF card header; manifest: hybrid state binds via ExecutorMetadata). litert-torch 0.9.3 writes no ExecutorMetadata section, so it is retrofitted at bundle time (the standing repair) — 48 state buffers (36 linear-attn + 12 kv), prefer_activation_type = fp32 declared (FINDINGS.md Conversion).",
    "Android/Mali is pending: no Android gate has been run for this build. The northmv law says fp16 vision reboots Mali, so the Android path would be the un-shipped int8-vision variant, whose parts are kept locally (int8/int8 document-fixture corr 0.994-0.998) — not measured on a phone (FINDINGS.md post-ship housekeeping).",
    "Vision input is static 512x512 and the runtime resizes the attachment to it (pre-render pages at 512x512 to control the aspect ratio). Dense academic pages (9 pt text) exceed what 512x512 can resolve: the opening is transcribed correctly, then the model invents and eventually repeats. The HF fp32 original does the same on the same input — every variant loops on that page, and the bundle matches the 1-D-position fp32 reference byte-for-byte through the entire legible region (first 522 chars), diverging only inside the invented continuation. Upstream's own reference code ships a repeat-cleanup post-processor for exactly this. Use pages legible at 512x512, or tile (HF card 'What it does'; FINDINGS.md HF fp32 anchors).",
    "Rail inheritance has two per-checkpoint costs. (a) The fp16-safe LayerNorm scales are calibrated per checkpoint — OvisOCR2's have the same shape as base 0.8B (block 6+ at 16, final 512) but absmax moved 2390.9 -> 2803.9, so recalibration is not optional. (b) Positions are 1-D under fast_vlm: the formula page matches the 1-D-position fp32 reference at similarity 1.0000 but the full M-RoPE reference at 0.9968 — a one-token $B$ vs B difference, the position-contract cost (HF card conversion notes + 'What it does'; FINDINGS.md)."
  ],
  "schema_version": "1.2"
}
