{
  "artifacts": [
    {
      "file": "InternVL3-1B.litertlm",
      "sha256": "7cf87c35cf364d04bd2e6f957e3a0366830ad0b46eca3dd81e0000891fd1284a",
      "size_mb": 703.158
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "TODO (the HF card states the bundle layout and rewrites, not the invocation)",
    "quantization": "TODO — the HF card states no quantization recipe for this bundle; sections: VISION_ENCODER [1,448,448,3]->[1,256,4096], VISION_ADAPTER [1,256,4096]->[1,256,896], externalized embedder, PREFILL_DECODE (HF card Conversion notes)",
    "tool": "litert-torch (LiteRT-LM fast_vlm bundle: VISION_ENCODER + VISION_ADAPTER + single-token EMBEDDER + PREFILL_DECODE with embeddings input) (HF card Conversion notes)",
    "tool_version": "TODO (the HF card names the converter but not its version; no export log in the sources)"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": 81.7,
          "delegated_ops": 1337,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 961 out of 1063 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 141 partitions for subgraph 0 (prefill_128).",
            "VERBOSE: Replacing 961 out of 1063 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 141 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 869 out of 972 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 146 partitions for subgraph 2 (decode).",
            "VERBOSE: Replacing 8 out of 8 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 3 (odml.rms_norm.impl).",
            "results block: prefill=358.66 decode=81.7 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 205.0,
            "init_s": 0.95637,
            "prefill_tokens": 239.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 358.66,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1392,
          "ttft_ms": 680.0
        },
        "source": "data/device_runs/0.16.0/2026-09-05/internvl3-1b__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": 68.14,
          "delegated_ops": null,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-vl-response-v1",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/NOTES.md — the table row for InternVL3-1B litert-lm-cpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_cpu.jsonl — 3 launch records; medians here are computed over the 3 passing launches",
            "session anchor right before the sitting: llama.cpp b8999 unsloth/Qwen3-0.6B-GGUF short-chat decode 108.0 / 107.0 / 110.7 tok/s, median 108.0, ratio 1.016 vs the 2026-09-24 reference (106.3, n=6) — ADMITTED under the dashboard rule (SESSION.json)",
            "screen off in every launch (provenance.deviceBefore.wakefulness = Dozing, stay_on_while_plugged_in = 0); thermal status nominal before and after all 36 launches; battery 30.9–35.0 °C; 45 s between launches, 60–90 s between cells; no CPU frequency cap before any launch (capped after 11 of them, flag only)",
            "litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_cpu_run1.stderr.log: VERBOSE: Replacing 961 out of 1063 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 141 partitions for subgraph 0 (prefill_128). (first partition line; the bundle has several graphs — vision, prefill, decode)",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it",
            "reply (identical across the 3 launches, text check pass): 'The image shows a close-up of a tabby cat lying on a dark fabric surface.'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 19,
            "invalid_decode_lines": 0,
            "launches": 3,
            "load_s_median": 0.621,
            "passing_launches": 3,
            "prefill_tokens": 311,
            "response_s_max": 1.515,
            "response_s_median": 1.49,
            "response_s_min": 1.476,
            "ttft_s_median": 0.64,
            "vision_encoder_ms_median": 557.6103
          },
          "output_match": null,
          "peak_mem_mb": 1002.8,
          "prefill_tokens_per_s": 495.39,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 640.0
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-26/internvl3-1b__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-vl-response-v1",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": "BROADCAST_TO: Operation is not supported.",
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/NOTES.md — the table row for InternVL3-1B litert-lm-gpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_gpu.jsonl — 3 launch records; medians here are computed over no passing launch",
            "session anchor right before the sitting: llama.cpp b8999 unsloth/Qwen3-0.6B-GGUF short-chat decode 108.0 / 107.0 / 110.7 tok/s, median 108.0, ratio 1.016 vs the 2026-09-24 reference (106.3, n=6) — ADMITTED under the dashboard rule (SESSION.json)",
            "screen off in every launch (provenance.deviceBefore.wakefulness = Dozing, stay_on_while_plugged_in = 0); thermal status nominal before and after all 36 launches; battery 30.9–35.0 °C; 45 s between launches, 60–90 s between cells; no CPU frequency cap before any launch (capped after 11 of them, flag only)",
            "litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_gpu_run1.stderr.log: Dynamically loaded GPU accelerator(libLiteRtGpuAccelerator.so) registered. — accelerator that computed: LiteRT GPU (libLiteRtGpuAccelerator.so, ML Drift OpenCL; LiteRT-LM prebuilt/android_arm64 at 1dadd00c, sha256 88e716f4ff79970b377597fb4207e930dc4905bc23ddb08479cc3f317627274a)",
            "litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_gpu_run1.stderr.log: VERBOSE: Replacing 1063 out of 1063 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_128). (first partition line; the bundle has several graphs — vision, prefill, decode)",
            "BROADCAST_TO: Operation is not supported.",
            "GATHER_ND: Operation is not supported.",
            "└ Some ops are not accelerated. Add kLiteRtHwAcceleratorCpu to the compilation accelerator set to allow using the CPU to run those.",
            "exit 13 on all 3 launches (FAILURES.txt); the vision executor compiles with kGpu only, so the engine is never created"
          ],
          "failure_class": "engine_create_failed",
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": false,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "launches": 3,
            "passing_launches": 0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-26/internvl3-1b__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-24",
          "decode_tokens_per_s": 85.42,
          "delegated_ops": 1063,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 1063 out of 1063 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_128).",
            "VERBOSE: Replacing 1063 out of 1063 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 972 out of 972 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (decode).",
            "results block: prefill=1455.95 decode=85.42 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 115.0,
            "init_s": 1.38797,
            "prefill_tokens": 203.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 1455.95,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1063,
          "ttft_ms": 150.0
        },
        "source": "data/device_runs/0.16.0/2026-08-24/internvl3-1b__galaxy-s26.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-10-01",
          "decode_tokens_per_s": 78.7,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max-vl-response-v1",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-10-01-vl-response-v1-m4max-mac/NOTES.md — the table row for InternVL3-1B litert-lm-cpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-10-01-vl-response-v1-m4max-mac/litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_cpu.jsonl — 3 launch records; medians here are computed over the 3 passing launches",
            "session anchor first in the sitting: mlx-swift Qwen3-0.6B-4bit short-chat cold 558.7, warm 544.4 / 560.9 tok/s (warm median 552.6, 0.988 of the 2026-09-24 VL leg's anchor; not throttled); thermal nominal and no CPU speed limit at every launch",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it",
            "reply (identical across the 3 launches, text check pass): 'The image shows a close-up of a tabby cat lying on a dark fabric surface.'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 19,
            "launches": 3,
            "load_s_median": 0.176,
            "passing_launches": 3,
            "prefill_tokens": 311,
            "response_s_max": 1.025,
            "response_s_median": 1.0,
            "response_s_min": 0.991,
            "ttft_s_median": 0.44,
            "vision_encoder_ms_median": 328.1061
          },
          "output_match": null,
          "peak_mem_mb": 1035.9,
          "prefill_tokens_per_s": 736.23,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 440.0
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-10-01/internvl3-1b__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-10-01",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max-vl-response-v1",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": "BROADCAST_TO: Operation is not supported.",
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-10-01-vl-response-v1-m4max-mac/NOTES.md — the table row for InternVL3-1B litert-lm-gpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-10-01-vl-response-v1-m4max-mac/litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_gpu.jsonl — 3 launch records; no passing launch",
            "session anchor first in the sitting: mlx-swift Qwen3-0.6B-4bit short-chat cold 558.7, warm 544.4 / 560.9 tok/s (warm median 552.6, 0.988 of the 2026-09-24 VL leg's anchor; not throttled); thermal nominal and no CPU speed limit at every launch",
            "litert-lm_litert-community_InternVL3-1B_vl-describe-catcouch-gen64_gpu_run1.stderr.log: Dynamically loaded GPU accelerator(libLiteRtWebGpuAccelerator.dylib) registered. — accelerator that computed: GPU WebGPU (libLiteRtWebGpuAccelerator.dylib, Dawn with the Metal backend; GPU Metal registered in none of the launch logs)",
            "ERROR: Following operations are not supported by GPU delegate:",
            "BROADCAST_TO: Operation is not supported.",
            "GATHER_ND: Operation is not supported.",
            "1307 operations will run on the GPU, and the remaining 13 operations will run on the CPU.",
            "└ Some ops are not accelerated. Add kLiteRtHwAcceleratorCpu to the compilation accelerator set to allow using the CPU to run those.",
            "exit 13 on all 3 launches (FAILURES.txt); the same two ops and the same 1307 / 13 split as the Galaxy S26 OpenCL leg of 2026-09-26; each launch log also carries one Dawn validation error 'Buffer size (67954944) wrapping host-mapped memory was not aligned to 4096.' before the vision compile",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it"
          ],
          "failure_class": "engine_create_failed",
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": false,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "launches": 3,
            "passing_launches": 0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-10-01/internvl3-1b__mac-studio-m4-max.json"
      }
    ]
  },
  "model": {
    "family": "internvl3",
    "id": "internvl3-1b",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/InternVL3-1B",
    "task": "image-text-to-text"
  },
  "pitfalls": [
    "Known limitation — one image per conversation on the GPU backend: on the GPU (Metal) backend a second image in the same conversation truncates the answer; on the CPU backend multi-image works (verified); the truncation reproduces with other fast_vlm models, so it is the runtime's GPU fast_vlm path, not this bundle. For reliable multi-image, run on the CPU backend (HF card Known limitation).",
    "The vision encoder bakes ImageNet normalization + the NCHW transpose into the graph (the runtime feeds a [0,1] NHWC image), and the InternViT attention is rewritten 4D-clean (qkv split before the head reshape — no GPU-rejected 5D reshape), numerically identical (corr ≈ 1.0) (HF card Conversion notes).",
    "Decoder exported with an externalized embedder; InternVL's dynamic-NTK rope_scaling is stripped to base RoPE, valid since the export cache <= the base context window (HF card Conversion notes).",
    "License: MIT (the InternVL model) + Apache-2.0 (the Qwen2.5 language component); converted artifacts released under the same terms (HF card License).",
    "LiteRT-LM .litertlm bundle: LiteRT.js cannot run it, so delegation stays null and there is no browser block; device rows come from data/device_runs/ (Pi 5 / S26 / Pixel 8a / Mac / iPhone as measured)."
  ],
  "schema_version": "1.2"
}
