{
  "artifacts": [
    {
      "file": "SmolVLM2-500M.litertlm",
      "sha256": "0dfb6fb881eb16e5ef2b2be04de5476caf939b7d9ae601fdee308bbc5462fd55",
      "size_mb": 344.326
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "TODO (owner) — not recorded in the source card",
    "quantization": "vision: SigLIP encoder + pixel-shuffle x4 + Linear connector int8 -> 64 image tokens; decoder SmolLM2-360M int4 weights (blockwise-32 + OCTAV); tied embedding INT8 (externalized); integer compute; KV cache 2048. File SmolVLM2-500M.litertlm, ~361 MB",
    "tool": "TODO (owner) — the card describes the fast_vlm bundle layout and graph edits but does not name the converter tool",
    "tool_version": "TODO (owner) — not recorded in the source card"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": 72.41,
          "delegated_ops": null,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-vl-response-v1",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/NOTES.md — the table row for SmolVLM2-500M litert-lm-cpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_cpu.jsonl — 3 launch records; medians here are computed over the 3 passing launches",
            "session anchor right before the sitting: llama.cpp b8999 unsloth/Qwen3-0.6B-GGUF short-chat decode 108.0 / 107.0 / 110.7 tok/s, median 108.0, ratio 1.016 vs the 2026-09-24 reference (106.3, n=6) — ADMITTED under the dashboard rule (SESSION.json)",
            "screen off in every launch (provenance.deviceBefore.wakefulness = Dozing, stay_on_while_plugged_in = 0); thermal status nominal before and after all 36 launches; battery 30.9–35.0 °C; 45 s between launches, 60–90 s between cells; no CPU frequency cap before any launch (capped after 11 of them, flag only)",
            "litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_cpu_run1.stderr.log: VERBOSE: Replacing 1290 out of 1423 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 189 partitions for subgraph 0 (prefill_128). (first partition line; the bundle has several graphs — vision, prefill, decode)",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it",
            "reply (identical across the 3 launches, text check pass): 'A close-up photograph of a tabby kitten, lying on a dark fabric surface. The kitten, with its distinctive striped fur, is gently nuzzling it'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 64,
            "invalid_decode_lines": 0,
            "launches": 3,
            "load_s_median": 1.224,
            "passing_launches": 3,
            "prefill_tokens": 82,
            "response_s_max": 1.756,
            "response_s_median": 1.47,
            "response_s_min": 1.464,
            "ttft_s_median": 0.29,
            "vision_encoder_ms_median": 292.0477
          },
          "output_match": null,
          "peak_mem_mb": 875.2,
          "prefill_tokens_per_s": 294.92,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 290.0
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-26/smolvlm2-500m__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": 76.72,
          "delegated_ops": 1794,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 1290 out of 1423 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 189 partitions for subgraph 0 (prefill_128).",
            "VERBOSE: Replacing 1290 out of 1423 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 189 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 1158 out of 1292 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 194 partitions for subgraph 2 (decode).",
            "VERBOSE: Replacing 8 out of 8 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 1 partitions for subgraph 3 (odml.rms_norm.impl).",
            "results block: prefill=417.03 decode=76.72 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 212.0,
            "init_s": 0.61011,
            "prefill_tokens": 203.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 417.03,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1864,
          "ttft_ms": 500.0
        },
        "source": "data/device_runs/0.16.0/2026-09-05/smolvlm2-500m__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-26",
          "decode_tokens_per_s": 54.48,
          "delegated_ops": null,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-vl-response-v1",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/NOTES.md — the table row for SmolVLM2-500M litert-lm-gpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-26-vl-response-v1-s26-android/litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_gpu.jsonl — 3 launch records; medians here are computed over the 3 passing launches",
            "session anchor right before the sitting: llama.cpp b8999 unsloth/Qwen3-0.6B-GGUF short-chat decode 108.0 / 107.0 / 110.7 tok/s, median 108.0, ratio 1.016 vs the 2026-09-24 reference (106.3, n=6) — ADMITTED under the dashboard rule (SESSION.json)",
            "screen off in every launch (provenance.deviceBefore.wakefulness = Dozing, stay_on_while_plugged_in = 0); thermal status nominal before and after all 36 launches; battery 30.9–35.0 °C; 45 s between launches, 60–90 s between cells; no CPU frequency cap before any launch (capped after 11 of them, flag only)",
            "litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_gpu_run1.stderr.log: Dynamically loaded GPU accelerator(libLiteRtGpuAccelerator.so) registered. — accelerator that computed: LiteRT GPU (libLiteRtGpuAccelerator.so, ML Drift OpenCL; LiteRT-LM prebuilt/android_arm64 at 1dadd00c, sha256 88e716f4ff79970b377597fb4207e930dc4905bc23ddb08479cc3f317627274a)",
            "litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_gpu_run1.stderr.log: VERBOSE: Replacing 1423 out of 1423 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_128). (first partition line; the bundle has several graphs — vision, prefill, decode)",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it",
            "reply (identical across the 3 launches, text check pass): 'A close-up photograph of a tabby kitten, lying on a dark fabric surface. The kitten, with its distinctive striped fur, is lying on its side,'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 64,
            "invalid_decode_lines": 0,
            "launches": 3,
            "load_s_median": 2.4,
            "passing_launches": 3,
            "prefill_tokens": 82,
            "response_s_max": 1.563,
            "response_s_median": 1.465,
            "response_s_min": 1.447,
            "ttft_s_median": 0.09,
            "vision_encoder_ms_median": 200.0972
          },
          "output_match": null,
          "peak_mem_mb": 467.6,
          "prefill_tokens_per_s": 1115.94,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 90.0
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-26/smolvlm2-500m__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-24",
          "decode_tokens_per_s": 76.73,
          "delegated_ops": 1423,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 1423 out of 1423 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_128).",
            "VERBOSE: Replacing 1423 out of 1423 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 1292 out of 1292 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (decode).",
            "results block: prefill=1371.43 decode=76.73 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 282.0,
            "init_s": 1.26452,
            "prefill_tokens": 204.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 1371.43,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1423,
          "ttft_ms": 160.0
        },
        "source": "data/device_runs/0.16.0/2026-08-24/smolvlm2-500m__galaxy-s26.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-24",
          "decode_tokens_per_s": 95.53,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max-vl-response-v1",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-24-vl-response-v1-m4max-mac/NOTES.md — the table row for SmolVLM2-500M litert-lm-cpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-24-vl-response-v1-m4max-mac/litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_cpu.jsonl — 3 launch records; medians here are computed over the 3 passing launches",
            "session anchor: mlx-swift Qwen3-0.6B-4bit short-chat warm 559.6 tok/s, 2 % above the 527–550 band of the previous five Mac sessions (not throttled)",
            "the host was shared (Xcode builds, Chrome and Storage bursts); every launch samples foreign processes >= 20 % CPU into provenance.hostBefore/After; the Gemma 4 pair was re-taken in a quieter window (NOTES.md 'Spread')",
            "peak_mem_mb = VmHWM (Android) / ps rss (Mac) of the CLI process, host memory only — a GPU arm's device heap is outside it",
            "reply (identical across the 3 launches, text check pass): 'A close-up photograph of a tabby kitten, lying on a dark fabric surface. The kitten, with its distinctive striped fur, is gently nuzzling it'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 64,
            "invalid_decode_lines": 0,
            "launches": 3,
            "load_s_median": 0.159,
            "passing_launches": 3,
            "prefill_tokens": 82,
            "response_s_max": 0.991,
            "response_s_median": 0.975,
            "response_s_min": 0.972,
            "ttft_s_median": 0.19
          },
          "output_match": null,
          "peak_mem_mb": 831.2,
          "prefill_tokens_per_s": 464.71,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 190.0
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-24/smolvlm2-500m__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-24",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max-vl-response-v1",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-1dadd00c",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": "reply = 64 x <|endoftext|> on every launch; the engine printed rates (TTFT 0.00 s, prefill ~46,000 tok/s) but no text — text check fails, the rates are not a measurement",
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-24-vl-response-v1-m4max-mac/NOTES.md — the table row for SmolVLM2-500M litert-lm-gpu (median of 3 launches, range in brackets)",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-24-vl-response-v1-m4max-mac/litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_gpu.jsonl — 3 launch records; medians here are computed over no passing launch",
            "session anchor: mlx-swift Qwen3-0.6B-4bit short-chat warm 559.6 tok/s, 2 % above the 527–550 band of the previous five Mac sessions (not throttled)",
            "the host was shared (Xcode builds, Chrome and Storage bursts); every launch samples foreign processes >= 20 % CPU into provenance.hostBefore/After; the Gemma 4 pair was re-taken in a quieter window (NOTES.md 'Spread')",
            "litert-lm_litert-community_SmolVLM2-500M_vl-describe-catcouch-gen64_gpu_run1.stderr.log: Dynamically loaded GPU accelerator(libLiteRtWebGpuAccelerator.dylib) registered. — accelerator that computed: GPU WebGPU (libLiteRtWebGpuAccelerator.dylib, Dawn with the Metal backend; NOT libLiteRtMetalAccelerator — PROVENANCE.md erratum 2026-09-26)"
          ],
          "failure_class": "degenerate_output",
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "exit_code": 0,
            "launches": 3,
            "passing_launches": 0
          },
          "output_match": null,
          "peak_mem_mb": 254.7,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/selfbuilt-1dadd00c/2026-09-24/smolvlm2-500m__mac-studio-m4-max.json"
      }
    ]
  },
  "model": {
    "family": "smolvlm",
    "id": "smolvlm2-500m",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/HuggingFaceTB/SmolVLM2-500M-Video-Instruct",
    "task": "image-text-to-text"
  },
  "pitfalls": [
    "One image per conversation on GPU: a second image in the same conversation may degrade — a GPU-delegate trait shared across fast_vlm models. CPU handles multi-image; start a new conversation for a different image.",
    "Gallery import: in the Import Model dialog you must check \"Support image\" or image input will not work; Gallery v1.0.16+ can also import directly from Hugging Face inside the app.",
    "Vision-only bundle, no audio tower — on the Swift runtime load with the vision tower enabled (Modality.textImage / [.vision]).",
    "Very small (500M) model: keep a sensible max_tokens and use sampling (e.g. top-p); at pure greedy it can be repetitive/verbose.",
    "Image input is resized to 512x512, with the (x-0.5)/0.5 normalization and NCHW transpose baked into the vision-encoder graph.",
    "The vision encoder uses the static arange(1024) position-embedding path — the model's dynamic bucketize position logic is bypassed, numerically identical only for a full 512x512 frame. Single-image, no high-res splitting: a fixed 64 soft tokens.",
    "The SigLIP vision tower converts bit-faithfully (float CPU-parity corr ~ 1.0); single-image VQA is CPU-verified as coherent and image-grounded.",
    "This is the image path of SmolVLM2-500M-Video-Instruct — the package is image+text; video is not the converted path."
  ],
  "schema_version": "1.2"
}
