{
  "artifacts": [
    {
      "file": "LFM2.5-VL-3B_int4.litertlm",
      "sha256": "ebb563f3587feb0cfa6867ff53fba556043ee459410a6a0276f3245e20ba0a59",
      "size_mb": 2243.068
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "python convert_lfm25_vl3b.py ../src_models/lfm25-vl-3b out_vl3b_fp --fp && python quantize_vl.py out_vl3b_fp/model.litertlm LFM2.5-VL-3B_int4.litertlm",
    "quantization": "text int4 blockwise-32 OCTAV linears + int8 embedding/lm_head; vision tower int8 dynamic (wi4b32_wi8 + embedder wi8 pass)",
    "tool": "litert-torch export_hf --task image_text_to_text (via litertlm-convert lfm25vl_work/convert_lfm25_vl3b.py + quantize_vl.py; public mirror hf-to-litertlm lfm_work/convert_lfm25_vl.py + quantize_lfm25_vl.py)",
    "tool_version": "0.9.3 (.venv-vl093: transformers 5.14.1, torch 2.12.1, torchvision 0.27.1)"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-24",
          "decode_tokens_per_s": 27.04,
          "delegated_ops": 937,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=471.56 decode=27.04 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 161.0,
            "init_s": 10.11537,
            "prefill_tokens": 205.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 471.56,
          "provenance": "measured",
          "runs": true,
          "total_ops": 937,
          "ttft_ms": 470.0
        },
        "source": "data/device_runs/0.16.0/2026-08-24/lfm25-vl-3b-int4__galaxy-s26.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-08-13",
          "decode_tokens_per_s": 5.77,
          "delegated_ops": null,
          "env": {
            "device": "pixel-8a",
            "machine_label": "pixel-8a-cl-pinned",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Google Tensor G3",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "results block: prefill=14.12 decode=5.77 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 1109.0,
            "init_s": 44.67743,
            "prefill_tokens": 292.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 14.12,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 20850.0
        },
        "source": "data/device_runs/0.16.0/2026-08-13/lfm25-vl-3b-int4__pixel-8a.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-13",
          "decode_tokens_per_s": 10.79,
          "delegated_ops": 937,
          "env": {
            "device": "pixel-8a",
            "machine_label": "pixel-8a-cl-pinned",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Google Tensor G3",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 937 out of 937 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=91.63 decode=10.79 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 1168.0,
            "init_s": 39.1153,
            "prefill_tokens": 292.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 91.63,
          "provenance": "measured",
          "runs": true,
          "total_ops": 937,
          "ttft_ms": 3280.0
        },
        "source": "data/device_runs/0.16.0/2026-08-13/lfm25-vl-3b-int4__pixel-8a.json"
      }
    ]
  },
  "model": {
    "family": "lfm2_vl",
    "id": "lfm25-vl-3b-int4",
    "license": "lfm-open-license-v1.0",
    "source_url": "https://huggingface.co/LiquidAI/LFM2.5-VL-3B",
    "task": "image-text-to-text"
  },
  "pitfalls": [
    "Pin transformers==5.14.1 for the export: 5.15.0 renamed Lfm2ShortConv.L_cache -> conv_kernel_size and litert-torch 0.9.3's subclass still reads the old attribute (AttributeError at model load). torchvision must stay 0.27.x to keep torch under litert-torch's <2.13 pin (REPRODUCE.md).",
    "The metadata pbtext needs llm_model_type { lfm2 {} } — the empty message is correct (runtime Lfm2DataProcessor proto defaults equal this family's processor config). The chat template must render a bare <image> per image part and never boi/eoi: the runtime splits on the marker and inserts <|image_start|> + pixels + <|image_end|> itself (REPRODUCE.md).",
    "The externalized VLM embedder dodges every text quantization recipe: an --fp text export leaves it float32 (1 GB at the 128k vocab) and a post-hoc recipe applied to prefill_decode never reaches it — quantize it int8 as a separate pass (1049 -> 264 MB; REPRODUCE.md / quantize_vl.py).",
    "VLM bundles have >1 tflite section: single-tflite post-processing tools (quantize_litertlm.py, fix_zero_block_scales.py main()) rebuild via litert-lm-builder with one tflite and silently DROP the vision sections — post-process through litert-lm pack/unpack instead (RESULTS.md trap).",
    "The zero-scale int4 wall is inherited from the dense LFM2.5 lineage: this 3B checkpoint carries 746,432 all-zero 32-blocks across 26 tensors, so the zero-scale fix is mandatory or XNNPACK refuses the file (REPRODUCE.md; same signature as LFM2.5-2.6B).",
    "iPhone GPU is blocked for the ShortConv family at engine creation (upstream LiteRT-LM#3129, re-verified on 0.16.0) — iOS runs CPU; Android OpenCL and macOS GPU delegate fully (937/937 nodes, HF card).",
    "On a nearly-full Android device the GPU path's multi-GB compile-cache write can fail mid-initialization and masquerade as a kernel failure — run with --disable_cache=true (bare --disable_cache does not take) or free storage (RESULTS.md trap, hit on Pixel 8a)."
  ],
  "schema_version": "1.2"
}
