{
  "artifacts": [
    {
      "file": "LFM2.5-1.2B-Instruct_int4.litertlm",
      "sha256": "a28b5c59ac204e2e51c1f98d2d6db6982f0e12da59a268fe498edcb33237e906",
      "size_mb": 701.919
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "python convert_lfm25.py LiquidAI/LFM2.5-1.2B-Instruct out_lfm25_12b_fp --fp && python ../minicpm_work/quantize_litertlm.py apply out_lfm25_12b_fp/model.litertlm lfm25_int4.litertlm --recipe wi4b32_wi8 --algo octav",
    "quantization": "int4 blockwise-32 + OCTAV linears, int8 embedding, convs float",
    "tool": "litert-torch",
    "tool_version": "0.9.1"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-23",
          "decode_tokens_per_s": null,
          "delegated_ops": 536,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": "runtime/core/engine_advanced_impl.cc:308",
          "evidence": [
            "VERBOSE: Replacing 536 out of 579 node(s) with delegate (LITERT_CL) node, yielding 2 partitions for subgraph 0 (prefill_1024).",
            "ADD: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "GATHER_ND: Operation is not supported.",
            "GREATER_EQUAL: Can't parse inputs with const tensors.",
            "LESS_EQUAL: Can't parse inputs with const tensors.",
            "SUM: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "536 operations will run on the GPU, and the remaining 43 operations will run on the CPU.",
            "runtime/core/engine_advanced_impl.cc:308"
          ],
          "failure_class": "engine_create_failed",
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": false,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": 579,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.0/2026-08-23/lfm2.5-1.2b-instruct-int4__galaxy-s26.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-06",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.17.0",
            "soc": null,
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "compat_check status ok (runtime /Users/USER/code/litertlm-convert/.qa-venvs/litert-lm-0.17.0/bin/litert-lm 0.17.0); fixed-question answer: '17 + 25 = 42'"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.17.0/2026-09-06/lfm2.5-1.2b-instruct-int4__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-07-22",
          "decode_tokens_per_s": null,
          "delegated_ops": 536,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.14.0",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": "E0000 00:00:1784700818.954375 23966061 engine.cc:807] Failed to create engine: INTERNAL: ERROR: [third_party/odml/litert_lm/runtime/executor/llm_litert_compiled_model_executor.cc:1978]",
          "evidence": [
            "ADD: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "GATHER_ND: Operation is not supported.",
            "GREATER_EQUAL: Can't parse inputs with const tensors.",
            "LESS_EQUAL: Can't parse inputs with const tensors.",
            "SUM: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "536 operations will run on the GPU, and the remaining 43 operations will run on the CPU.",
            "E0000 00:00:1784700818.954375 23966061 engine.cc:807] Failed to create engine: INTERNAL: ERROR: [third_party/odml/litert_lm/runtime/executor/llm_litert_compiled_model_executor.cc:1978]"
          ],
          "failure_class": "engine_create_failed",
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": false,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": 579,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.14.0/2026-07-22/lfm2.5-1.2b-instruct-int4__mac-studio-m4-max.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-08-12",
          "decode_tokens_per_s": null,
          "delegated_ops": 536,
          "env": {
            "device": "Pixel 8a",
            "machine_label": "pixel-8a-cl-pinned",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Google Tensor G3",
            "vendor_sdk": null
          },
          "error": "F0000 00:00:1786530335.514764   32727 litert_lm_main.cc:187] Check failed: MainHelper(argc, argv) is OK (UNKNOWN: ERROR: [runtime/executor/llm_litert_compiled_model_executor_factory.cc:200]",
          "evidence": [
            "VERBOSE: Replacing 536 out of 579 node(s) with delegate (LITERT_CL) node, yielding 2 partitions for subgraph 0 (prefill_1024).",
            "ADD: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "CAST: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM …[trace truncated]",
            "GATHER_ND: Operation is not supported.",
            "GREATER_EQUAL: Can't parse inputs with const tensors.",
            "LESS_EQUAL: Can't parse inputs with const tensors.",
            "SUM: Tensor type(INT64) is not supported. litert_torch.generative.export_hf.core.exportable_module.LiteRTExportableModuleForDecoderOnlyLMPrefill/transformers.models.lfm2.modeling_lfm2.Lfm2ForCausalLM_ …[trace truncated]",
            "536 operations will run on the GPU, and the remaining 43 operations will run on the CPU.",
            "F0000 00:00:1786530335.514764   32727 litert_lm_main.cc:187] Check failed: MainHelper(argc, argv) is OK (UNKNOWN: ERROR: [runtime/executor/llm_litert_compiled_model_executor_factory.cc:200]"
          ],
          "failure_class": "engine_create_failed",
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": false,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": false,
          "total_ops": 579,
          "ttft_ms": null
        },
        "source": "data/device_runs/0.16.0/2026-08-12/lfm2.5-1.2b-instruct-int4__pixel-8a.json"
      },
      {
        "device": "raspberry-pi-5",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-01",
          "decode_tokens_per_s": 9.29,
          "delegated_ops": null,
          "env": {
            "device": "Raspberry Pi 5 Model B Rev 1.1",
            "machine_label": "raspberry-pi-5-cooled-52c",
            "os_build": "Linux-6.18.34+rpt-rpi-2712-aarch64-with-glibc2.41",
            "runtime": "litert-lm",
            "runtime_version": "0.16.1",
            "soc": "Broadcom BCM2712",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "pi5 LLM sweep row: `litert-lm benchmark --backend cpu --cpu-thread-count 4 -p 256 -d 256 --runs 1 --cache memory` (wave-2 driver pi5_llm_bench.py; --cache memory rather than the house --cache no, which OOM-kills every >=1.2B file on the 8 GB Pi — equivalence measured on granite-350m int8, +2-3%), 3 invocations per file with cool-down to <=52 C between them, vcgencmd measure_temp + get_throttled logged per invocation, peak RSS polled from /proc; throughput = median of the three invocations (spread in metrics); a row counts as measured only when the real-generation gate (`litert-lm run`, degenerate-output check) passed and every invocation exited 0 with get_throttled 0x0",
            "versions: cpu=Raspberry Pi 5 Model B Rev 1.1, litert-lm=0.16.1, litert-lm-api=0.16.1, platform=Linux-6.18.34+rpt-rpi-2712-aarch64-with-glibc2.41, python=3.13.5",
            "cache mode 'memory'; -p 256 -d 256 --runs 1 --cpu-thread-count 4",
            "gate ('What is 17 plus 26? Answer with the number only.'): status pass, exit 0, wall 54.5 s, output head '53'",
            "invocation 0: exit=0 wall_s=92.0 temp 49.9->53.2C throttled=0x0 prefill_tps=54.19 decode_tps=9.34 ttft_s=4.831 init_s=34.2405 peak_rss_mb=1455",
            "invocation 1: exit=0 wall_s=100.1 temp 50.5->53.8C throttled=0x0 prefill_tps=54.09 decode_tps=9.26 ttft_s=4.841 init_s=33.8695 peak_rss_mb=1455",
            "invocation 2: exit=0 wall_s=100.0 temp 51.0->53.8C throttled=0x0 prefill_tps=53.67 decode_tps=9.29 ttft_s=4.8774 init_s=33.81 peak_rss_mb=1454"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 256.0,
            "decode_tps_max": 9.34,
            "decode_tps_min": 9.26,
            "init_s": 33.8695,
            "init_s_max": 34.2405,
            "init_s_min": 33.81,
            "invocations": 3.0,
            "prefill_tokens": 256.0,
            "prefill_tps_max": 54.19,
            "prefill_tps_min": 53.67,
            "runs_per_invocation": 1.0,
            "threads": 4.0,
            "ttft_s_max": 4.8774,
            "ttft_s_min": 4.831
          },
          "output_match": null,
          "peak_mem_mb": 1455.0,
          "prefill_tokens_per_s": 54.09,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 4841.0
        },
        "source": "data/device_runs/0.16.1/2026-09-01/lfm2.5-1.2b-instruct-int4__raspberry-pi-5.json"
      }
    ]
  },
  "model": {
    "family": "lfm2.5",
    "id": "lfm2.5-1.2b-instruct-int4",
    "license": "lfm-open-license-v1.0",
    "source_url": "https://huggingface.co/litert-community/LFM2.5-1.2B-Instruct",
    "task": "text-generation"
  },
  "pitfalls": [
    "This artifact cannot use a GPU delegate: it comes from the litert-torch 0.9.1 export lineage, whose ShortConv block emits GATHER_ND and INT64 ops that GPU delegates reject. The delegate takes 536 of 579 operations and engine creation then fails with 'Hint fully delegated to single delegate is set, but the graph is not fully delegated' — measured on a Pixel 8a with litert_lm_main built from the litert-lm v0.16.0 tag, and the same signature on macOS. The repo also ships LFM2.5-1.2B-Instruct_int4_gpu.litertlm, a litert-torch 0.9.3 re-export that delegates fully; this record is about the CPU-only file.",
    "The 0.9.1 exporter needs the ShortConv prefill-pad fix that convert_lfm25.py applies: the stock block saves its conv state from the padded columns of a prefill chunk, corrupting the first generated token of nearly every reply. It is easy to miss — the model recovers after about one token and GSM8K still parses answers, it just loses roughly 20 points.",
    "Quantize convs at export time only. Post-hoc ALL_SUPPORTED int8 through ai-edge-quantizer kills the conv layers (no output); post-hoc recipes must stay on linears and the embedding (wi8fc, wi4b32_wi8).",
    "litert-lm >= 0.15 needs an ExecutorMetadata section for this hybrid: files exported before that run on 0.14 but fail at inference on 0.15 with 'missing some output TensorBuffers'. The published files were repaired in place on 2026-08-04.",
    "The device-run records for this id (0.14.0/2026-07-22, gallery-1.0.15/2026-07-23) were measured on the pre-repair artifact (sha256 e51f10c7...), before the 2026-08-04 in-place ExecutorMetadata repair that produced the currently published file (sha256 a28b5c59...). The repair appends a metadata section and leaves weights and graph byte-identical, so throughput and quality carry over; the sha256 in this card is the published file's.",
    "Decode speed depends strongly on the KV budget: --max-num-tokens 4096 costs about 24% of decode against 1024 on the same file.",
    "The artifact name here is the published one. The sha256 is the HF LFS oid of litert-community/LFM2.5-1.2B-Instruct/LFM2.5-1.2B-Instruct_int4.litertlm; the local working copy of the same bytes is named LFM2.5-1.2B-Instruct_int4_0150fix.litertlm after the 2026-08-04 in-place ExecutorMetadata repair, which is why that filename appears in the device-run and gate records."
  ],
  "schema_version": "1.2"
}
