---
family: lfm2.5
license: lfm-open-license-v1.0
model_id: lfm2.5-1.2b-instruct-int4
source_url: https://huggingface.co/litert-community/LFM2.5-1.2B-Instruct
task: text-generation
---

# lfm2.5-1.2b-instruct-int4

| | |
|---|---|
| **Task** | text-generation |
| **Family** | lfm2.5 |
| **Source** | https://huggingface.co/litert-community/LFM2.5-1.2B-Instruct |
| **License** | lfm-open-license-v1.0 |

## Conversion

| | |
|---|---|
| **Tool** | litert-torch 0.9.1 |
| **Command** | `python convert_lfm25.py LiquidAI/LFM2.5-1.2B-Instruct out_lfm25_12b_fp --fp && python ../minicpm_work/quantize_litertlm.py apply out_lfm25_12b_fp/model.litertlm lfm25_int4.litertlm --recipe wi4b32_wi8 --algo octav` |
| **Quantization** | int4 blockwise-32 + OCTAV linears, int8 embedding, convs float |

## Artifacts

| File | SHA-256 | Size (MB) |
|---|---|---|
| `LFM2.5-1.2B-Instruct_int4.litertlm` | `a28b5c59ac204e2e51c1f98d2d6db6982f0e12da59a268fe498edcb33237e906` | 701.919 |

## Performance

No benchmark data yet.

## Device runs (NPU / LiteRT-LM)

Model-level on-device results — measured behavior of this exact artifact on the recorded device, only meaningful together with that environment. Throughput is conditional on the measured prompt length (Prompt (tokens); '-' = the source stated none): rows differing only there are different measurements, not re-runs. Full records (failure classes, log evidence): `data/device_runs/0.14.0/2026-07-22/lfm2.5-1.2b-instruct-int4__mac-studio-m4-max.json`, `data/device_runs/0.16.0/2026-08-12/lfm2.5-1.2b-instruct-int4__pixel-8a.json`, `data/device_runs/0.16.0/2026-08-23/lfm2.5-1.2b-instruct-int4__galaxy-s26.json`, `data/device_runs/0.16.1/2026-09-01/lfm2.5-1.2b-instruct-int4__raspberry-pi-5.json`, `data/device_runs/0.17.0/2026-09-06/lfm2.5-1.2b-instruct-int4__mac-studio-m4-max.json`.

| Device | Accelerator | Status | Full delegation | Output match | Prompt (tokens) | Latency p50 (ms) | Prefill tok/s | Decode tok/s | TTFT (ms) | Peak mem (MB) | Environment | Date | Provenance |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| galaxy-s26 | gpu | load_failed | no | - | - | - | - | - | - | - | Galaxy S26 (SM-S942Q) · Qualcomm SM8850 · litert-lm 0.16.0 · Android 16 | 2026-08-23 | measured |
| mac-studio-m4-max | cpu | pass | - | - | - | - | - | - | - | - | Mac Studio (M4 Max) · litert-lm 0.17.0 | 2026-09-06 | measured |
| mac-studio-m4-max | gpu | load_failed | no | - | - | - | - | - | - | - | Mac Studio (M4 Max) · Apple M4 Max · litert-lm 0.14.0 | 2026-07-22 | measured |
| pixel-8a | gpu | load_failed | no | - | - | - | - | - | - | - | Pixel 8a · Google Tensor G3 · litert-lm 0.16.0 | 2026-08-12 | measured |
| raspberry-pi-5 | cpu | pass | - | - | 256 | - | 54.09 | 9.29 | 4841.0 | 1455.0 | Raspberry Pi 5 Model B Rev 1.1 · Broadcom BCM2712 · litert-lm 0.16.1 · Linux-6.18.34+rpt-rpi-2712-aarch64-with-glibc2.41 | 2026-09-01 | measured |

## Pitfalls

- This artifact cannot use a GPU delegate: it comes from the litert-torch 0.9.1 export lineage, whose ShortConv block emits GATHER_ND and INT64 ops that GPU delegates reject. The delegate takes 536 of 579 operations and engine creation then fails with 'Hint fully delegated to single delegate is set, but the graph is not fully delegated' — measured on a Pixel 8a with litert_lm_main built from the litert-lm v0.16.0 tag, and the same signature on macOS. The repo also ships LFM2.5-1.2B-Instruct_int4_gpu.litertlm, a litert-torch 0.9.3 re-export that delegates fully; this record is about the CPU-only file.
- The 0.9.1 exporter needs the ShortConv prefill-pad fix that convert_lfm25.py applies: the stock block saves its conv state from the padded columns of a prefill chunk, corrupting the first generated token of nearly every reply. It is easy to miss — the model recovers after about one token and GSM8K still parses answers, it just loses roughly 20 points.
- Quantize convs at export time only. Post-hoc ALL_SUPPORTED int8 through ai-edge-quantizer kills the conv layers (no output); post-hoc recipes must stay on linears and the embedding (wi8fc, wi4b32_wi8).
- litert-lm >= 0.15 needs an ExecutorMetadata section for this hybrid: files exported before that run on 0.14 but fail at inference on 0.15 with 'missing some output TensorBuffers'. The published files were repaired in place on 2026-08-04.
- The device-run records for this id (0.14.0/2026-07-22, gallery-1.0.15/2026-07-23) were measured on the pre-repair artifact (sha256 e51f10c7...), before the 2026-08-04 in-place ExecutorMetadata repair that produced the currently published file (sha256 a28b5c59...). The repair appends a metadata section and leaves weights and graph byte-identical, so throughput and quality carry over; the sha256 in this card is the published file's.
- Decode speed depends strongly on the KV budget: --max-num-tokens 4096 costs about 24% of decode against 1024 on the same file.
- The artifact name here is the published one. The sha256 is the HF LFS oid of litert-community/LFM2.5-1.2B-Instruct/LFM2.5-1.2B-Instruct_int4.litertlm; the local working copy of the same bytes is named LFM2.5-1.2B-Instruct_int4_0150fix.litertlm after the 2026-08-04 in-place ExecutorMetadata repair, which is why that filename appears in the device-run and gate records.

---

_Rendered by `edge-card` from `card.json` (the source of truth); edit the inputs, not this file._
