Core AI model zoo

LFM2.5-VL (450M / 3B, vision-language) — Core AI

The smallest VLM in this catalog by a factor of three, and its detail-tier sibling. Core AI ports of LiquidAI/LFM2.5-VL-450M and -3B: image + text → text on the pipelined-engine fast path.

Which one: the 450M (658 MB, iPhone-gated at 112 tok/s) is the one that fits beside an app — every other VLM here starts at ~2 GB resident. The 3B (int8 3.9 GB on Mac, int4 2.8 GB on iPhone at ~20 tok/s) is the one that answers with detail: where the 450M says “two cats on a pink couch”, the 3B says “the cat on the left is smaller, with a gray and black striped coat, while the cat on the right is larger with a brown and black striped pattern.” One exporter builds both.

Architecture (model_type: lfm2_vl), 450M: a SigLIP2-NaFlex tower (hidden 768, 12 layers) whose patch embedding is a Linear over pre-flattened 16×16×3 patches and whose 16×16 position table is bilinearly resized to the actual patch grid; a 2-layer projector (pixel_unshuffle(2) → linear_1 → gelu → linear_2, no LayerNorm); and the LFM2 hybrid text decoder already shipped in this repo (hidden 1024, 16 layers = 10 short-conv + 6 GQA attention, vocab 65 536, tied head, RoPE θ 1e6), reached by nothing more than a model.language_model. key prefix. Only the tower and the projector were authored fresh — knowledge/lfm2.5-vl-port.md.

⬇️ Converted .aimodel bundles: mlboydaisuke/LFM2.5-VL-450M-CoreAIgpu-pipelined/lfm2_5_vl_450m_vision_fp16/ + gpu-pipelined/lfm2_5_vl_450m_decode_int8lin/ (the pair), …_textcore/ (the same decoder with no image input), and ios-h18p/ AOT variants of all three, each gated on an iPhone 17 Pro. LFM Open License v1.0.

What it is for

A 450M VLM answers scene-level questions — what is in the picture, where it is, what colours dominate — and misses fine-grained geometry. That is the checkpoint, not the port: the same weights on LiteRT/Pixel show the same split. Treat it as a caption/triage model that fits beside an app, not as a document reader.

Numbers (M4 Max, macOS 27.0 26A5378n, Xcode 27.0 27A5218g, coreai-torch 0.4.1)

artifact size measured numerics
vision fp16 (ship) 181 MB 18.0 ms/image (median of 12, 512×512 → 256 tokens) image_embeds cos 0.999996 vs fp32 HF
vision int8lin 97 MB 21.8 ms/image — slower, see below cos 0.999729
decoder int8lin (ship) 477 MB text core: 609.2 prompt / 387.2 decode tok/s suite 7/9 cases token-exact; coreai_gate.py PASS 16/16
decoder fp16 717 MB not benchmarked logits cos 0.999994; suite 8/9
decoder int4lin 349 MB not benchmarked suite 0/9NO-GO, see below

llm-benchmark -p 128 -g 256 -n 3, COREAI_CHUNK_THRESHOLD=1. The Mac tok/s row is the text core (the same weights exported without the image input): llm-runner cannot bind the VLM bundle’s image_embeds buffer, so the text core is the Mac proxy — the same substitution the MiniCPM-V-4.6 card documents.

iPhone 17 Pro (AOT h18p, PipelinedBench, settled)

bundle prefill decode numerics
decode_int8lin (the VLM bundle, image bound) 123.2 112.0 nat 16/16 + image oracle 24/24
decode_int8lin_textcore 122.1 110.6 nat 16/16 + oracle 16/16
decode_int8lin, PB_G=1024 122.4 108.6 nat 16/16, no collapse
vision_fp16 33.6 ms/image cos 0.999995 vs the Mac tower’s own output

The tower’s first-ever encode costs ~860 ms of on-device MPSGraph compile; every encode after that is 33.6 ms (medians of 9 and 15 runs across two launches: 33.6 / 36.3). Warm it with a dummy encode at load and the user’s first photo is the warm number — the same lesson the MiniCPM-V-4.6 port wrote down at ~2.7 s. Mac is 18.0 ms, so a phone pays 1.9x.

xcrun coreai-build compile … --platform iOS --preferred-compute gpu --architecture h18p (no --expect-frequent-reshapes: on iOS it makes the runtime discard the AOT specialization and compile on device, which SIGSEGVs with no log). Engine ready in 0.5 s warm / 4.2 s cold.

The VLM bundle and the text core measure the same speed within noise — binding a 256×1024 fp16 buffer costs nothing per step — which is what makes the Mac text-core proxy legitimate rather than convenient.

The PB_G=1024 row is mandatory, not extra: the iOS compiler miscompiles KV specializations at seq ≥ 2048 and g ≤ 256 cannot see it (zoo PR #6). This bundle is clean at 1024.

On device the image path is right but not token-identical to fp32. The device describes the same picture and drops one adjective at a near-tie: fp32 says “two tabby cats … is stretched out on its side with its p[aws]”, the device “two cats … is lying on its side with its head resting on the” — two forks, both word choice, with the ten tokens between them identical. The gate therefore checks the device-verified sequence (recorded in PipelinedBench with the fp32 one beside it in _smoke/lfm25vl_ref/vl_ref.json), which is the same thing the MiniCPM-V-4.6 card does for the same class of fp16 near-tie.

The tower is gated on device against its own Mac output, not against fp32: what the phone has to reproduce is the encode the decoder was gated with. cos 0.999995 says the h18p tower and the h16c one compute the same image.

Gates

The port is gated seam by seam against an fp32 transformers-5 oracle before anything was exported (_smoke/lfm25vl_ref.py_smoke/test_lfm25vl_torch_ladder.py), then again through the engine:

Cosines are computed in float64; in float32 the reduction over ~10⁶ elements prints cos 1.000088 for two identical tensors.

What compression costs, read rather than scored

int8 moves this decoder’s logits much more than the family’s larger members (logits_last cos 0.9866 vs 0.99992 for the 1.2B) — a 350M decoder has less redundancy per block. The fp16 bundle at cos 0.999994 through the identical path is what proves that is compression, not a port bug.

Token counts against an fp16 baseline, not fp32 alone, because greedy decoding turns any near-tie into a different tail: the fp16 bundle itself diverges on one of the nine cases. Both int8 divergences are re-wordings (“contrast with the cats.” → “contrast with the cats’ fur.”).

int4lin is the cliff — 0/9, and the failure mode is not broken repetition but fluent drift: a kitchen becomes “a traditional Italian kitchen” where fp32 says “historical or rustic”, and one answer reports a kitchen’s “Beige - visible in the carpeted floor”. No loss curve catches that; reading the generations does.

The int8 vision tower is a Mac anti-optimization: 97 MB instead of 181 MB, but cos 0.999729 instead of 0.999996, 21.8 ms instead of 18.0, and 6/9 suite cases instead of 7/9. The tower is not bandwidth-bound at this size, so dequant is a net loss. It stays exported because a phone’s bandwidth budget is a different question — one no measurement here answers.

How the image reaches the decoder

Convert / verify

# vision tower + VLM decoder (the published pair)
python conversion/export_lfm25vl_pipelined.py int8lin --vision-mode fp16

# the text-core proxy: same weights, no image input (benchmarks + coreai_gate)
python conversion/export_lfm25vl_pipelined.py int8lin --text-core --skip-vision

# oracles (transformers >= 5 ONLY — 4.x applies a projector LayerNorm this checkpoint does not have)
~/code/litertlm-convert/.venv-vl093/bin/python _smoke/lfm25vl_ref.py --resize 512x512
~/code/litertlm-convert/.venv-vl093/bin/python _smoke/lfm25vl_suite_ref.py

# gates
python _smoke/test_lfm25vl_torch_ladder.py --ref _smoke/lfm2_5_vl_450m_ref_512x512.npz
python _smoke/test_lfm25vl_aimodel_gate.py
python _smoke/test_lfm25vl_suite_gate.py --mode int8lin --show-text
python conversion/coreai_gate.py <text-core bundle> LiquidAI/LFM2.5-VL-450M -n 16

The 3B

Same command with --hf-id LiquidAI/LFM2.5-VL-3B, and the model code needed no change: a wider tower (hidden 1152, 27 layers) and a 30-layer 128k-vocab decoder come off the config, and vision_block_size drops int8 to per-block-16 on its own because the tower’s 4304-wide MLP is not divisible by 32.

artifact size measured (M4 Max) numerics
vision fp16 815 MB 75.7 ms/image image_embeds cos 0.999995 vs fp32 HF
decoder int8lin (Mac ship) 3.1 GB suite 7/9; logits_last cos 0.999970
decoder int4lin 2.0 GB suite 7/9 — same as int8 and as the fp16 baseline
decoder fp16 (baseline) 5.2 GB 48/48 on the fixture, cos 0.999999; suite 7/9
text core int8lin 3.1 GB 120.9 prompt / 105.3 decode tok/s the Mac speed proxy

The torch ladder is cos 1.000000 at every seam with 48/48 token-exact greedy, at fp32 and at fp16, and the NumPy host path is 48/48 too.

int4 costs the 3B nothing — 7/9, the same cases as fp16 — which is the opposite of the 450M (0/9, fluent drift). Same family, same recipe, opposite verdict: int4 tolerance is a property of the model’s size, and the only way to know is to read the generations of the one in front of you.

The 3B reaches a phone, at int4. Its int8lin AOT resources.bin measures 3.13 GiB and does not load; int4lin’s measures 2.03 GiB and does — 24/24 on the image oracle, nat 16/16, 27.5 prefill / 19.3–22.8 decode tok/s on an iPhone 17 Pro, clean at PB_G=1024. That is worth spelling out because this repo’s own note put the wall at 2 GiB (2^31) from an earlier 1.96 ✅ / 3.92 ❌ bracket, and 2.03 GiB was written up here as expected to fail. It did not. The bracket is now 2.03 ✅ / 3.92 ❌ and the lesson is to try the phone rather than infer from 30 MiB.

int8lin remains the Mac ship (bigger and no wall to clear there); int4lin is the iPhone build, and it costs nothing on the suite (7/9, the same cases as fp16).

python conversion/export_lfm25vl_pipelined.py int8lin --hf-id LiquidAI/LFM2.5-VL-3B
python conversion/export_lfm25vl_pipelined.py int4lin --hf-id LiquidAI/LFM2.5-VL-3B --skip-vision

Two host details do not carry over from the 450M, and both are silent when wrong: the 3B declares resample: 3 (PIL BICUBIC) where the 450M declares 2 (BILINEAR), and the 3B tokenizer’s post-processor does not prepend <|startoftext|> while the chat template starts with it — without BOS the 3B answers " F, F, F, F".