OpenBMB VoxCPM2 (2B) converted to Apple Core AI, running fully on-device on iPhone (A19 Pro / iPhone 17 Pro) and Mac — no network. The 2B, 48 kHz successor to VoxCPM-0.5B-CoreAI.
A tokenizer-free diffusion TTS: a MiniCPM4 28-layer text-semantic LM + an 8-layer residual acoustic LM drive a 12-layer LocDiT flow-matching diffusion head, decoded by a 48 kHz AudioVAE. Five Core AI bundles + a few host-side projections.
⚡ One line — run the kit’s task op on this model
(import CoreAIOps; no session, no model plumbing, downloads on first use):
let audio = try await CoreAI.speak(text, options: .model("voxcpm2-2b"))
Twenty ops, one shape — Cookbook.
▶️ Run it (source) — the Speak runner (GUI + CLI, one app for every text-to-speech model in the catalog):
git clone https://github.com/john-rocky/coreai-kit
open coreai-kit/Examples/Speak/Speak.xcodeproj
# → Run, then pick "VoxCPM2 2B" in the model picker
# agents / headless (macOS):
cd coreai-kit/Examples/Speak
swift run speak-cli --model voxcpm2-2b --text "Hello from Core AI." --output hello.wav
💻 Build with it — complete; the glue is kit API, copy-paste runs:
import CoreAIKit
let speaker = try await KitSpeaker(catalog: "voxcpm2-2b")
let audio = try await speaker.synthesize(text)
// audio.samples: 48 kHz mono PCM in [-1, 1] — play it or write a WAV
The take-home is Examples/Speak/Sources/QuickStart.swift
— this exact code as one typed function, no UI; the CLI is an argument shell over it, and
the GUI drives the same KitSpeaker(catalog:) and plays the samples.
Live playback? synthesizeStreaming(_:onChunk:) hands you ~0.5 s chunks as they decode,
so audio starts before the whole clip exists. The WAV container is your app’s territory
(the runner ships a 20-line writer).
Integration checklist
https://github.com/john-rocky/coreai-kit → product CoreAIKitdownloadProgress callback)| dir | contents |
|---|---|
macos/ |
JIT .aimodel bundles (Mac): int8 base/res decode + prefill, fp16 feat_decoder / feat_encoder / vocoder |
ios/ |
AOT .aimodelc bundles (iOS h18p, GPU): same five + the two int8 prefill bundles |
voxcpm2_host_glue/ |
embed table + projections / FSQ-512 / stop-head / fusion (.bin + manifest) |
tokenizer/ |
the VoxCPM2 tokenizer (Llama fast) |
The backbone LMs are weight-only int8 (the size driver); the diffusion + VAE stay fp16 (the continuous-feedback path is quant-sensitive — same split mlx-community uses).
Runs through coreai-kit VoxCPM2TTS, wired into the
coreai-model-zoo coreai-audio app (“Voice 2B” tab).
Conversion + gates + export scripts: coreai-model-zoo/conversion/voxcpm/ (*_v2.py).
let tts = try await VoxCPM2TTS(paths: .standard(artifactsRoot: root, lm: .int8))
let wav = try await tts.synthesize("On device speech synthesis, running entirely on your iPhone.") // 48 kHz Float PCM
Reimplemented in exportable Core AI overlays and gated end-to-end against the official model: backbone / feat_decoder / feat_encoder cos 1.0, full chain magspec 0.996; every exported bundle engine-gated cos ≥ 0.9999.
Apache-2.0 (commercial OK), inherited from openbmb/VoxCPM2. Not affiliated with OpenBMB or Apple. Community port.
⬇️ Download: 🤗 mlboydaisuke/VoxCPM2-CoreAI — this card and the
model page are the same document; scripts/gen-cards keeps the Use it block in sync.
Reproduce it: python3 conversion/zoo_convert.py show voxcpm2-2b prints the command and its
prerequisites; recipe.toml is the record.