Engineering notes from porting Edge0/Audio8-TTS-Preview-0.6b
(Apache-2.0; 601M + a 337M codec) to Core AI: a DualAR text-to-speech model of the Fish Audio S2 Pro design — a
Qwen2.5-shaped slow AR (24 layers, 896 wide, 14 heads / 2 KV heads) that predicts one semantic token per 46 ms
frame, a 4-layer fast AR that predicts the frame’s nine other codebooks one after another, and a 44.1 kHz
DAC-style codec (semantic book 4,096 × 8, nine residual books 1,024 × 8, a window-128 transformer, a SEANet decoder
to 2,048 samples per frame). Eleven languages, zero-shot voice cloning from a 0.5–30 s reference. The zoo’s first
DualAR / semantic-token TTS and its first sampler inside a graph. Card:
models/audio8-tts/README.md; scripts:
conversion/audio8_tts/.
| stage | Core AI form |
|---|---|
| prompt | host: the publisher’s segments encoded one at a time (<\|im_start\|>system\n, the system line, …, <\|voice\|>), packed as [11, P] — row 0 ids, rows 1..10 the reference’s codebooks under its semantic ids |
| slow AR prefill | prefill(codes [1,11,32], pos) in windows over one KV state pair [24,1,2,2048,64]; the last real row’s logits (4,097 rows: the semantic ids then eos — the only rows the sampler can pick) and post-norm hidden |
| one frame | frame(codes [1,11,1], pos, noise_slow [2,4097], window [10], noise_fast [9,4096], forced [11], use_forced): the slow step, the semantic draw with the RAS rule, the fast AR’s ten rows and nine codebook draws — one runtime call per frame |
| codec | main(codes [1,10,160]) -> wav [1, 327680], stateless; every op causal, so a stream decodes in 160-frame windows and keeps the last 32 |
| voice registration | main(audio [1,1,442368]) -> codes [1,10,216] (the encoder, 10 s bucket) |
topk(50)
bounds the top-p candidates (top-k removes everything else anyway), logsumexp gives the full-vocabulary softmax
normaliser, a cumulative sum over the 50 sorted probabilities gives the top-p cut with the best always kept, the
kept scores are divided by the temperature, and the draw is argmax(softmax(kept) / -log(u)) with u gathered at
the candidates’ indices — the Gumbel-max form ArkttsModel._sample uses. The randomness is the vector u, an
input, so a recorded u replays the publisher’s choice exactly on identical logits: the graph’s own draws match
the NumPy sampler 20/20 on random logits and the oracle’s tokens frame for frame under teacher forcing. The fused
frame runs in ~32 ms on the same machine under the same load (about a third of it the slow step); the fast
AR’s ten rows are unrolled with their keys concatenated in-graph, so it needs no state and no call of its own.ArkttsSemanticLogitsProcessor sets every
logit outside the semantic range and eos to -inf before sampling, so the head can be the embedding table’s 4,097
rows (7 MB) instead of 155,776 (279 MB fp16). Over 1,695 oracle steps the full-vocabulary argmax never fell outside
those rows either._precompute_rope rounds cos / sin to bf16 in the fp32 model too, so
the fp32 oracle carries ~3-digit angles. The port bakes the same bf16-rounded tables as fp32 constants (interleaved
pairs, GPT-J style — apply_rope_pairs, not rotate-half); swapping in exact fp32 angles moves the fp16 logits to
cos 0.9991 against the bf16-table build. Apple’s RoPE(interleaved=True) composite implements the same pairing
(max |Δ| 0.0 against the publisher’s rotation on the [B, H, S, D] layout).generate creates previous as zeros after step 0, so the
first frame’s semantic never enters the 10-token window; from step 2 the window rolls. The publisher’s ONNX runtime
appends from step 0 — it does not reproduce generate. The oracle here is generate (asserted identical to the
recorded loop with the same seed), and the graph takes the window as an input (-1s for “none”).fast_output has 4,096 rows for every codebook, but codebooks 1..9 of
the codec have 1,024 entries; the publisher’s decode() clamps. In 8,000 oracle draws one codebook sample exceeded
1,023, so the clamp is in the codec graph too.synthesize:
the codec once at the end, windows sharing 128 frames of context) brings the phone to about real time (RTF 0.92 at nominal, 0.93 on a
thermally serious phone — the frame is dispatch-bound either way). The codec, not the transformers, is where the phone’s next milliseconds are.Failed to allocate storage for NDArray … sk: ioSurface). One child interpreter per fixture
(pocket-tts-port.md, defect 2). The fused graph cuts a fixture to ~200 calls; the workers stay.V=~/code/coreai/coreai-models/.venv/bin/python; C=conversion/audio8_tts
$V $C/oracle_audio8.py --check-generate # fp32 oracle (publisher's code), every draw recorded
$V $C/prompt.py # tokenizer-only prompt == the processor's, 18/18
$V $C/parity_audio8.py # plain-torch re-author vs oracle, eager fp32
$V $C/export_audio8_frame.py --mode int8 # the fused asset (+ eager check, quick GPU gate)
$V $C/export_audio8.py --part codec # the codec decoder
$V $C/audio8_encoder.py --frames 216 # the encoder (voice registration)
$V $C/gate_audio8_frame.py --mode int8 # 18 fixtures: teacher-forced, codec, free run, ASR, speaker