{
  "artifacts": [
    {
      "file": "codec_decoder_fp16_T128.tflite",
      "sha256": "c2b326927b20ac1efe98f2736bfd1fff613c248e7dfe3ba9b678c6a82db1bab4",
      "size_mb": 249.484
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-28",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-28",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-28/audio8-tts-preview-0.6b__codec_decoder_fp16_t128.json"
  },
  "conversion": {
    "command": "T=128 SUFFIX=_g2 python export_codec.py (codes [1,10,128] int32 -> wav [1,1,262144] fp32, static T, fp32; checked against the oracle waveform) -> python quantize_codec.py out/codec/codec_decoder_fp32_T128_g2.tflite fp16 -> python assemble_ship.py out/ship (copied as codec_decoder_fp16_T128.tflite; built_from codec/codec_decoder_fp16_T128_g2.tflite) (REPRODUCE entry)",
    "quantization": "fp16 weights (ai-edge-quantizer FLOAT_CASTING), fp32 compute, fp32 RVQ codebook tables; exact vs the fp32 graph (corr 1.000000, max|d| 1.3e-3); the fp32 .tflite itself is within 1.9e-6 of the vendor's waveform (FINDINGS §2, §4; card Accuracy)",
    "tool": "litert-torch (fp32 export of the vendor codec module imported from the snapshot with three patches: torch.jit.script snake -> plain python, torch.polar rope table -> eager-cached (cos,sin) constants, attention rewritten with folded GQA heads and an eagerly built additive causal/window mask) + ai-edge-quantizer fp16 FLOAT_CASTING (quantize_codec.py fp16); repro = hf-to-litertlm audio8_tts_work/ (REPRODUCE entry; FINDINGS §2)",
    "tool_version": "litert-torch 0.9.4 (~/venvs/lt094dev, the interpreter named in every export/quantize script's usage line: torch 2.13.0, transformers 5.14.1, ai-edge-quantizer 0.9.0, ai-edge-litert 2.2.0 — pip show 2026-09-28; REPRODUCE entry: 'env with litert-torch 0.9.4, ai-edge-litert 2.2.0, ai-edge-quantizer'); PyTorch oracle transformers 4.57.6 / torch 2.12.1 CPU fp32 in ~/parakeet-env (FINDINGS §1)"
  },
  "cross_runtime": [],
  "delegation": {
    "backend": "gpu_mldrift",
    "blocking_ops": [
      "DEQUANTIZE",
      "SIN",
      "PAD",
      "TRANSPOSE_CONV",
      "SQUARED_DIFFERENCE",
      "DEPTHWISE_CONV_2D",
      "TANH",
      "LOGISTIC",
      "SLICE",
      "RESHAPE",
      "EMBEDDING_LOOKUP",
      "MAXIMUM",
      "MINIMUM"
    ],
    "coverage_ops_pct": 81.6,
    "lint_report_version": "1.1",
    "litert_version": "2.2.0",
    "matched_provenance_counts": {
      "measured": 970,
      "unmatched": 99
    },
    "partitions": 154
  },
  "model": {
    "family": "audio8-tts",
    "id": "audio8-tts-preview-0.6b__codec_decoder_fp16_t128",
    "license": "apache-2.0",
    "source_url": "https://huggingface.co/litert-community/Audio8-TTS-Preview-0.6b",
    "task": "text-to-speech"
  },
  "pitfalls": [
    "One call covers up to 128 frames = 5.9 s (128 x 2,048 samples at 44.1 kHz). The decoder is causal but its post-transformer stacks eight 128-frame sliding-window layers, so decoding in windows is not sample-exact against one offline decode: with 127 frames of left context the frames after a boundary differ by max|d| 0.08 (corr 0.9968), with 32 frames of context corr 0.70. The host loop makes one T128 call for utterances up to 5.9 s, one T192 call up to 8.9 s, and beyond that T192 windows with 128 frames of left context (the vendor's ONNX runtime uses context 128 + a 1-frame guard) (card Limitations; FINDINGS §7).",
    "A 256-frame decoder is not shipped: with --use_gpu on the Galaxy S26 it fails to prepare ('Dilated im2col buffer size overflowed', node 934 CONV_2D) (card Performance note; FINDINGS §7; NOTES S26 clean legs).",
    "Galaxy S26 GPU (OpenCL) runs 936 of 1,069 ops; the CPU remainder is the fp16-weight DEQUANTIZE and the codebook EMBEDDING_LOOKUP ('Empty quantization params'); the fp32 decoder delegates 971/971 at the same 924 ms with a 1,462 MB footprint (card Performance; NOTES). The earlier attention build (repeat_interleave GQA, interleaved rope) delegated only 422/1,246 ops: BROADCAST_TO 'Operation is not supported' and rank-5 CONCATENATION/RESHAPE 'Tensor dimensions must be less than 5' — folding the shared query heads into the matmul rows and pre-building the masks cleared them (FINDINGS §2; NOTES).",
    "Every index that reaches a gather is clamped in-graph so random-input benchmark tools cannot fault ('gather_nd index out of bounds' on the S26 before the clamps) (FINDINGS §2).",
    "Sampled model: quantized graphs sample a different but valid trajectory, so per-frame agreement with the reference is not a meaningful metric; the gate is the transcript (whisper large-v3-turbo WER/CER vs the input text) and the voice (TitaNet-L speaker cosine vs the reference clip) on 14 seeded sentences (6 en + 6 ja with a cloned voice, one of each without). PyTorch reference en WER 1.1% / ja CER 0.0% / cosine 0.66 / 0.74; int8 slow + int8 fast + fp16 codec 1.1% / 0.0% / 0.68 / 0.76; int4 slow 1.1% / 0.0% / 0.63 / 0.77; int8 codec on the reference codes 1.1% / 0.0% / 0.67 / 0.73. The one English error is shared with the reference (the model drops the first word of one sentence). The fp32 graphs reproduce the reference code sequence frame for frame on all 14 sentences (card Accuracy; FINDINGS §2, §5).",
    "The graphs are driven by a host loop (audio8_tts_litert.py, Python, ai-edge-litert Interpreter, CPU): prompt construction, chunked prefill, the vendor's sampler (temperature 0.7 / top-p 0.9 / top-k 50, repetition-aware re-draw), 10 fast-AR calls per frame, windowed codec decode, voice registration; on a phone the same graphs run from Kotlin/C++ through the CompiledModel API — no on-device functional run exists yet, the Galaxy S26 rows are per-graph benchmark_model latency + delegate coverage, functional parity is Mac-only on the same files (card Files + Performance; FINDINGS §8).",
    "Prompt length + generated frames must stay under 2,048 positions (the model's max_seq_len = the KV cache); the default cap is 512 frames (about 24 s) per call (card Limitations).",
    "GQA on the KV cache must not be expressed as repeat_interleave: it lowers to BROADCAST_TO (48 per step) and made decode 615 ms/step on the Mac (thread-count independent); folding the 7 query heads that share a kv head into the matmul row dimension removed it (16 ms/step fp32). RoPE: the vendor applies interleaved-pair rotation with a bf16-rounded table; the port permutes the q/k rows of wqkv to the rotate-half layout with the same bf16-rounded cos/sin as fp32 constants, bit-identical to the vendor buffers (FINDINGS §2).",
    "fp16 weights give the AR graphs no speed or RAM gain on the CPU (XNNPACK unpacks them to fp32 at init), so the AR graphs ship int8/int4 and only the codec ships fp16 (FINDINGS §4).",
    "Preview checkpoint: the vendor documents limited dialect coverage and sensitivity to noisy or mis-transcribed references. Generated speech can be misused for impersonation; obtain consent before cloning a voice and disclose synthetic audio (card Limitations)."
  ],
  "schema_version": "1.1"
}
