{
  "artifacts": [
    {
      "file": "parakeet_tdt_0.6b_v3_30s_i8_stateful.tflite",
      "sha256": "378935f8897f0c713ad2aa97d939c73c44e9f26546e12c1cd6af9c76f386f06a",
      "size_mb": 600.519
    }
  ],
  "benchmarks": [],
  "browser": {
    "backends": [
      {
        "backend": "wasm_xnnpack",
        "date": "2026-09-25",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": 7451.368,
        "loads": true,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": true
      },
      {
        "backend": "webgpu_mldrift",
        "date": "2026-09-25",
        "env": {
          "browser": "chromium",
          "browser_version": "151.0.7922.34",
          "headless": true,
          "jspi": true,
          "litertjs_core_version": "2.5.3",
          "machine_label": "mac-studio-m4-max",
          "os": "macOS",
          "os_version": "27.0.0",
          "webgpu_adapter": {
            "architecture": "metal-3",
            "description": "",
            "device": "",
            "vendor": "apple"
          }
        },
        "full_delegation": null,
        "latency_p50_ms": null,
        "loads": false,
        "max_rel_diff": null,
        "output_match": null,
        "provenance": "measured",
        "runs": false
      }
    ],
    "demo_url": null,
    "sweep_source": "data/sweep/2.5.3/2026-09-25/parakeet-tdt-0.6b-v3__parakeet_tdt_0.6b_v3_30s_i8_stateful.json"
  },
  "conversion": {
    "command": "python compiled_model_api/speech_recognition/convert/convert_to_tflite.py --model nvidia/parakeet-tdt-0.6b-v3 --input_sec 30 --stateful_after 4 --quant drq --sample_audio <wav> --output parakeet_tdt_0.6b_v3_30s_i8_stateful.tflite",
    "quantization": "int8 dynamic-range weights, fp32 activations (the pipeline's drq recipe); the same recipe as litert-community/parakeet-tdt-0.6b-v3's 5 s file, with a 30 s window",
    "tool": "litert-torch via the official litert-samples speech_recognition convert pipeline (compiled_model_api/speech_recognition/convert/convert_to_tflite.py at efca5805, 2026-05-15)",
    "tool_version": "litert-torch 0.9.4, NeMo 3.0.0"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-25",
          "decode_tokens_per_s": null,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (Apple M4 Max, 128 GB)",
            "machine_label": "mac-studio-m4-max",
            "os_build": "macOS 27.0 (26A428)",
            "runtime": "litert-lm",
            "runtime_version": "selfbuilt-66058c82+omni-eval-patches",
            "soc": "Apple M4 Max",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-25-omni-asr-wer-v1-m4max-mac/final/parakeet-tdt-0.6b-v3-30s-clean-cpu.json: \"model_name\": \"parakeet-tdt-0.6b-v3-30s\", \"backend\": \"cpu\", \"num_threads\": 4, \"utterances\": 2620, \"wer_percent\": 2.17, \"substitutions\": 865, \"deletions\": 195, \"insertions\": 92, \"reference_words\": 53000, \"hung_utterances\": [], \"empty_hypotheses\": 8, \"flush_empty\": 8, \"rtfx_processing\": 18.59, \"processing_seconds\": 1046.7, \"audio_hours\": 5.403, \"peak_rss_mb\": 1825.7",
            "https://github.com/john-rocky/edge-llm-bench/blob/main/results/raw/2026-09-25-omni-asr-wer-v1-m4max-mac/final/parakeet-tdt-0.6b-v3-30s-other-cpu.json: \"model_name\": \"parakeet-tdt-0.6b-v3-30s\", \"backend\": \"cpu\", \"num_threads\": 4, \"utterances\": 2939, \"wer_percent\": 5.35, \"substitutions\": 2041, \"deletions\": 508, \"insertions\": 277, \"reference_words\": 52841, \"hung_utterances\": [], \"empty_hypotheses\": 44, \"flush_empty\": 44, \"rtfx_processing\": 17.2, \"processing_seconds\": 1118.3, \"audio_hours\": 5.342, \"peak_rss_mb\": 1820.8",
            "Open ASR Leaderboard scoring (EnglishTextNormalizer + corpus WER, jiwer) over LibriSpeech test-clean (2620 utt, 5.403 h) and test-other (2939 utt, 5.342 h); harness = omni_eval_runner on LiteRT-LM OmniEngine (omni/asr, PushInputSource, one session per utterance, CPU num_threads 4), edge-llm-bench tools/omni-eval. WER is pipeline-level (engine front-end + model + decoder), not the .tflite alone; 'hung' = never returned from TdtDecoder::Decode, killed by the driver watchdog and scored as empty; RTFx = audio seconds / engine wall seconds.",
            "patches 01 + 03 (parakeet-tdt-0.6b-v3-30s metadata entry: inputMilliseconds 30000, nFrames 3000) + 06; NOTES.md table row: | 30 s export, Slaney mel + valid-length stop (patch 06) | 2.17 | 5.35 | (NeMo fp32, leaderboard, H200: 1.92 / 3.59)"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "empty_hypotheses_test_clean": 8,
            "empty_hypotheses_test_other": 44,
            "hung_utterances_test_clean": 0,
            "hung_utterances_test_other": 0,
            "num_threads": 4,
            "peak_rss_mb_test_clean": 1825.7,
            "peak_rss_mb_test_other": 1820.8,
            "rtfx_test_clean": 18.59,
            "rtfx_test_other": 17.2,
            "utterances_test_clean": 2620,
            "utterances_test_other": 2939,
            "wer_test_clean_pct": 2.17,
            "wer_test_other_pct": 5.35
          },
          "output_match": null,
          "peak_mem_mb": 1825.7,
          "prefill_tokens_per_s": null,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": null
        },
        "source": "data/device_runs/selfbuilt-66058c82+omni-eval-patches/2026-09-25/parakeet-tdt-0.6b-v3__parakeet_tdt_0.6b_v3_30s_i8_stateful__mac-studio-m4-max.json"
      }
    ]
  },
  "model": {
    "family": "parakeet",
    "id": "parakeet-tdt-0.6b-v3__parakeet_tdt_0.6b_v3_30s_i8_stateful",
    "license": "cc-by-4.0",
    "source_url": "https://huggingface.co/mlboydaisuke/Parakeet-TDT-0.6B-v3-LiteRT",
    "task": "automatic-speech-recognition"
  },
  "pitfalls": [
    "Three signatures: encode (log-mel [1,128,3000] -> [1,1024,375]) + decode (4-token stateless prefix [1,4], LSTM state 2 x [2,1,640]; logits [1,375,4,8198]) + decode_1 (single step [1,1]; logits [1,375,1,8198]); 123 subgraphs (measured with the repo parser and ai-edge-litert 2.2.0).",
    "Why 30 s: through LiteRT-LM omni/asr every utterance is cut into standalone windows of the export's length, and the original model returns no text for 37 % of standalone 5 s windows that start mid-utterance (267 of 724, NeMo fp32, LibriSpeech test-clean), so the 5 s file scores 17.67 / 18.25 WER (test-clean / test-other) against NeMo's 1.92 / 3.59. With a 30 s window a LibriSpeech utterance is one window (measured 2026-09-25).",
    "Through LiteRT-LM omni/asr at main@66058c82 this file needs three patches (edge-llm-bench tools/omni-eval/patches): 01 Slaney/power mel, 03 a parakeet-tdt-0.6b-v3-30s metadata entry (inputMilliseconds 30000, nFrames 3000, this file's URL), 06 TdtDecoder stops at the last encoder frame that carries audio, with a 10-symbols-per-step cap. Without 06 a short clip is decoded over up to 28 s of zero padding and the TDT decoder adds text there: 2.69 / 10.54 WER; with 06: 2.17 / 5.35 (Mac CPU, 4 threads, Open ASR Leaderboard scoring; measured).",
    "Cost of the long window: every utterance pays a 30 s encoder pass, RTFx 17-19 on an M4 Max CPU against 34 for the 5 s file. What remains at 2.17 / 5.35 is clips under 5 s on test-other (10.2 % WER, 44 of 1419 empty), the model's own behaviour on short standalone input (measured).",
    "Mel front-end must match NeMo: Slaney scale, power spectrum, floor 5.96e-8, per-feature normalization over the whole padded window. The export was trained on padded batches without a length mask, so normalizing over the valid frames only and zero-filling the rest sends short clips from 2.40 to 11.42 WER on the 5 s file (measured).",
    "Engine note, not the model: through omni/asr on Metal (Mac GPU) RSS grows about 10 MB per session on the 5 s file (decoder output buffers not released); reuse one session for a batch. The 30 s file was measured on CPU only."
  ],
  "schema_version": "1.2"
}
