{
  "artifacts": [
    {
      "file": "sarashina2.2-0.5b-instruct-v0.1_int8.litertlm",
      "sha256": "b1e7410d9f7b9c465e37da313ca4951eb012b3e022712ec5bfee11064dae11e3",
      "size_mb": 830.518
    }
  ],
  "benchmarks": [],
  "conversion": {
    "command": "convert_sarashina.py <export_dir> <out> templates/sarashina_simple.jinja <recipe> — NO_START_TOKEN=1, HF tokenizer.json embedded (vocab_file cleared so export falls through to save_pretrained legacy_format=False -> HF_Tokenizer_Zlib), CACHE 4096, prefill ladder 1024..1 = 11 signatures; no post-processing (FINDINGS Recipe)",
    "quantization": "export-time dynamic int8 on linears + embedding (dynamic_wi8_afp32); no post-processing",
    "tool": "litert-torch (released wheels only; repro = hf-to-litertlm sarashina_work/convert_sarashina.py wrapping scripts/export_simple_template.py, family driver with structured prompt_templates)",
    "tool_version": "0.9.3 (transformers 5.14.1; ai-edge-quantizer named but unversioned in the sources — HF card Conversion notes; FINDINGS env ~/venvs/ltconv040dev; Mac gates on the litert-lm 0.16 lineage)"
  },
  "cross_runtime": [],
  "delegation": null,
  "device": {
    "records": [
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-02",
          "decode_tokens_per_s": 21.04,
          "delegated_ops": 1292,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=348.77 decode=21.04 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 497.0,
            "init_s": 0.30043,
            "prefill_tokens": 321.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 348.77,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1348,
          "ttft_ms": 970.0
        },
        "source": "data/device_runs/0.16.0/2026-09-02/sarashina2.2-0.5b-instruct-v0.1-int8__galaxy-s26.json"
      },
      {
        "device": "galaxy-s26",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-02",
          "decode_tokens_per_s": 36.46,
          "delegated_ops": 1066,
          "env": {
            "device": "Galaxy S26 (SM-S942Q)",
            "machine_label": "galaxy-s26-cl-pinned",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Qualcomm SM8850",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=895.06 decode=36.46 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 643.0,
            "init_s": 2.01947,
            "prefill_tokens": 321.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 895.06,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1066,
          "ttft_ms": 390.0
        },
        "source": "data/device_runs/0.16.0/2026-09-02/sarashina2.2-0.5b-instruct-v0.1-int8__galaxy-s26.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-02",
          "decode_tokens_per_s": 47.05,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "mac-studio-m4-max",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": null,
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "results block: prefill=1029.39 decode=47.05 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 1029.39,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 529.7
        },
        "source": "data/device_runs/0.16.0/2026-09-02/sarashina2.2-0.5b-instruct-v0.1-int8__mac-studio-m4-max.json"
      },
      {
        "device": "mac-studio-m4-max",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-02",
          "decode_tokens_per_s": 197.1,
          "delegated_ops": null,
          "env": {
            "device": "Mac Studio (M4 Max)",
            "machine_label": "m4max-quiet-300s-gpu-rest",
            "os_build": null,
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": null,
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "results block: prefill=5722.09 decode=197.1 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": null,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 5722.09,
          "provenance": "measured",
          "runs": true,
          "total_ops": null,
          "ttft_ms": 58.2
        },
        "source": "data/device_runs/0.16.0/2026-09-02/sarashina2.2-0.5b-instruct-v0.1-int8__mac-studio-m4-max.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "cpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": 8.78,
          "delegated_ops": 1292,
          "env": {
            "device": "Pixel 8a",
            "machine_label": "pixel-8a-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Tensor G3",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 963 out of 1066 node(s) with delegate (TfLiteXNNPackDelegate) node, yielding 143 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=73.34 decode=8.78 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": false,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 497.0,
            "init_s": 7.18535,
            "prefill_tokens": 321.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 73.34,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1348,
          "ttft_ms": 4490.0
        },
        "source": "data/device_runs/0.16.0/2026-09-05/sarashina2.2-0.5b-instruct-v0.1-int8__pixel-8a.json"
      },
      {
        "device": "pixel-8a",
        "run": {
          "accelerator": "gpu",
          "context_length": null,
          "date": "2026-09-05",
          "decode_tokens_per_s": 16.79,
          "delegated_ops": 1066,
          "env": {
            "device": "Pixel 8a",
            "machine_label": "pixel-8a-cold-cache-cooled",
            "os_build": "Android 16",
            "runtime": "litert-lm",
            "runtime_version": "0.16.0",
            "soc": "Tensor G3",
            "vendor_sdk": null
          },
          "error": null,
          "evidence": [
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (prefill_1024).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 1 (prefill_512).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 2 (prefill_256).",
            "VERBOSE: Replacing 1066 out of 1066 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 3 (prefill_128).",
            "results block: prefill=297.4 decode=16.79 tokens/s"
          ],
          "failure_class": null,
          "full_delegation": true,
          "latency_p50_ms": null,
          "loads": true,
          "max_abs_diff": null,
          "max_rel_diff": null,
          "metrics": {
            "decode_tokens": 464.0,
            "init_s": 28.92351,
            "prefill_tokens": 321.0
          },
          "output_match": null,
          "peak_mem_mb": null,
          "prefill_tokens_per_s": 297.4,
          "provenance": "measured",
          "runs": true,
          "total_ops": 1066,
          "ttft_ms": 1140.0
        },
        "source": "data/device_runs/0.16.0/2026-09-05/sarashina2.2-0.5b-instruct-v0.1-int8__pixel-8a.json"
      }
    ]
  },
  "model": {
    "family": "sarashina2.2",
    "id": "sarashina2.2-0.5b-instruct-v0.1-int8",
    "license": "mit",
    "source_url": "https://huggingface.co/litert-community/sarashina2.2-0.5b-instruct-v0.1",
    "task": "text-generation"
  },
  "pitfalls": [
    "Quality: 8-question gate (Apple M4 Max, litert-lm 0.16 lineage, one process per question, greedy) EN/JA — see HF card Correctness table; JCommonsenseQA (own harness, n=100): bf16 65 / int8 62 / int4 65 — all within paired noise (bf16-only 4 vs int8-only 1 / int4-only 4) (HF card Accuracy; FINDINGS).",
    "The bundle embeds the HF tokenizer.json, not the vendor tokenizer.model: sarashina's SentencePiece model types every chat special (<|user|> <|assistant|> <|system|> </s>) as a CONTROL piece, which a bare sentencepiece encoder never matches from text (<|user|> becomes 5 pieces, no id 9/8/2); the HF fast tokenizer matches them as added tokens. Engine-vs-HF tokenize parity 7/7 probes on every bundle (HF card Conversion notes; FINDINGS Recipe item 1).",
    "No start_token is written: add_bos_token is false and the official template emits no <s>, but the exporter would still write start_token from tokenizer.bos_token; measured in bf16 with the same rendered prompt +/- BOS, the 0.5b drops EN 6/8->5/8 and JA 8/8->7/8 with a BOS (the 1b is unchanged), so NO_START_TOKEN=1 (HF card; FINDINGS Recipe item 2).",
    "Chat template = structured prefix/suffix pairs (<|system|>…</s>, <|user|>…</s>, <|assistant|>…</s>, generation prompt <|assistant|>) byte-identical to what the official jinja renders for plain chat shapes; the tool-calling branch is not carried (HF card; FINDINGS Template).",
    "int8 is the recommended file: closest to bf16 and, on both the Galaxy S26 and the Mac, not slower than int4 on either backend — the 102,400-entry untied vocab makes embedding + lm_head the dominant per-token cost at this size (262M of ~500M params on the 0.5b); int4 CPU prefill is 2-2.6x slower; int4 is the size option (HF card; FINDINGS S26 + Mac cardbench readings).",
    "Decode is far below LFM2.5-230M's 122 tok/s on the same S26 GPU for the same reason (untied 102,400 vocab; the 230M has a 65k tied vocab) (HF card Performance; FINDINGS).",
    "JCommonsenseQA numbers (JGLUE v1.3 valid, first 100, greedy, own harness scored by gold-choice-text match) are comparable only within that table, not to published JGLUE scores (HF card Accuracy; FINDINGS).",
    "The English 8-question misses are the checkpoint's own: the bf16 PyTorch reference scores 6/8 EN (misses 'merci' and the rhyme) and 8/8 JA; all legs non-degenerate; emoji / rare-kanji streaming probe (😀 𠮷 ☔) shows no U+FFFD on any leg (ByteFallback streams 4 tokens per emoji, so the streaming check matters) (HF card Correctness; FINDINGS Mac gates).",
    "Galaxy S26 (litert_lm_advanced_main v0.16.0 kit): every GPU leg fully delegated to LITERT_CL (1066/1066 on prefill_2-1024, 947/947 prefill_1, 975/975 decode, zero XNNPack fallback); all gate legs answered the Japanese capital question (HF card; FINDINGS Device gates).",
    "iPhone 17 Pro rows are not yet taken: the phone was 'unavailable' in devicectl and the BenchmarkApp install was refused with CoreDeviceError 4016 (screen lock) on 2026-09-02 — the card ships with an append-later note (FINDINGS Status).",
    "Mac cardbench (litert-lm 0.16.0 CLI, -p 256 -d 256 --runs 3 --cache no, quiet machine, 300 s rest before each GPU cell); the per-cell log carries no token-count header, so the device-run cells for the Mac carry no prompt-length condition (mac_cardbench.log; DECISIONS #163(e)).",
    "The 0.5b int8's one JA miss is code-switching ('一週間には seven days あります'), which bf16 does too with a BOS; the EN 7/8 CPU / 6/8 GPU misses match the checkpoint's own (FINDINGS Mac gates)."
  ],
  "schema_version": "1.2"
}
