Core AI model zoo

RGBA-Image-2.1-Turbo — Core AI port of Qwen-Image-2.1-Turbo

RGBA-Image-2.1-Turbo is Qwen/Qwen-Image-2.1-Turbo (revision d65dbc9) converted to Core AI .aimodel graphs with coreai-torch 0.4.1 and coreai-core 1.0.0b2. On an Apple M4 Max (macOS 27.0 26A428, measured 2026-10-11) the 7B DiT runs a 1024² step in 4.72 s, so an 8-step image takes about 38 s of DiT time.

A Qwen3-VL-8B text encoder (text path only) conditions a 7B single-stream DiT (32 blocks, block-causal attention). A 64-channel VAE decodes to four channels, RGBA. The sampler is the checkpoint’s own: 8 FlowMatch Euler steps on a fixed sigma grid, no CFG. Ask for a transparent background in the prompt and the fourth channel is a real alpha.

Sizes: 256², 512² and 1024², square. The DiT and the text encoder have dynamic axes (see Graph contracts); the VAE has one graph per size. The upstream 1:1 preset of 2048² needs 16,384 image tokens, outside this DiT’s 64–4096-token axis.

What changed from RGBA-Image-2.1

The graph code and the host code are those of RGBA-Image-2.1. The differences, and why the rest is the same:

The Turbo DiT is a distilled checkpoint. On the same noise and prompt, the fp32 Turbo pipeline’s 8-step image and the fp32 base pipeline’s 40-step image are 21.75 dB (256²) and 23.35 dB (512²) apart (white-composited RGB PSNR).

Bundle

🤗 mlboydaisuke/RGBA-Image-2.1-Turbo-CoreAI

file what it is size
qi21t_dit_full_bf16_dyn_iofp32.aimodel the DiT: bf16 weights and compute, fp32 inputs and outputs 14.23 GB (13.25 GiB)
qi21t_dit_full_int8lin_dyn_iofp32.aimodel the same DiT with int8 Linear weights (one bf16 scale per 32 inputs), bf16 compute 7.56 GB (7.04 GiB)
qi21_encoder_dynL_w16a32_ids_iofp32.aimodel the text encoder: bf16 weights, fp32 compute; token ids in, embed_tokens inside 15.14 GB (14.10 GiB)
qi21_vae_{256,512,1024}_fp32.aimodel the VAE decoder, fp32, one per size 1.01 GB (0.94 GiB) each
host/ per-axis RoPE tables, scheduler.json (the 8-step grid), host_contract.md  
tokenizer/ the source’s processor/ files, unchanged  
config.json the source’s transformer/config.json, unchanged  
LICENSE, NOTICE the Qwen Research License and the change notice  

The two DiT bundles have one interface; a run uses one of them. On every gate below the int8 DiT stays inside the bar, at the bf16 DiT’s speed (within 0.1 %). Compiled, the bf16 DiT is 26.52 GiB and the int8 DiT 20.30 GiB.

Compile before you run. On macOS 27.0 (26A428) the Python runtime crashes on this DiT unless it is compiled ahead of time with --expect-frequent-reshapes (step 2 below). A plain ahead-of-time compile fails at the initial call (ANERegion.mm:414 … Code=-19); so did a just-in-time compile of the same graph with the base weights on 2026-09-25. The compiled copies need their own disk: 26.52 GiB for the bf16 DiT or 20.30 GiB for the int8, and 14.10 GiB for the encoder.

Use it

There is no Swift app for this model yet. The Python engine conversion/qwenimage21/pipeline_engine.py runs the whole loop on the three bundles: tokenize, encode, 8 DiT steps, decode, write a PNG. --turbo selects the 8-step grid.

# 1. the bundles (the tokenizer is in tokenizer/)
hf download mlboydaisuke/RGBA-Image-2.1-Turbo-CoreAI --local-dir RGBA-Image-2.1-Turbo-CoreAI

# 2. compile for the Mac GPU (once per bundle; the gates used --architecture h16c on an M4 Max)
cd RGBA-Image-2.1-Turbo-CoreAI
for m in qi21_encoder_dynL_w16a32_ids_iofp32 qi21t_dit_full_bf16_dyn_iofp32 qi21_vae_256_fp32 qi21_vae_1024_fp32; do
  xcrun coreai-build compile $m.aimodel --output aot/$m --platform macOS --architecture h16c \
      --preferred-compute gpu --expect-frequent-reshapes
done
R=$PWD; A=$PWD/aot

# 3. generate, from a checkout of the zoo
#    (Python with coreai-core 1.0.0b2, torch, numpy, tokenizers, pillow)
cd /path/to/coreai-model-zoo/conversion/qwenimage21
python pipeline_engine.py --turbo \
    --prompt "This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent." \
    --size 1024 --seed 42 --tag dragon \
    --tokenizer-json $R/tokenizer/tokenizer.json \
    --encoder $A/qi21_encoder_dynL_w16a32_ids_iofp32/qi21_encoder_dynL_w16a32_ids_iofp32.h16c.aimodelc \
    --dit $A/qi21t_dit_full_bf16_dyn_iofp32/qi21t_dit_full_bf16_dyn_iofp32.h16c.aimodelc \
    --vae $A/qi21_vae_1024_fp32/qi21_vae_1024_fp32.h16c.aimodelc
# -> _work/samples/dragon.png (RGBA) and dragon_rgb.png (composited on white)

For the int8 DiT, compile qi21t_dit_full_int8lin_dyn_iofp32 the same way and pass its .aimodelc as --dit. The prompt above is the upstream card’s recommended form for transparent images: “This is an RGBA image with transparency. {subject}. The image has alpha channel and the background is transparent.” The noise is torch.randn on a CPU generator seeded with --seed, the same draw the reference pipeline makes with a CPU generator and that seed.

To check the port against the fp32 reference, record the reference once and run the engine in oracle mode. capture_oracle.py runs the diffusers-main pipeline in fp32 on the CPU, 17 s of generation at 256² on the M4 Max. It needs the full Qwen/Qwen-Image-2.1-Turbo snapshot and a venv with diffusers main 85c9fa4 and transformers 5.17.

python capture_oracle.py --snapshot <Qwen-Image-2.1-Turbo snapshot dir> --model-id Qwen/Qwen-Image-2.1-Turbo \
    --size 256 --steps 8 --tag turbo256                    # -> oracle/turbo256/
python pipeline_engine.py --turbo --oracle oracle/turbo256 \
    --tokenizer-json $R/tokenizer/tokenizer.json \
    --encoder $A/qi21_encoder_dynL_w16a32_ids_iofp32/qi21_encoder_dynL_w16a32_ids_iofp32.h16c.aimodelc \
    --dit $A/qi21t_dit_full_bf16_dyn_iofp32/qi21t_dit_full_bf16_dyn_iofp32.h16c.aimodelc \
    --vae $A/qi21_vae_256_fp32/qi21_vae_256_fp32.h16c.aimodelc
# prints the latent corr vs the reference after every step, then the RGBA and white-composited PSNR

Which image model should I use?

The zoo’s Mac text-to-image models are not ranked; pick by trade-off. The two RGBA-Image models write RGBA natively, so they fit images that need an alpha channel.

  params sampler time @1024 precision
FLUX.2 klein 4B 4 steps, guidance-distilled (no CFG) ~11 s fp16 or int8
Z-Image-Turbo 6B 8 steps + CFG (16 forwards) ~70 s bf16, near-lossless
GLM-Image 16B (9B AR + 7B DiT) AR prior + 20-step DiT ~208 s int8
RGBA-Image-2.1 7B 40 steps, no CFG ~190 s bf16
RGBA-Image-2.1-Turbo 7B 8 steps, distilled, no CFG ~38 s bf16 or int8

Times are the ones each card reports, on an M4 Max. For the two RGBA-Image models they are the DiT steps at the warm median per forward: 40 × 4.757 s and 8 × 4.721 s. They leave out the encoder, the VAE and loading.

Graph contracts

Every graph has one function, main, and fp32 inputs and outputs except input_ids. Both DiT bundles have the DiT row’s contract.

graph inputs output
encoder input_ids [1,Lfull] int32, Lfull 16..512 hidden [1,Lfull,4096]: the last layer’s residual stream, before the final norm
DiT img_tokens [1,N,64], txt_feats [1,L,4096], timestep [1], txt_cos/txt_sin [1,L,64], img_cos/img_sin [1,N,64]; L 8..512, N 64..4096 vel [1,N,64]
VAE latents_packed [1,N,64], the sampler’s latent unchanged image [1,4,S,S], RGBA in [-1, 1]

The full contract, with the formulas a Swift host needs: host/host_contract.md.

Measured

M4 Max (128 GiB), macOS 27.0 (26A428), coreai-core 1.0.0b2, 2026-10-11. Bundles compiled ahead of time with --expect-frequent-reshapes (--architecture h16c), run with SpecializationOptions.default(). Speed runs held the Mac’s measurement window, which holds back other sessions’ heavy jobs. The same form timed in different windows moved by up to 1.6 %. The timings were taken on the bundles before their debug locations were stripped; the published files compute the same outputs bit for bit (Gates).

DiT speed (bench_dit.py, text L = 40, bf16 DiT):

size image tokens initial call s/forward (warm median of 5) 8 steps
256² 256 2.27 s 0.341 s 2.7 s
512² 1024 1.10 s 1.097 s 8.8 s
1024² 4096 4.72 s 4.721 s 37.8 s

The int8 DiT against the bf16 one in the same process (ab_dit.py, L = 40, 8 alternating pairs per size): 0.3409 / 1.0923 / 4.7808 s per forward against 0.3410 / 1.0926 / 4.7850 s.

One image end to end (pipeline_engine.py --turbo, bf16 DiT, inside the measurement window, compile cache already warm):

prompt, size encoder load / call DiT load / initial step / 8 steps VAE load / call total
dragon sticker, 512² 17.0 / 0.40 s 2.0 / 3.20 / 10.9 s 1.2 / 0.38 s 32.0 s
dragon sticker, 1024² 3.4 / 0.37 s 2.0 / 6.63 / 39.6 s 0.5 / 1.45 s 47.3 s
apple, 1024² 3.5 / 0.36 s 1.9 / 6.64 / 39.3 s 0.3 / 1.46 s 47.0 s

Load times vary between processes on this shared Mac: 1.9–43.3 s for the DiT and 2.5–55.8 s for the encoder across the gate runs. The cold load of a new .aimodelc adds a compile-cache entry the size of the compiled bundle (26 GB for the bf16 DiT, 20 GB for the int8).

Memory. Peak memory footprint 4.80 GB at 512² and 17.77–18.16 GB at 1024² (/usr/bin/time -l). Its “maximum resident set size”, 58.3 GB with the bf16 DiT and 51.6 GB with the int8, counts the mapped weight pages of the encoder and the DiT.

Fidelity against the fp32 diffusers Turbo pipeline (prompt “a red apple on a wooden table, studio lighting”, seed 1234, the same noise, 8 steps):

size bf16 DiT: white-composited RGB PSNR / final latent corr int8 DiT: the same alpha max|Δ|
256² 43.79 dB / 0.999961 41.76 dB / 0.999944 1/255
512² 42.45 dB / 0.999800 49.28 dB / 0.999976 1/255
1024² 42.50 dB / 0.999926 39.32 dB / 0.999873 1/255

The reference decodes with the Turbo bf16 VAE file and the bundles carry the fp32 weights; the band between the two (below) is inside these numbers. With the reference’s prompt embeddings in place of the encoder bundle, 256² scores 46.57 dB (bf16 DiT) and 42.56 dB (int8 DiT).

The VAE band. The same final latent decoded with the fp32 weights (these bundles) and with the Turbo pipeline’s bf16 copy:

size max|Δ| on [-1, 1] uint8 PSNR max levels, RGB / alpha bundle vs a fp32 decode of its own weights
256² 0.074 56.95 dB 10 / 1 max|Δ| 3.0e-5
512² 0.106 58.95 dB 14 / 1 max|Δ| 3.2e-5
1024² 0.105 58.15 dB 13 / 1 max|Δ| 4.2e-5

Gates:

Other forms tried (s/forward against the shipped bf16 DiT in the same window):

form result
int8 Linear weights, AOT with --expect-frequent-reshapes compiles in 301 s, passes every gate, −0.02 / −0.02 / −0.09 % at 256² / 512² / 1024²: ships as the second DiT
image axis fixed at N = 4096 −0.50 % at L = 40, +0.58 % at L = 18: the dynamic graph ships
--preferred-compute neural-engine 0 Neural Engine regions; runs on the GPU, outputs equal bit for bit, +0.01 / 0.00 / +0.72 %
plain AOT, no --expect-frequent-reshapes crashes at the initial call (ANERegion.mm:414 … Code=-19), dynamic and fixed image axis alike
text K/V computed once per image the 32 text tokens between L = 8 and L = 40 cost 1.59 % of a forward at 512² and 0.51 % at 1024²

Lessons

  1. A probe’s compile result does not carry to the full graph, either way. The 2-layer int8 probe with random weights fails the --expect-frequent-reshapes compile (Pass failed: MPSMemrefAllocFusion), on 2026-09-25 and again on 2026-10-11. The full 32-layer int8 DiT with the Turbo weights compiles in 301 s and passes every gate. On 2026-09-25 it went the other way: the probe ran where the full bf16 DiT crashed. Judge compilation on the full graph.
  2. A sibling checkpoint can carry a bf16 copy of a component. The Turbo vae/ is the base fp32 VAE rounded to bf16, so the fp32 bundles sit 0.074–0.106 from the Turbo pipeline’s decode, 7–10× the 1e-2 bar. Gate a bundle against a fp32 decode of the weights it holds, and report the rounding band beside it. Alpha is near-constant on an opaque image (0.976–1.000), so its correlation drops to 0.966–0.991 under that band: report alpha in levels, not correlation.
  3. Read a dB gap between two DiT forms next to a nudge control. The 8-step sampler turns a 1.25e-5 relative change of the prompt embeddings into 29.03 dB (512²) and 32.08 dB (1024²) between dragon stickers. The int8 and bf16 stickers are 34.42 and 28.99 dB apart, the same order. 78–88 % of the pixels whose alpha moves by more than 2 levels lie within 3 px of the silhouette: alpha max|Δ| on a sticker (227–255) measures the edge.
  4. --preferred-compute neural-engine can compile to 0 regions without an error. This DiT compiled that way has no Neural Engine region and runs on the GPU, equal bit for bit to the gpu compile. Count the *ANE_region* entries before timing an ANE form.
  5. An exported .aimodel carries the export machine’s paths. coreai-torch writes every op’s source location into main.mlirb, and grep -I skips that binary. coreai-torch’s strip_debug_info removes them without touching the ops or the weights (strip_bundle.py); the result is a new asset, so gate it again.

The lessons of the base port (the Neural Engine region, the encoder’s fp32 compute, the bf16 band): RGBA-Image-2.1. Port notes, every gate and the dead ends: knowledge/qwenimage21-port.md. Scripts: conversion/qwenimage21/.

Licence

The weights are Qwen Materials under the Qwen RESEARCH LICENSE AGREEMENT (release date 2026-09-20), included unchanged as LICENSE. A summary follows; the Agreement is what binds.

NOTICE begins with the attribution text §3(c) requires:

Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.