Depth Anything 3 (ByteDance, Apache-2.0) is the
zoo’s first depth model: a DINOv2 ViT backbone + DPT-style dense head that predicts a relative
depth map from one RGB image. DA3 is an “any-view” model (1→N views); fed a single view (S=1) it
is a monocular depth estimator. HF: mlboydaisuke/Depth-Anything-3-CoreAI (small + base × fp16/fp32).
Conversion: conversion/export_da3.py. Sample: knowledge/scripts/depth_anything_3_sample.py.
da3-small = DINOv2 ViT-S (alternating cross-view attention + 2D RoPE + QK-norm + a camera
token) → DualDPT head (depth + confidence + a ray aux head). da3-base = ViT-B, same shape.
da3mono-large = ViT-L + plain DPT (depth only, no cross-view). We export only the depth path —
backbone + head → depth, depth_conf; the camera decoder, ray aux head and sky post-processing
are dropped (the ray branch is dead-code-eliminated by optimize() because only depth/depth_conf are
named graph outputs).
Why S=1 just works as a static graph: the cross-view global attention collapses to self-attention (s=1), the reference-view reorder is statically dead (it needs S ≥ a threshold), and the camera token is a fixed parameter — no data-dependent control flow survives.
Exported as ONE static graph: image [1,3,504,504] RGB raw [0,1] → depth [1,504,504] (+
depth_conf). R = 504 = 36×14 matches DA3’s default process_res; the DINOv2 pos-embed bicubic
interpolation is over fixed sizes so it folds to a constant at export (no runtime bicubic).
The ExportWrapper folds ImageNet mean/std into the graph (x = (image - mean) / std). The
runtime input must therefore be raw [0,1] RGB. Feeding an already-ImageNet-normalized tensor
double-normalizes and silently corrupts the depth.
This bug masqueraded as model/engine problems for a long time. A comparison harness that fed
normalized input to the engine produced: a fake cos ≈ 0.9 engine-vs-torch, “the engine output
looks noisy”, “non-square exports are broken (cos 0.9)”, “letterbox padding breaks attention”. All
of it was the double-normalization. With raw [0,1], the engine is cos 1.000000 vs torch at any
fixed shape — square AND non-square (verified on diverse images, relmax ~1e-5 fp32 / ~1e-2 fp16).
Lesson: when an on-device vision graph “looks subtly wrong”, first confirm where normalization lives
(in-graph vs host) before blaming the conversion. (CoreAIKit’s DepthEstimator uses an identity
preprocessor — ImagePreprocessor(mean: 0, std: 1) — for exactly this reason.)
The bundle is a fixed square graph. Host: resize the image to 504×504 (cv2 INTER_AREA), feed
raw [0,1], run, then resize the depth back to the original H×W. Depth is relative, so the brief
aspect squash is recovered by the resize-back.
Spectral colormap
(far = red, near = blue).In conversion/export_da3.py, applied at import:
RotaryPositionEmbedding2D sizes its cos/sin table by
int(positions.max()) + 1 — a Python int pulled from a traced tensor → a data-dependent guard
under torch.export. The grid is static, so we bake the length.torch.export
poisons those dicts with fake tensors, and a later eager run in the same process then dies
with GuardOnDataDependentSymNode. Recompute every call._add_pos_embed builds a fp32 sin/cos grid
(make_sincos_pos_embed hard-casts .float()); under fp16 it upcasts the feature map and the next
conv sees Half weights vs float input. Cast it back to the feature dtype.No GPU-delegate op workarounds were needed (unlike RF-DETR): bilinear upsample, 2D RoPE, SDPA and ConvTranspose all lower as-is.
.half(), not autocastDepth is robust: fp16 holds cos 1.000000 (≤ ~1% per-pixel). Two rules:
copy.deepcopy(wrapper).to(fp16), never wrapper.to(fp16) in place — the caller reuses the
fp32 module as the verify oracle, and .to() mutates in place.torch.autocast — under autocast LayerNorm stays fp32, so its output collides with the Half
conv that follows (Input type (float) and bias type (c10::Half)). Half the whole model, run Half.CoreAIKit feeds .float32(pixels) regardless; makeNDArray converts float32→fp16 from the input
descriptor, so the same Swift path drives the fp16 bundle.
| variant · dtype | params | size | parity | M4 Max GPU |
|---|---|---|---|---|
| small · fp32 | 34.3M | 105 MB | cos 1.000000 (cpu+gpu) | 17.7 ms · 56.5 FPS |
| small · fp16 | 34.3M | 54 MB | cos 1.000000, relmax 7e-3 | 15.2 ms · 65.7 FPS |
| base · fp16/fp32 | 135.4M | 202 / 402 MB | cos 1.000000 | 37.7 / 43.4 ms |
| mono-large · fp32 | 334.2M | 1.34 GB | cos 1.000000 | 118 ms |
small · fp16 is the on-device hero. App: apps/CoreAIDepth (photo + live
camera, iOS + macOS). Card: models/depth-anything-3/README.md.