Packaged engagement · for AI teams shipping to iPhone

iOS On-Device ML
Production Sprint

2-week fixed scope: take your PyTorch / Hugging Face model and ship it on iPhone, in production. Apple Neural Engine, quantized, benchmarked, TestFlight-ready.

US$15,000 – US$25,000 / engagement · 2 weeks · async-first

Who it's for

You've trained a model in PyTorch or pulled one from Hugging Face. It works on your server. Now you need it running on iPhone / iPad — at the latency, thermal, and memory budget the product needs — and you don't have a senior Apple Silicon specialist on the team.

That's what this is for. I take your model and ship it on Apple Neural Engine in 2 weeks, with a benchmark you can sign off on.

What you get

Deliverables — fixed

  • Model conversion. PyTorch / HF / ONNX → CoreML, MLX, GGUF, or ExecuTorch — whichever path is fastest and most maintainable for your model family.
  • Apple Neural Engine optimization. Graph hand-tuning + quantization (FP16 / INT8 / 4-bit), KV-cache packing for autoregressive models, palettization where applicable.
  • iOS SDK or sample-app integration. Swift / SwiftUI wrapper, AVCapture / VisionKit / ARKit glue where relevant. Reusable in your existing Xcode project.
  • Benchmark report. Latency (cold + warm), thermal sustain, memory peak, tokens/sec or FPS — measured on iPhone 15 Pro and iPad M4. Written report you can send to your CTO / investor.
  • TestFlight-ready build. Signing, provisioning, build settings, privacy manifest — ready for App Store review.

Pricing & timeline

Engagement2 weeks fixed scope · async-first · JST + a few hours US-Pacific evening / EU morning overlap
PriceUS$15,000 – US$25,000 per sprint, scope-dependent (model family complexity, target devices, integration surface)
Payment50% on kickoff, 50% on benchmark report delivery
DiscoveryOptional 1-week feasibility study (US$5,000) — written go/no-go before committing to the full sprint
Smaller workA single model conversion, a latency investigation, or a benchmark on your target device — welcome as standalone pieces. Tell me the scope and I'll quote it.
StartSlot available within 1–2 weeks of agreement

Good fit / not a fit

👍 You'll get value

  • You have a working model in PyTorch / HF and need it on iPhone or iPad
  • You're an AI startup or research lab without an in-house Apple Silicon specialist
  • Latency, thermals, or ANE perf is the thing slowing you down
  • You want benchmark data before committing to a longer engagement
  • You're shipping LLM, VLM, ASR, CV, or diffusion to iPhone in 2026

👎 Skip this if

  • You only have a research idea, not a trained model
  • You want full-app development end-to-end (this is the inference layer, not the UI / UX)
  • Android-only, React Native-only, or Flutter-only deployment target
  • Strict CET / PT / ET overlap requirement — I'm async-first from JST

Why me

Built specifically for this work

  • Maintainer of CoreML-Models (1.8k+★, 170+ OSS repos) — the de-facto iOS Core ML model zoo. Ports of YOLO, SAM, LaMa, Tesseract, on-device LLM, Stable Diffusion.
  • Ex-Ultralytics (the YOLOv5 / YOLOv8 team) — mobile deployment and Apple Neural Engine optimization.
  • Upstream contributor to Hugging Face swift-transformers, PyTorch executorch, and Microsoft onnxruntime.
  • 10 apps shipped on the App Store under my own Developer account — SnapMeasure (LiDAR), super2x, AIPainter, Mask2Face, Memosh, AnimateU, Blur., others.
  • Currently solo-lead on production iOS apps combining ARKit + LiDAR + YOLOv8 + on-device OCR + on-device LLM/VLM at 30 FPS on iPad / iPhone. Detailed case studies under NDA.
  • 7+ years iOS / on-device ML engineering.

Stack

CoreML coremltools MLX llama.cpp / GGUF ExecuTorch ONNX Runtime TFLite PyTorch Hugging Face diffusers Apple Neural Engine Metal Vision VisionKit ARKit RealityKit Swift SwiftUI Python C / C++

Inquire

Email is fastest. Tell me what model and what target device — that's enough to scope.

✉ rockyshikoku@gmail.com

I reply within a day · async-first · plain text is fine.