Packaged engagement · for AI teams shipping to iPhone
iOS On-Device ML
Production Sprint
2-week fixed scope: take your PyTorch / Hugging Face model and ship it on iPhone, in production. Apple Neural Engine, quantized, benchmarked, TestFlight-ready.
US$15,000 – US$25,000
/ engagement · 2 weeks · async-first
Who it's for
You've trained a model in PyTorch or pulled one from Hugging Face. It works on your server. Now you need it running on iPhone / iPad — at the latency, thermal, and memory budget the product needs — and you don't have a senior Apple Silicon specialist on the team.
That's what this is for. I take your model and ship it on Apple Neural Engine in 2 weeks, with a benchmark you can sign off on.
What you get
Deliverables — fixed
- Model conversion. PyTorch / HF / ONNX → CoreML, MLX, GGUF, or ExecuTorch — whichever path is fastest and most maintainable for your model family.
- Apple Neural Engine optimization. Graph hand-tuning + quantization (FP16 / INT8 / 4-bit), KV-cache packing for autoregressive models, palettization where applicable.
- iOS SDK or sample-app integration. Swift / SwiftUI wrapper, AVCapture / VisionKit / ARKit glue where relevant. Reusable in your existing Xcode project.
- Benchmark report. Latency (cold + warm), thermal sustain, memory peak, tokens/sec or FPS — measured on iPhone 15 Pro and iPad M4. Written report you can send to your CTO / investor.
- TestFlight-ready build. Signing, provisioning, build settings, privacy manifest — ready for App Store review.
Pricing & timeline
| Engagement | 2 weeks fixed scope · async-first · JST + a few hours US-Pacific evening / EU morning overlap |
| Price | US$15,000 – US$25,000 per sprint, scope-dependent (model family complexity, target devices, integration surface) |
| Payment | 50% on kickoff, 50% on benchmark report delivery |
| Discovery | Optional 1-week feasibility study (US$5,000) — written go/no-go before committing to the full sprint |
| Smaller work | A single model conversion, a latency investigation, or a benchmark on your target device — welcome as standalone pieces. Tell me the scope and I'll quote it. |
| Start | Slot available within 1–2 weeks of agreement |
Good fit / not a fit
👍 You'll get value
- You have a working model in PyTorch / HF and need it on iPhone or iPad
- You're an AI startup or research lab without an in-house Apple Silicon specialist
- Latency, thermals, or ANE perf is the thing slowing you down
- You want benchmark data before committing to a longer engagement
- You're shipping LLM, VLM, ASR, CV, or diffusion to iPhone in 2026
👎 Skip this if
- You only have a research idea, not a trained model
- You want full-app development end-to-end (this is the inference layer, not the UI / UX)
- Android-only, React Native-only, or Flutter-only deployment target
- Strict CET / PT / ET overlap requirement — I'm async-first from JST
Why me
Built specifically for this work
- Maintainer of CoreML-Models (1.8k+★, 170+ OSS repos) — the de-facto iOS Core ML model zoo. Ports of YOLO, SAM, LaMa, Tesseract, on-device LLM, Stable Diffusion.
- Ex-Ultralytics (the YOLOv5 / YOLOv8 team) — mobile deployment and Apple Neural Engine optimization.
- Upstream contributor to Hugging Face
swift-transformers, PyTorch executorch, and Microsoft onnxruntime.
- 10 apps shipped on the App Store under my own Developer account — SnapMeasure (LiDAR), super2x, AIPainter, Mask2Face, Memosh, AnimateU, Blur., others.
- Currently solo-lead on production iOS apps combining ARKit + LiDAR + YOLOv8 + on-device OCR + on-device LLM/VLM at 30 FPS on iPad / iPhone. Detailed case studies under NDA.
- 7+ years iOS / on-device ML engineering.
Stack
CoreML
coremltools
MLX
llama.cpp / GGUF
ExecuTorch
ONNX Runtime
TFLite
PyTorch
Hugging Face
diffusers
Apple Neural Engine
Metal
Vision
VisionKit
ARKit
RealityKit
Swift
SwiftUI
Python
C / C++
Inquire
Email is fastest. Tell me what model and what target device — that's enough to scope.
✉ rockyshikoku@gmail.com
I reply within a day · async-first · plain text is fine.