The speech design (SPEECH_API.md) came out of one observation: nothing is
live. Asking the same question of the other domains gives the same answer in a different shape,
and it is worth naming once because it decides most of the remaining work.
Every op accepts the unit the model accepts. Every app holds a larger unit. The distance between the two is the job that was supposed to disappear.
| Domain | Ops accept | The app actually has | What the adopter writes today |
|---|---|---|---|
| Speech | one file, or [Float] |
a microphone; a two-hour recording | capture loop, chunking, silence detection, partial results |
| Vision | one CGImage |
a camera; a photo library; a video | frame loop, throttling, orientation, mapping boxes into view space |
| Documents | one image | a multi-page PDF; a scan session | rasterising pages, ordering, stitching output |
| Text | one String |
a 50-page document; a chat history; a corpus | chunking, map-reduce, context budgeting |
Two sub-problems recur in every row: live (results as they arrive) and scale (input larger than the model’s window). Speech happens to need both at once, which is why it surfaced first.
CoreAIKitVision/CameraFeed.swift already vends AsyncStream<CGImage>, throttled to a target
frame rate. No op takes it. CoreAI.detect(in: image) is per-frame, so every adopter writes
the loop, decides the throttle, handles device orientation, and converts detection rectangles
into view coordinates — the last of which everyone gets wrong at least once.
for try await s in CoreAI.watch() {
s.detections // [Detection]
s.image // the frame they came from
}
watch() is to the camera what listen() is to the microphone: it owns permission, the
session, orientation, and interruption.
A Detection must carry a normalised rect (0…1 in the frame). Drawing a box over a preview
is then a multiplication, not a coordinate-system puzzle. If detections come back in pixel
coordinates of a rotated buffer, the job was not removed — it was renamed.
Depth has the same shape and the same live need:
for try await d in CoreAI.watchDepth() { … } // DepthCamera example does this by hand today
let summary = try await CoreAI.summarize(fiftyPageDocument)
This is already the code someone writes. It currently sends the whole string at the model and
exceeds its context — there is no chunking anywhere in Sources/CoreAIOps/.
The fix adds no API. It makes the existing one true:
The same applies to translate, proofread, extract and redact: all of them are given a
String today and all of them break at length.
CoreAI.read(documentAt: url) takes an image. Nothing in the package references CGPDF or a
page count. A real document is a multi-page PDF or a scan session, so the adopter rasterises
pages, runs each, and stitches the markdown — while the op’s name implies it already did.
Again: no new API, a truer one.
let markdown = try await CoreAI.read(documentAt: contractPDF) // all pages, in order
for try await page in CoreAI.readStream(documentAt: contractPDF) { … } // long documents
Most of it is not new API surface. Three of the four rows are existing calls learning to accept the real unit — which is the better kind of change, because nothing has to be discovered or documented for an adopter to benefit.
New verbs are needed only where the input is genuinely live and has no representation today:
| New | Why |
|---|---|
CoreAI.listen() |
the microphone has no op |
CoreAI.watch() |
the camera has a stream and no op |
*Stream variants |
progress on inputs large enough that the wait is visible |
And one requirement that is not an API at all: normalised coordinates on Detection, because
the alternative hands a coordinate-system problem back to the person we told would not need one.
summarize / translate / proofread / extract. No new
API, no new models, no device work, and it fixes calls that fail today.watch() — the stream already exists; this is wiring plus the coordinate contract.listen() + VAD — the largest demand, but it needs a new component and the iOS ports to
be affordable (see SPEECH_API.md).Nothing here needs a new model. Every part of it is the distance between what the models take and what apps hold.