What this is
A recording goes in; text comes out, split into segments with timestamps and, where the source separates them, speakers. It is the audio layer of the platform: a call recording, a voice message, a meeting.
It is not a separate product bolted on. Speech recognition is a model type in the same registry as chat and embedding models, called through the same single transport, charged against the same credits, and bound by the same provider policy. See Models and providers.
Who transcribes is a setting
Recognition speaks one widely supported contract, so the choice of engine is a routing decision, not code:
| Route | When |
|---|---|
| A hosted speech model | Fast to start, when audio may leave the contour |
| Any Whisper-compatible server | Your own recognition server inside your network |
A language hint is optional; without it the model detects the language. Fallback models apply as they do for chat: if the first route fails, the next one takes over.
Metered in seconds, not tokens
Audio is priced by its length. Credits are reserved on an estimate from the file size when the request starts and settled on the duration the model reports, so a long recording cannot overrun a budget unnoticed. It counts against the same budgets as every other model call.
Audio stays inside by default
Sending speech to an external provider is denied by default, separately from chat. A company that allows cloud chat models has not thereby allowed its calls to be sent to a cloud recognition service; that is a decision of its own. See Privacy and on-prem.
Transcription is also licensed separately from the rest of AI, so it is switched on where it is wanted and nowhere else.
Limits worth knowing
Audio files only, up to a size cap, with a timeout generous enough for a long meeting. One module in the platform is allowed to call a recognition endpoint, and a build-time test enforces it - the same rule that keeps every chat call on one transport. See The invocation contour.
Next
Transcription is the foundation for what is being built on top of it: an assistant that follows a sales call, a support conversation or a working meeting, and writes what it heard into the deal, the ticket or the project - uploaded recordings and voice messages first, then a live microphone. The platform's own speech service will join the routes above as a third option: recognition on local open models installed with the platform, sized to the installation's hardware.