Get started
Models we use
Exactly which speech and language models run on your computer and in Infina Cloud.
Once Infina is running, your dictation works on your computer without an internet connection. Dictation is processed entirely on the device — your audio and your words never leave your computer. Infina Cloud is an optional add-on that runs much larger models on our servers when you want fully polished text — grammar, punctuation, and formatting handled for you. Here is what powers each mode.
On your computer, included with your license
These models download once and then run fully on-device, with no internet connection needed once Infina is running, included with your Infina plan.
- Parakeet TDT 0.6B v2 turns your speech into text. It is NVIDIA's Parakeet model at 0.6 billion parameters. On a Mac it runs on the Apple Neural Engine; on Windows the same model runs on the laptop's own processor, which is why Windows asks for at least 4 processor cores and 8 GB of memory.
- Silero VAD v5 decides when you have started speaking and when you have stopped. It listens to the shape of the sound rather than the words, which is what lets Infina open on your first syllable and close as soon as you finish, instead of making you press anything. It runs on your computer alongside the others.
- Qwen3 0.6B is the small on-device language model behind the voice assistant and voice commands. Mac only. It needs Apple Silicon, which is why Windows matches spoken commands against a set list instead of working out what you meant.
- Parakeet Realtime EOU 120M is the on-device listener for hands-free mode. It is NVIDIA's real-time Parakeet model at 120 million parameters, built to follow speech as it happens and to notice when you have finished talking. It listens for your trigger words like type, send it, open and Infina, so nothing is sent anywhere until you speak one. On a Mac it runs with Core ML on Apple Silicon (macOS 14 or newer); on Windows it runs on the laptop's own processor. Dictation itself still uses Parakeet TDT above.
- pyannote segmentation 3.0 and WeSpeaker CAM++ are the two models that let hands-free mode tell your voice apart from every other voice in the room. The first works out how many people are talking and when; the second turns a voice into a fingerprint it can compare. Together they are what stops a video, a TV or someone sitting next to you from typing into your text, and what lets Infina stop when you stop even if the room carries on. Your voice fingerprint is built on your computer, stored on your computer, and never uploaded anywhere.
Infina Cloud, optional add-on
Some models are simply too big to run well on a laptop. With Infina Cloud they run on our servers instead, so your dictation comes back fully polished in every app, with grammar, punctuation and formatting handled for you.
- OpenAI's Whisper — its biggest, most accurate versions handle speech recognition. Whisper's best models are simply too large to run on a laptop, which is exactly why they live in the cloud.
- Llama, OpenAI and Qwen — large versions of these models clean up and format your text and power the assistant. They are far beyond what your own computer can run on its own.
Cloud is entirely optional. On-device dictation and the assistant come with your Infina plan, run locally and stay private, and you can switch between on-device and cloud whenever you want.