Skip to main content
SpeechKit uses three strict modes. Choose the smallest mode that matches the task so the expected processing and output stay clear.

Dictation

Dictation turns speech into text. It does not perform LLM rewriting, execute utilities, or generate spoken replies. Use it for notes, forms, messages, and other direct transcription tasks.

Assist

Assist turns speech or text into one bounded result. That result may come from a deterministic replacement, a utility, or an LLM, and may optionally be read aloud. Use it when you want transformation or a single answer rather than a live conversation.

Voice Agent

Voice Agent is a realtime audio conversation. It is intended for interactive brainstorming, support, and follow-up questions where turn detection and spoken output matter.

Hands-Free and customization

Hands-Free controls activation, microphone capture, automatic end-of-speech, and optional speaker output over the three modes. Words teach SpeechKit terms to recognize; Replacements apply deterministic text, command, snippet, synonym, or template transformations.

Local and external providers

Provider choice determines the data path: SpeechKit does not make an external provider local. Review provider configuration and privacy terms before sending sensitive audio.

Install the Windows client

Download and verify the current public beta.