# Voice Interaction

Polyscope supports voice input for writing prompts, so you can talk to your coding agents instead of typing. There are two dictation engines: the browser's built-in Web Speech API, or a local Whisper model that runs entirely on your machine.

## Using Dictation

Click the **microphone** button in the prompt input toolbar to start recording. Speak your prompt, then click again to stop. Polyscope transcribes your speech and inserts the text into the prompt input.

<!-- ![Dictation button in the prompt toolbar](/images/docs/TODO.png) -->

The button shows a red highlight while recording and a spinner while transcribing.

## Dictation Engines

You can choose between two engines in **Settings > Dictation**.

### Web Speech

The default engine uses the browser's built-in speech recognition. It works out of the box with no setup — just click the mic button and start talking.

### Whisper (Recommended)

For faster, more accurate, and fully offline transcription, switch to the Whisper engine. Whisper runs a local AI model on your machine, so all audio processing stays on-device. On Apple Silicon it's hardware-accelerated with Metal for near-instant results.

#### Downloading a Model

Before using Whisper, you need to download a model. Go to **Settings > Dictation**, select **Whisper** as the engine, and pick a model size:

| Model | Size | Speed | Accuracy |
|-------|------|-------|----------|
| **Base** | ~142 MB | Fastest | Good for clear speech |
| **Small** | ~466 MB | Fast | Better accuracy |
| **Medium** | ~1.5 GB | Moderate | Best accuracy |

Click **Download** next to your chosen model. A progress bar shows the download status. Models are downloaded from Hugging Face and stored locally — you only need to download once.

<!-- ![Whisper model settings](/images/docs/TODO.png) -->

You can download multiple models and switch between them, or delete models you no longer need.

#### Why Use Whisper

- **Privacy** — Audio never leaves your machine
- **Speed** — Local inference is fast, and hardware-accelerated with Metal on Apple Silicon
- **Accuracy** — Whisper handles technical terminology and code-related language well
- **Offline** — Works without an internet connection

## Tips

- **Speak naturally** — Both engines handle conversational language well. You don't need to dictate punctuation.
- **Use voice for high-level instructions** — Voice works best for describing what you want ("add a login page with email and password fields") rather than dictating specific code.
- **Combine with typing** — Start with voice for the overall instruction, then type to add specific details like file paths or code snippets.
- **Start with Base** — The Base model is a good default. Switch to Small or Medium if you find transcription accuracy lacking.
