How to Prompt a Voice Agent: Writing for the Ear (2026)
Master the art of conversational AI by learning specialized prompting techniques designed for auditory clarity and natural voice agent interactions.
Summary
Writing for the Ear
When prompting a language model for voice conversations, the instructions must account for how people listen rather than read. The core instruction for any voice system prompt should be: "You're speaking, not writing."
Key Rules for Voice Prompts
- Keep turns short: Limit responses to one idea in one or two sentences, then let the caller speak. Long answers risk being interrupted or forgotten. Ask only one question at a time.
- Avoid formatting: Instruct the model to avoid bullet points, bold text, emojis, tables, and links. For lists, use spoken transition words like "first," "then," and "finally," keeping the list to about three items. Read codes slowly, one character at a time.
- Format numbers and dates clearly: Avoid ambiguous date formats (like "4/11") and have the model say the full date (e.g., "November fourth"). Depending on the voice provider, either write out numbers in words or use normal digits and rely on text normalization.
- Account for messy transcripts: In cascaded pipelines, the model receives live transcripts from speech recognition, which often contain errors. Instruct the model to ask for clarification if something is unclear and to repeat names and numbers back to the caller before using them. Use an idle timer to handle silence politely.
- Manage tool calls and silence: Looking up information takes time. Have the agent fill the silence by saying something like "I'll check that now" before executing a tool. Once the result is retrieved, present only the necessary information rather than reading raw data fields. Always confirm details before executing actions that change state, such as bookings or payments.
- Establish guardrails: Define the agent's scope and instruct it to politely steer the conversation back or hand off to a human if a request is out of bounds. The agent must never pretend to be human and should state that it is an AI at the start of the call. Critical restrictions, such as preventing unauthorized refunds, must be enforced in the code rather than relying solely on prompts.
Related Content

How to Test a Voice Agent: Simulated Callers and Evals (2026)

Why Speech-to-Text Gets It Wrong, and How to Fix It (2026)

How Speech-to-Text Works: From Sound Waves to Words (2026)

Cascaded vs Speech-to-Speech Voice Agents, Explained (2026)

Turn-Taking in Voice AI: How Agents Know When to Talk (2026)

Voice AI Latency Explained: Where the Time Goes (2026)

SIP Explained: How Phone Calls Reach AI Voice Agents (2026)

How WebRTC Works, Explained Simply (2026)