Find all the Voice AI Startup programs:
    Videoby Voice AI Space

    How Voice Agents Use Tools: Function Calling Mid-Call

    Learn to implement mid-call function calling, enabling voice agents to trigger external tools and APIs dynamically during natural human interactions.

    Summary

    How Voice Agents Use Tools

    A language model on its own can only write words and cannot perform external actions like opening a calendar. To enable these actions, developers use tools, which act like request forms with a name, description, and blanks (called parameters). When a user makes a request, the model fills in these blanks with arguments. This process is known as function calling or tool calling. The developer's code then executes the request, retrieves the result, and passes it back to the model to generate a spoken response.

    Managing the Silence Gap

    Because tool execution requires the model to run twice and relies on external systems, it introduces a delay. While humans typically respond in a fifth of a second, external systems can take several seconds, creating a gap of silence. Developers can manage this gap using three methods:

    • Preambles: A short spoken line before the tool runs (e.g., "Let me check that for you").
    • Progress sounds: Audio cues like quiet typing to indicate the system is still active.
    • Status updates: Spoken updates if the process runs long (e.g., "Still checking, the system is a little slow today").

    Alternatively, some platforms support running tools in the background while the conversation continues. This is suitable for tasks where the immediate next sentence does not depend on the tool's output, such as sending a confirmation text.

    Handling Interruptions and Safeguarding Actions

    If a caller interrupts while a tool is running, different systems handle the active request differently—some let it finish, some cancel it, and others notify the code. To secure actions that cannot be undone, developers should block interruptions during execution and confirm details before acting.

    Tools are split into two categories:

    • Lookup tools: Safe to repeat (e.g., checking order status).
    • Action tools: Not safe to repeat (e.g., making a booking or charging a card). These require confirmation first and should be designed to be idempotent, meaning duplicate requests do not result in duplicate actions.

    Addressing Voice-Specific Challenges

    Voice interactions introduce transcription errors, where misheard names or numbers result in incorrect tool arguments. To mitigate this, systems should read back critical details for confirmation, allow keypad input for numbers, and validate arguments within the code.

    Additionally, importing external text into a conversation exposes the system to prompt injection, where malicious instructions in the retrieved data attempt to override the agent's rules. Security measures include instructing the model to treat tool results strictly as information rather than instructions, applying the principle of least privilege by limiting each agent's tools, and verifying user permissions within the backend code rather than relying on the prompt.

    Standardizing Connections with MCP

    The Model Context Protocol (MCP) is an open standard that simplifies connecting AI models to external systems, acting like a universal port for tools. When using MCP, developers should keep tool descriptions concise to prevent bloated requests and implement an approval step for any tools that modify data.