← Library

Build · intermediate

Test a live voice agent before you trust the demo

Use ADK's live evaluation loop to test turn-taking, tool calls, handoffs and spoken conversations before shipping a voice agent.

A voice agent can sound excellent in one hand-picked conversation and still fail as soon as a caller interrupts it, changes their mind, or gives an answer in an unexpected form. The hard part is not making one demo work. It is collecting enough repeatable evidence to know when a change made the conversation better or worse.

On August 24, Google announced native live evaluation in the Agent Development Kit (ADK). The new loop can drive a live agent with a simulated user, turn that user’s messages into audio, score the conversation with rubrics, and keep the results inspectable in ADK Web. Read the announcement for the full feature walkthrough.

This tutorial turns that announcement into a practical starting point. You will use Google’s live_workflow sample, then adapt its evaluation shape to your own agent.

What you are actually testing

Do not begin with “does the voice sound natural?” That question is too vague to protect a release. A useful live evaluation checks the whole path:

  • Does the agent ask for information in the required order?
  • Does it call the right tool with the right value?
  • Does state survive a handoff between agents?
  • Does it recover when the user asks a side question or changes direction?
  • Does the spoken answer satisfy the intent, even when the exact words differ?

ADK’s sample is deliberately small: a greeter confirms a caller’s name, a second agent verifies a date of birth with a tool, and a final agent shares the appointment details. The three stages form a graph, so the test covers both conversation and orchestration. The sample’s records are mocked; keep any real personal or customer data out of a development evaluation.

Start from the official sample

You need Python 3.11+, uv, and credentials that can access both the Live API and Gemini text-to-speech. The sample and its README contain the current setup details.

  1. Clone the ADK repository. Open the live_workflow sample and clone the repository so the relative paths in the commands resolve.

  2. Install evaluation support. From the ADK Python checkout, run uv pip install -e “.[eval]” in the environment you use for the sample.

  3. Configure credentials. Add the required Vertex AI credentials as described by the sample README. Live evaluation needs access to the live model and the Gemini TTS model.

  4. Read the fixtures before running them. The important files are agent.py, live_workflow.evalset.json, and test_config.json. The test is easier to extend when you can see which behavior belongs in each file.

Build a test that can fail

An eval case should describe a situation, not praise the agent. The sample uses a conversation scenario with a starting prompt, a plan for the simulated user, and a persona. That lets the simulator decide natural turn-taking while still giving the test a clear destination:

{
  "eval_id": "caller_changes_their_mind",
  "conversation_scenario": {
    "starting_prompt": "Hi, I need help with my appointment.",
    "conversation_plan": "Confirm the caller's name. Give the date of birth when asked. Ask whether the appointment time can be changed before confirming any details.",
    "user_persona": "NOVICE"
  },
  "session_input": {
    "app_name": "live_workflow",
    "user_id": "test_user_id",
    "state": {}
  }
}

Use a scenario when you want to test the agent’s ability to lead a conversation. Use a fixed conversation when an exact turn sequence matters, such as a regression for a tool call or a safety boundary. Keep both kinds in the same suite: scenarios explore; fixed cases pin down behavior that must not move.

Set a turn limit with max_allowed_invocations. A simulated user should be allowed to be unpredictable, but an evaluation must still have a bounded cost and a definite end.

Turn on audio and judge intent

The test_config.json file connects the conversation to live audio. The key distinction is between the model that decides what the simulated user says and the audio model that speaks those turns:

{
  "criteria": {
    "rubric_based_multi_turn_trajectory_quality_v1": {
      "threshold": 0.7,
      "judge_model_options": {
        "judge_model": "gemini-3.7-flash"
      },
      "rubrics": [
        {
          "rubric_id": "verify_before_disclosure",
          "rubric_content": {
            "text_property": "The agent confirms identity and validates the date of birth before sharing appointment details."
          }
        }
      ]
    }
  },
  "live_model_config": {
    "timeout_seconds": 300
  },
  "user_simulator_config": {
    "type": "llm_audio",
    "model": "gemini-3.7-flash",
    "max_allowed_invocations": 10,
    "audio_model": "gemini-3.1-flash-tts-preview"
  }
}

The exact config can include voice and language settings as well; see the live workflow sample for the complete fixture. Test those settings when accent, pronunciation, or barge-in behavior matters to your product.

Natural-language rubrics are useful here because a correct spoken answer can have many valid phrasings. Judge the invariant, not one sentence. For example, “verified identity before disclosure” is a stronger rubric than “said the appointment sentence exactly.” Add a separate rubric for the tool execution if the order or arguments are business-critical.

Run the eval and inspect evidence

From the ADK Python checkout, run the sample command:

uv run adk eval \
  contributing/samples/live/live_workflow \
  contributing/samples/live/live_workflow/live_workflow.evalset.json \
  --config_file_path contributing/samples/live/live_workflow/test_config.json

The run generates evidence for the conversation, not just a final score. Open the results in ADK Web to inspect each turn’s transcript and playable audio. Listen for interruptions and awkward pauses, then compare those observations with the rubric result. A high score with an obvious turn-taking failure is a signal to improve the rubric or add a fixed regression case, not a reason to ignore the failure.

Once the local loop is useful, call the same evaluation programmatically with AgentEvaluator and run it in CI. The ADK evaluation guide also documents dataset generation, grading, comparison, and the eval-fix loop.

A small release gate

Before shipping a voice workflow, keep a short suite that covers these paths:

  1. The happy path completes with the expected tool calls and handoffs.
  2. The caller interrupts or asks an unrelated question.
  3. The caller changes an answer after the agent reads it back.
  4. The agent refuses or pauses before a protected action.
  5. A model or prompt change does not lower the score on a known regression.

Run the suite after changes to prompts, tools, model versions, voice settings, and workflow edges. Those changes can alter behavior even when application code does not move. Keep thresholds modest at first, review failures manually, and raise the bar when the suite becomes representative.

The takeaway

The useful part of Google’s announcement is not that an agent can produce a voice demo. It is that spoken, multi-turn behavior can enter the same inspect, score, fix, and rerun loop as ordinary software.

Start with one business-critical conversation. Make its failure observable with a rubric, cap the simulated turns, inspect the audio, and keep one fixed case for every regression you discover. That is enough to replace “it sounded fine to me” with evidence you can review before the next prompt or model change.

Sources

Knowledge retrieval

What are you working through?

Start typing to search every guide, video, tool and snippet.