Test an AI voice agent by checking the complete journey from a candidate's first sentence to the recruiter's next action. A pleasant voice is only one part of acceptance. The caller must be able to correct information, the agent must respect its limits, and the resulting record must reach the right person without invented details or false promises.
For a staffing agency, the useful question is: could our morning team work from this call without repairing it? Use the checklist below to answer that before routing public traffic to a new setup. It is an operational testing method, not a certification or a guarantee of results.
Set the pass conditions before making calls
Choose one call flow, such as evening enquiries about warehouse work. Write down what the agent may collect, which statements it may make, and what the recruiter should receive. The broader voice agent implementation guide covers pilot scope; this test checks whether the configured flow actually meets that scope.
For each test, record the scenario, expected outcome, actual outcome, defect owner and retest result. Use fictional candidates and controlled phone numbers. Do not copy genuine applicant records into a demonstration environment simply to make it look realistic.
Define critical failures separately from cosmetic issues. Losing a callback request, exposing another candidate's information or confirming an unapproved job offer should block that flow. An awkward pause may still need attention, but it should not carry the same severity as a lost enquiry.
Build a small set of difficult candidate calls
Corrections during the conversation
Have the caller change a fact: “I said Monday, but I can start Wednesday.” Check both the spoken confirmation and the saved field. A summary that contains both dates without identifying the corrected one leaves the recruiter guessing.
Repeat this with a phone number, preferred location and callback window. The agent should ask for clarification when necessary, not choose whichever answer seems convenient. Uncertainty must remain visible if the caller disconnects before resolving it.
Questions outside the approved brief
Ask whether accommodation is guaranteed or a particular shift is definitely available when neither has been confirmed. The expected answer should acknowledge the limit and arrange an appropriate human review. The test fails if the agent fills the gap with a confident promise.
Try the same request politely and insistently. “Another recruiter promised it” should not cause the system to treat an unverified statement as an approved condition. Record the claim as the caller's statement when it matters to follow-up.
Interruptions and incomplete answers
Interrupt a question, remain silent, speak from a moderately noisy room and end the call halfway through intake. Check whether the flow can recover without repeating the entire script or silently discarding information already captured.
A partial enquiry needs an explicit outcome. It may be insufficient for recruiter action, but it must not appear as a completed qualification. Also test a caller who simply wants a human; they should not have to finish unnecessary screening questions first.
Test languages independently
A successful English call does not establish that the Dutch or Polish version works. Ask fluent speakers to test normal phrases used by your candidates, including shift descriptions, place names and corrections. Spanish flows need their own review if they are included in the launch.
For a Dutch logistics desk, a tester might say “volgende woensdag” and then name the date. For a Polish-speaking desk, test whether “popołudniówka” is captured as the intended afternoon shift rather than general availability. These are test prompts, not assumptions about every candidate's vocabulary.
Check language changes halfway through a call. The expected result depends on the supported setup: continue accurately or offer an appropriate handoff. Never approve a language solely because the greeting sounds convincing. Inspect its saved fields and recruiter summary too.
Follow the record into the recruiter's queue
After every call, inspect the actual destination. Confirm the candidate details, language, unresolved question, next action and responsible desk. Do this in the interface the recruiter uses, rather than relying only on a provider's success message.
Use a controlled failure to test an unavailable destination or rejected CRM write. Agree how that test will be run with the implementer so it cannot disrupt live records. A caller must not hear “your callback is booked” if no booking or recoverable request exists.
Then replay a failed delivery. Check whether recovery creates one usable task or two competing tasks. If the setup cannot prevent duplicates automatically, document the manual reconciliation step and assign an owner before launch. These are acceptance requirements to discuss, not claims that every CRM provides the same controls.
Check transfer failure and closed-office behaviour
Run a transfer when a recruiter answers, when the line is busy and when nobody answers. Confirm the caller hears what is happening and is not left in an endless loop. Test the configured next-working-day rule with the office closed, including the time zone used for any promise.
Use your actual staffing rota when assessing the proposed fallback. A technically valid transfer to an unstaffed desk is still an operational failure. The escalation rules guide can help define the intended destination before this check.
Decide what prevents launch
Review the evidence with a recruiter and the implementation owner. For each critical defect, fix the cause and repeat the original scenario plus a normal call that uses the same part of the flow. A change to date handling, for example, deserves a fresh check in every supported language.
Keep a dated record of the tested configuration. If the script, routing or CRM mapping changes later, rerun the affected cases. Start live use within the tested scope and agree who can pause the flow if an unexpected problem appears.
Common mistakes include testing only cooperative callers, approving audio without inspecting records, averaging critical failures into a reassuring overall score, and treating a supplier's demo as proof of your own configuration. Each removes evidence from the launch decision.
Practical acceptance checklist
- Corrected information replaces the earlier answer clearly.
- Unsupported requests lead to an honest, usable handoff.
- Every launched language has been checked by a fluent reviewer.
- Partial calls and failed transfers have defined outcomes.
- Delivery failure and recovery leave visible, owned work.
- Critical defects have passed a documented retest.
AI JOB AGENCY's candidate intake service connects first contact with recruiter workflow. Bring a difficult fictional call and your expected CRM result to a workflow discussion so the acceptance criteria can be included in the scope.
FAQ
How many test calls are enough?
There is no universal number. Cover every supported outcome, language and failure path, then repeat cases where behaviour varies. A large set of easy calls cannot replace missing exception tests.
Should we measure call duration?
Yes, but interpret it alongside accuracy and the next action. A short call that loses a correction is not a successful intake.
Can our recruiters perform the tests?
Yes. Recruiters are well placed to assess useful handoffs. Include fluent language reviewers and the implementer for technical failure scenarios.
What if a CRM connection is not ready?
Test the agreed interim destination and manual ownership process. Do not approve or describe an end-to-end CRM flow that has not been demonstrated.
When should we test again?
After relevant configuration changes, a newly observed failure, or expansion to another language or desk. Keep the original cases so improvements can be compared against the same expectations.
