Every argument in the chat-agent market is about the definition of a word. Zendesk bills $1.50 per automated resolution on committed volume and decides what counts using its own model after a 72-hour silence window. Intercom’s Fin bills $0.99 per outcome. Whether a conversation was “resolved” is a negotiation you re-open at every renewal.
Put the agent on the phone and that argument disappears, because nobody disputes what a minute is. Parloa bills per minute and per API call, and published the case for it in January 2026: a minute is measurable, a resolution is a claim. That is the honest meter. It is also the one that bills you for your own latency.
This stack is built around the two things voice changes. First, your invoice is a function of average handle time, so a verbose prompt is a budget line. Second, a caller does not go quiet for three days — they hang up. On chat, a bad handoff produces a mildly annoyed customer. On voice, it produces an abandoned call, and you paid for every second leading up to it.
The shape
- Parloa answers the phone. Voice is the product rather than a channel bolted onto a chat platform, which is the whole reason it is here: its integration surface is the CCaaS layer a phone-heavy org already runs — Avaya, Five9, Genesys, NICE, Twilio, Verint — plus direct SIP. The compliance set is built for regulated phone work: ISO 27001:2022, SOC 2 Type 1 and 2, PCI DSS, HIPAA, and DORA. DORA is the one that decides the deal if you are an EU financial-services buyer.
- Zendesk Contact Center is the human side of the line and the system of record. It adds native inbound and outbound calling, conversational IVR, real-time transcription and call recording on top of any Suite plan, at $83 per agent per month paid yearly, with telephony running on Amazon Connect. Every call the AI touches lands as a ticket with the transcript attached, because the audit you will eventually run against Parloa’s invoice is a ticket query, not a vendor dashboard export.
- Assembled staffs the humans against a forecast that already knows what the AI contained. This is the layer most voice deployments skip and then discover they needed. Containment moving ten points in a quarter changes how many humans you need and in which intervals — and on voice that surfaces as hold time within the hour, not as a backlog you work down overnight. Assembled forecasts and schedules in-house agents, BPO seats, and AI agents against one number, and pulls telephony data from Amazon Connect, Five9, Genesys Cloud, NICE, Talkdesk, and UJET.
- Solidroad grades the calls — Parloa’s and the humans’ — against one scorecard. It reads closed conversations out of the helpdesk and scores phone alongside chat and email, then splits the finding: a scorecard-and-prompt revision when the bot failed, a practice simulation when the person did. It does not sell the voice agent it is grading, which is the entire reason it holds this seat. When the vendor defining the outcome also writes the report card, the report card is not evidence.
Named handoffs
- Call arrives → Parloa attempts the intent → the transfer fires with state attached. Transcript, detected intent, verification status, and actions already taken land in the agent desktop. A caller asked to repeat their account number after four minutes with the bot is the failure mode this stack exists to prevent.
- Transfer lands → Zendesk Contact Center opens the ticket. Contained or escalated, the call exists as a record with transcription. An AI that resolves calls invisibly is an AI whose invoice you cannot check.
- Containment rate moves → Assembled re-forecasts → schedules change for the affected intervals. This is the loop that keeps a containment gain from turning into an abandoned-call spike.
- Call closes → Solidroad scores it against your policy, routing bot failures to a prompt change and human failures to coaching.
- Month closes → reconcile minutes billed against minutes recorded in Zendesk. Parloa’s meter and your call records are two independent counts of the same thing. Run them against each other before you pay.
The seam that decides the build
Do not assume these two products plug into each other. Parloa names Zendesk as an integration, but its named CCaaS and telephony list does not include Amazon Connect — and Amazon Connect is exactly what Zendesk Contact Center’s telephony runs on. The Zendesk integration is the CRM and helpdesk surface; the call path is a separate question.
That leaves three routes, and you need to pick one before signing anything: front the whole thing with Twilio, which both sides support; run Parloa against a CCaaS it does name and use Zendesk purely as the ticketing record rather than the phone system; or get Amazon Connect support committed in writing during the trial. Discovering this after signature is how a voice deployment loses a quarter.
Cost baseline
A 30-agent support org running 600,000 voice minutes a year:
- Zendesk Suite Team at $55 per agent per month: $19,800/year floor. Suite Professional is $115 if you need it for other reasons.
- Zendesk Contact Center at $83 per agent per month: $29,880/year.
- Minutes Blocks at $33 per agent per month, one block per agent covering 1,000 minutes: $11,880/year for 360,000 minutes of IVR, transcription, and recording. The trap is the granularity — additional blocks must be bought in the same quantity for every Contact Center user, so exceeding your allowance costs another $11,880, not a marginal top-up. Size this against your human-handled minutes, not your total.
- Amazon Connect telephony and AWS services: consumption-billed on top, separately.
- Assembled Pro at $45 per agent per month: $16,200/year. Vendr’s marketplace data across 49 purchases puts the median Assembled contract at $29,400/year, so list arithmetic here is close to reality.
- Solidroad: unpublished. Third-party estimates of $50 to $150 per user per month span $18,000 to $54,000/year at this seat count — too wide to approve. Negotiate against Zendesk’s Workforce Engagement bundle at $50 per agent per month, which is published and absorbs WFM too.
- Parloa: unpublished. Secondary buyer analyses put the practical entry near $300,000/year with average contract values above $350,000. Treat that as an estimate; Parloa confirms no number publicly.
The arithmetic that matters: everything except Parloa lands near $96,000/year, and Parloa is roughly three quarters of the total. This stack is a bet on one vendor, and the other three exist to keep that vendor honest. Parloa’s own ROI band starts above roughly 500,000 annual voice minutes for the same reason — below that, the per-minute commitment cannot be repaid by containment.
Variations and when to swap
- Swap Parloa for the agent attached to your existing CCaaS when voice is under 40% of your contact mix. You are buying a specialist for a problem you do not have, and the deployment cost alone will exceed the containment gain. The chat-first composition is the AI support agent stack.
- Drop Assembled below roughly 15 agents, where a schedule is a spreadsheet and one person can see the whole queue. Add it back the moment BPO seats enter the mix — vendor management is where it stops being a scheduling tool and starts auditing an invoice nobody else can check.
- Drop Solidroad and use Zendesk QA when you are already buying the Workforce Engagement bundle and your AI agent is Zendesk’s own. The rule for keeping Solidroad is narrow and specific: keep it when the agent being graded is sold by someone other than the grader, which is exactly the case here with Parloa. Third-party scoring across both is the support quality assurance stack.
What this stack does not replace
- It is not an IVR redesign. It automates intents; it does not decide which intents belong on a phone line at all. Half the calls worth eliminating should have been a self-service flow.
- It does not replace human capacity. If the agent contains 40%, your humans take the remaining 60% — and they now get the harder 60%, which raises average handle time on the human side and shows up in your Assembled forecast as more staffing, not less.
- It is not a knowledge base. Every layer assumes the answers are already correct. On voice the cost of a wrong answer is higher, because there is no thread to scroll back through.
- It is not a CCaaS migration. Nothing here replaces your telephony contract; it sits on top of it.
Watch-outs, each with a guard
- Per-minute billing charges you for latency and for a talkative agent. A 150-second resolution costs 40% more than a 90-second one at identical volume. Guard: instrument average handle time per intent from week one against your human baseline, fix the per-minute rate for the contract term, and write in a review trigger if AHT crosses a named threshold instead of trusting that tuning will happen.
- A volume floor you miss is the standard way a voice contract becomes shelfware. Parloa is priced for growth against reported ARR above $50M at a $3B valuation, which means multi-year terms and commitments in the first proposal. Guard: cap the first contract at one year and get shortfall rollover in writing rather than accepting a use-it-or-lose-it minimum.
- Abandonment is the voice equivalent of a silent chat, and it is faster. A caller who hangs up during a transfer was contained by nobody. Guard: track abandon rate during handoff as a paired metric alongside containment, and treat any handoff over about 20 seconds as a defect — see CSAT for the satisfaction side of the same measurement.
- Recording and transcription obligations are stricter on voice than on chat. Two-party consent jurisdictions and AI-disclosure rules apply to the call itself. Guard: put the disclosure in the opening prompt, not the terms of service, and confirm per-jurisdiction consent handling during the trial.
- A containment gain and an unchanged schedule produce longer hold times, not savings. Guard: make the Assembled forecast consume the containment rate as an input on a weekly cadence, and treat a containment change without a schedule change as an open incident. When a handoff goes badly, run the escalation RCA.
Match rules
Right pick when: phone carries the majority of your contact volume, you run past roughly 500,000 voice minutes a year, and the intents are repeatable — order status, account lookups, appointment changes, balance enquiries. It fits best at 20 to 100 support seats in industries where callers still call: insurance, utilities, healthcare, travel, financial services. The compliance set is the tiebreaker if DORA or HIPAA comes up in the first meeting.
Wrong pick when: voice is a minority channel, or your volume cannot repay a six-figure first-year commitment, or your calls are bespoke investigation rather than repeatable questions. Also wrong when nobody owns the knowledge base — that gap breaks this stack faster than it breaks a chat deployment, because a wrong answer spoken aloud is the one your customer repeats back to you.
If you can only do one thing: measure your current average handle time per intent before you talk to any vendor. It is the denominator of every number in this stack — the containment case, the staffing forecast, and the invoice — and it is the one number no vendor can produce for you.