ooligo
ENTRY TYPE · definition

Conversation intelligence

By Marius Bughiu Last updated 2026-08-19 RevOps

Conversation intelligence is software that records sales calls and meetings, transcribes them, and runs analysis across the whole corpus of recordings — not just the one you were on. That analysis produces three things a recording by itself cannot: topic detection you can search across every call your team has ever had, deal-risk signals rolled up to the pipeline, and coaching scorecards tied to named rep behaviours. The recording is the input. The cross-call analysis is the product, and it is the only part worth paying enterprise prices for.

It is not an AI notetaker. Fathom, Granola, and Otter capture and summarize one meeting for the people who were in it, and the artifact ends at the meeting boundary. Conversation intelligence answers questions that only exist across hundreds of calls: how often reps hit the competitor’s new pricing last quarter, and what happened to those deals. It is also not revenue intelligence — conversation data is one input to a forecast, not the forecast — and it is not contact-center speech analytics, which scores support interactions against a QA and compliance rubric rather than against deal progression.

The four layers

Capture and transcription. A bot joins from the calendar invite, or the vendor captures through the conferencing provider’s API without one. Speaker diarization splits the transcript by participant. This layer is close to commodity — every vendor in the category clears it, and it is not where the price difference lives.

Topic and tracker detection. This is where a recorder becomes conversation intelligence. Gong splits it into two mechanisms: keyword trackers match the literal term, while smart trackers are models that detect the concept regardless of wording, so “is that the best you can do?” registers as a discount request without the word discount appearing (Gong Help Center). The distinction matters at buying time, because a keyword-only implementation looks identical in a demo and misses most of what it is supposed to catch.

Deal risk. The platform reads behavioural signals off the conversation and activity record — no mutual next step scheduled, single-threaded engagement with one contact, a buyer who has gone quiet, a competitor named late in the cycle, a stage that has stalled — and surfaces them against open pipeline rather than against individual calls.

Coaching scorecards. A manager scores a call against a rubric, and the scores aggregate by rep, team, and question. AI now fills these in: Gong’s automatic review scorecards review every call matching a filter, with a 90-day lookback ceiling, available to teams on its Enable Essentials or Enable plans (Gong Help Center).

What it costs

Two pricing shapes, and mixing them up is how budgets get missed.

Mid-market vendors publish a seat price that excludes the analysis. Avoma lists Startup at $19 per user per month annually, Organization at $24, and Enterprise at $39 with a 10-seat minimum — then sells Conversation Intelligence as a $29 per-seat add-on ($35 billed monthly), with Revenue Intelligence another $29 (Avoma pricing, checked 19 August 2026). A rep who needs the coaching and deal layers costs $48 to $77 per seat, not $19. Check every quote in this category for the same split.

Enterprise conversation intelligence is quote-gated. Vendr’s benchmark data puts the median Gong contract at $54,900 per year across 593 recorded purchases, with an observed range of $11,184 to $204,033 (Vendr marketplace, last updated March 2026). Ask for the quote broken out into platform fee and per-seat licences before comparing it to anything, and ask what the renewal uplift is in writing.

How accurate is it, really

There are two accuracy questions and buyers usually only ask the first one.

Transcription accuracy is measurable and mostly solved. A peer-reviewed comparison of 11 ASR services on English lecture audio measured word error rates from 2.9 percent (Whisper large-v2) to 20.1 percent (Google), averaging 7.0 percent — and 10.2 percent for non-native English speakers against 7.0 percent for native speakers (Measuring the Accuracy of Automatic Speech Recognition Solutions). Lecture audio is one speaker with a decent microphone. A four-way call with crosstalk and a laptop mic is harder, so treat any vendor accuracy figure as a claim about the easy case and test on your own recordings, weighted toward the accents and product vocabulary your team actually uses.

Analysis accuracy is the one that costs money, and the vendor documentation says the quiet part out loud. Gong’s own page on automatic scorecard review states that “sometimes, AI produces different outputs for identical inputs,” and calls it expected behaviour. A rubric score that moves between runs on the same call cannot be the input to a performance conversation without a human confirming it. Ask every vendor to score the same 20 calls twice and show you the spread.

Recording a call in a US all-party-consent state without every participant’s consent is a wiretap question, not a policy preference. Roughly a dozen states require all-party consent, and the rules are not uniform within them — Connecticut treats phone calls differently from in-person conversations, Oregon treats oral communications differently from electronic ones. For a distributed sales team the safe assumption is that some call every week has a participant in one of those states.

The live test case is In re Otter.AI Privacy Litigation, No. 5:25-cv-06911 (N.D. Cal.), consolidated from suits filed starting August 2025, alleging that the OtterPilot notetaker recorded meetings on the host’s consent alone under the federal Wiretap Act and the California Invasion of Privacy Act. Judge Eumi K. Lee heard the motion to dismiss on 20 May 2026 and had not ruled as of 19 August 2026, so no court has yet held these practices lawful or unlawful. Three guards do not depend on the outcome: turn on the vendor’s pre-meeting recording notice, give participants a decline path that actually stops the capture rather than just muting the summary, and read the clause governing whether your call data trains the vendor’s models.

What changed in 2026

Gong announced Gong Enable on 25 February 2026 and the Gong Revenue Harness on 24 June 2026 — an agentic execution layer built on its Agent Studio and MCP support, with Custom Agents that RevOps can configure without engineering (Gong press). ZoomInfo exposes Chorus call data through a conversation_intelligence tool in its MCP server, so an agent can reason over indexed calls, emails, and transcripts.

The direction is consistent across vendors: the call corpus is turning into a queryable data source for agents instead of a library managers visit to watch clips. That moves the buying question from whose transcript is cleanest to who will let your agents read the corpus, under which scopes — the read-side twin of the decision in MCP write access for CRM.

Common pitfalls

  • Buying it as a recorder and never configuring trackers. The platform records, nobody searches, and renewal arrives with usage data that justifies a downgrade. Guard: ship five trackers in week one tied to live deals — top competitor, pricing objection, security review, champion departure, procurement — and review them in the weekly pipeline meeting.
  • Trackers set once and never tuned. Guard: review tracker hit rates quarterly. A tracker firing on under 1 percent of calls is measuring nothing; one firing on over 60 percent is measuring the language everyone uses anyway.
  • Scorecards wired into performance management before calibration. Guard: score the same sample of 20 calls with two managers and the AI, and publish the agreement rate before any score reaches a review. Where they disagree, the rubric is the problem, not the rep.
  • Adoption measured by recording rate. Every call gets recorded on day one; that number proves the bot works, not that anyone learned anything. Guard: track searches run and clips shared per manager per week.
  • No retention or exclusion policy. Recordings of pricing negotiations, terminations, and legal discussions sit in the corpus indefinitely and become discoverable. Guard: set a retention window, and exclude meeting types by calendar keyword at capture time.