Sales Automation

    Conversation Intelligence: What the Recording Layer Can and Cannot Tell You

    Talk ratio is arithmetic. Sentiment is a guess about words. And the biggest blind spot is the calls that never happened, which no recorder can see.

    Editorial illustration for Conversation Intelligence
    August 23, 2026Updated August 18, 20267 min read
    Share:
    The short answer

    Conversation intelligence records calls, transcribes them, and analyses the text. Talk ratio, call counts and tracked-phrase detection are measurements. Sentiment, deal risk and suggested next steps are inferred from the transcript. Nothing in the layer sees the conversations that never happened, which is usually the binding constraint on a quarter.

    Key takeaways

    • Microsoft's Dynamics 365 conversation intelligence FAQ states the system does not process mono-channel recordings and supports stereo files only, so the telephony stack decides whether the product can ingest anything at all.
    • Competitor-mention reporting is a lookup against trackers you supply at onboarding, so a rival nobody listed returns a zero that reads exactly like genuine absence.
    • Microsoft documents that sentiment is generated from the transcribed text rather than from the audio, which means tone, pace and pauses are not what is being scored.
    • The layer is blind to coverage: it observes conversations that happened, and says nothing about the accounts nobody reached.

    Reviewed and updated August 18, 2026

    Microsoft's conversation intelligence for Dynamics 365 Sales will not process a mono-channel call recording at all. Its FAQ, last updated 7 July 2026, says the system "doesn't process mono-channel call recording files. It only supports stereo-type call recording files," meaning one speaker per channel. That single requirement explains most of the category: almost everything conversation intelligence sells is downstream of knowing, with certainty, which of the two people was talking.

    Conversation intelligence is the layer that records a call or meeting, turns the audio into text, and runs analysis over that text: who spoke for how long, which topics came up, which competitors were named, what the tone of the language was, and what should happen next. Every major seller of it publishes the same list of components. Gong's own guide names automatic recording and transcription, keyword and topic detection, sentiment and talk-ratio analysis, deal tracking from conversation data, and integration with the CRM.

    The useful question for a team buying it is narrower than that list. It is which of those outputs are measurements and which are estimates, because the two get presented on the same dashboard in the same font.

    The pipeline, and where each output comes from

    1. Step 1Capture

      The call or meeting is recorded, ideally with one speaker per audio channel

    2. Step 2Transcribe

      Audio becomes text, with speaker labels attached from the channel split

    3. Step 3Detect

      Trackers, topics and competitor mentions are matched against the transcript

    4. Step 4Derive

      Sentiment, talk ratio, deal risk and next steps are computed from the text and the metadata

    The four stages of a conversation intelligence pipeline. Only the first two are measurement; the third and fourth are derived from text.

    Stage one and stage two are close to measurement. A recording either exists or it does not, and a transcript is checkable against the audio by anyone who wants to listen.

    Stage three is a lookup against a list you supplied. Microsoft's onboarding documentation is unusually direct about this: the documentation states plainly: "You must also share trackers that you care about, along with your competitive brands" and products. A tracker is a phrase the system watches for. This matters more than it sounds, because the report titled competitor mentions is a count of the competitors you thought to list. A rival nobody in your building has heard of yet produces a zero, and the zero looks identical to genuine absence.

    Stage four is inference. Microsoft's FAQ states how sentiment is produced: the system "transcribes the calls into text and generates sentiment from the text in the conversation." Sentiment is therefore a judgement about words, not about voices. A buyer who says "that sounds great" flatly and a buyer who says it with real enthusiasm produce the same string, and the pause before the sentence produces nothing at all.

    What the recording layer can tell you well

    Four things, and they are worth having.

    Who talked. Talk ratio is the most robust output in the category because it is arithmetic on timestamps rather than an opinion about language. It is also the output managers most often ignore in favour of the softer ones.

    What was actually said, as opposed to what was logged. The gap between a rep's CRM note and the transcript is frequently the most valuable thing on the platform, and it needs no AI at all to read.

    Whether a named thing came up. Given the tracker list is maintained, phrase detection is reliable, and it answers questions like whether pricing was raised before the demo or after it.

    Where a call is in relation to other calls. Aggregation across hundreds of conversations makes patterns visible that no individual reviewer would notice, which is the genuine coaching argument for the category and the reason Gong prices the way it does.

    What it cannot tell you, however good the model is

    Section illustration: What it cannot tell you, however good the model is

    Anything about the calls that did not happen. This is the largest blind spot and it is structural. A recording layer observes conversations that reached the stage of being a conversation. The accounts nobody reached, the dials that ended in voicemail, the buyer who never took the meeting: none of that is in the corpus, and coverage is usually the thing actually constraining the quarter. That half of the problem lives in the dialling and sending layers, and the arithmetic behind a calling programme is where it gets measured.

    Whether the reason given is the real reason. Loss reasons stated on a call are self-reported by a person who often has an interest in a comfortable answer. The transcript records the sentence accurately, and the sentence may still be untrue.

    Whether a coaching change caused a result. Platforms present coaching insights next to outcome metrics, and the correlation is real. The causal claim needs an experiment nobody in a live sales team ever gets to run cleanly, because the reps who adopt the coaching are usually the reps who were going to improve anyway.

    Whether the deal is real. Deal-risk scoring reads conversation signals. It cannot see the buyer's internal budget conversation, which happens in a meeting nobody in the vendor's account was invited to. A mutual action plan is a cheaper instrument for that specific question because it asks the buyer to commit to something in writing.

    MeasuredCheckable against the recording
    • Talk and listen ratio
    • Call duration and count
    • Whether a tracked phrase occurred
    • Who was on the call
    • What the transcript says versus what the CRM note says
    EstimatedDerived from the transcript by a model
    • Sentiment and buyer enthusiasm
    • Deal risk and pipeline health scores
    • Suggested next steps
    • Topic categorisation beyond exact trackers
    AbsentNot in the corpus at all
    • Conversations that never happened
    • The buyer's internal discussion
    • Whether the stated objection is the actual one
    • Accounts nobody contacted
    Which dashboard outputs are measurement and which are estimation. Both appear in the same interface.

    The operating constraints that decide whether it works

    These sit in documentation rather than on pricing pages, and they change implementation plans.

    Channel format. The stereo requirement above is not a preference. A team recording through a system that produces a single mixed channel has bought a product that cannot ingest its data, and finding that out during rollout is expensive.

    Warm-up volume. Microsoft's FAQ states you need "at least 10 call recording files" before the system will process and display data. Aggregate analysis needs a corpus, so the first weeks of any deployment produce dashboards that mean very little.

    Refresh lag. The same FAQ says conversation intelligence data "can take up to 12 hours to appear in the app." That rules out same-day coaching on this particular implementation, and it is the sort of detail that only appears once a manager asks why this morning's call is missing.

    Retention. Microsoft states the system "deletes call recordings as soon as it processes the audio file," and offers a choice between Microsoft-provided storage and your own Azure storage for the insights. Retention policy varies sharply by vendor, and it is a compliance question rather than a feature question, so it belongs in the security review rather than in the demo.

    Consent and disclosure. Recording law is jurisdictional and applies to the call rather than to the software. The tool does not make the call lawful, and any programme recording externally needs its own legal read rather than a vendor assurance.

    Verify before signing, in the order that saves the most money
    • Yes: Confirm your telephony or meeting stack emits two-channel stereo recordings
    • Yes: Ask which outputs are computed from the transcript and which from the audio
    • Yes: Get the tracker maintenance owner named, in writing, before rollout
    • Yes: Read the retention and storage terms with whoever owns your security review
    • Depends: Establish the data refresh interval before promising same-day coaching
    • No: Assume competitor-mention counts are complete
    • No: Treat sentiment scores as a measurement of buyer feeling
    The buying checks that separate a working deployment from an expensive dashboard.

    The manual read the platform will not do for you

    Section illustration: The manual read the platform will not do for you

    One habit is worth building alongside any deployment, and it costs an hour a week.

    Pick five transcripts at random, not five the platform flagged, and read them end to end against the CRM record for the same opportunity. The exercise answers three questions no dashboard answers: whether the stage recorded in the CRM matches what the buyer said, whether the objection logged is the objection raised, and whether the next step agreed on the call was ever written down anywhere a colleague could find it.

    Teams that do this find the same thing repeatedly, which is that the platform was right about the words and wrong about the meaning, and that the correction is a management conversation rather than a configuration change. It also keeps the tracker list honest, because the phrases a real buyer used this month are visible in a way that a quarterly review of the tracker configuration never makes them.

    The random sample matters more than the sample size. A platform surfaces the calls its model found interesting, and reading only those measures the model rather than the pipeline.

    The category is consolidating into other products

    Standalone conversation intelligence is increasingly a feature of something larger. ZoomInfo runs it inside a data platform, the CRMs ship their own, and dialler vendors bundle a version with the calling seat. The practical consequence is that the buying decision is rarely a comparison of recorders. It is a decision about which system of record should own the transcript, and the answer usually follows whichever platform already owns the meeting.

    That also means the price is rarely a line item you can read off a page. Gong quotes rather than publishes, which is why what the quote contains is the useful question to prepare for.

    Where this sits in an outbound programme

    Section illustration: Where this sits in an outbound programme

    RevenueFlow runs cold email and LinkedIn, not phone, so we do not operate a recording layer and everything above is described from vendor documentation rather than from our own use. The equivalent instrument in a written channel is the reply itself, which arrives already transcribed and needs no analysis layer to read.

    The one point worth carrying across channels is the coverage blind spot. A recorder makes existing conversations legible, and it does nothing about the number of them. Teams that install one and find the dashboards thin usually have a top-of-funnel problem being displayed as a coaching problem. If that is the shape of it, we will build the first campaign and you can compare the volume of conversations against what the recorder is currently seeing.

    Platform behaviour verified against Microsoft's Dynamics 365 conversation intelligence FAQ (page last updated 7 July 2026) and Gong's conversation intelligence guide, both fetched 17 August 2026. Vendor documentation changes; verify current terms before relying on them.

    Questions

    Frequently asked questions.

    Frequently asked questions
    What does conversation intelligence software actually do?
    It records a call or meeting, converts the audio to text, and runs analysis over that text. Gong's own guide lists the standard components as automatic recording and transcription, keyword and topic detection, sentiment and talk-ratio analysis, deal tracking from conversation data, and integration with the CRM and sales engagement tools.
    Is conversation intelligence sentiment analysis accurate?
    It is accurate about language rather than about feeling. Microsoft's FAQ states that the system transcribes calls into text and generates sentiment from the text in the conversation, so a flat delivery and an enthusiastic one produce the same string. Treat the score as a signal to go and listen, not as a measurement.
    What do I need before conversation intelligence will work?
    Two-channel stereo recordings, a maintained tracker list, and enough volume to aggregate. Microsoft's documentation requires stereo files, asks for your trackers and competitive brands at onboarding, and states you need at least ten recording files before data will process and display. Refresh can take up to twelve hours.
    Does conversation intelligence keep the recordings?
    It depends on the vendor, and it is a security-review question rather than a feature question. Microsoft states that its system deletes call recordings as soon as it processes the audio file, and offers a choice between Microsoft-provided storage and your own Azure storage for the resulting insights. Other vendors retain audio far longer.
    Conversation IntelligenceSales AutomationSales ToolsCall RecordingSales Coaching
    Byline

    About the author.

    RevenueFlow Team

    B2B cold email experts helping companies generate qualified leads through done-for-you outreach campaigns.

    RevenueFlow Team

    Your next move

    Ready to scale your outreach?

    We build GTM engines that book real meetings. See the receipts.

    Further reading

    Related articles.

    Sales Automation

    Gong as a Sales Tool: What It Does and What the Quote Will Contain

    Gong publishes no price, but its pricing page publishes the model: per-user licences plus a platform fee, integrations free. How to turn that into a comparable number.

    7 min readRead →
    Sales Automation

    Fathom AI Note Taker: What It Records, and Where the Notes Should Land

    Fathom records, transcribes and summarises calls, and hides a conversation intelligence product on its upper tier. What it publishes, and the decisions to make first.

    7 min readRead →
    Sales Automation

    Avoma Pricing: Recorder Seats, Free Viewers, and the Add-Ons

    Avoma charges only for people whose meetings get recorded and makes viewers free, so the seat count to price is smaller than the team and the add-ons are larger.

    7 min readRead →
    Sales Automation

    Agentforce Sales Coach: Which Editions Carry It, and the Four Dependencies Underneath

    The Salesforce sales coach is a product, not a person. What Agentforce Sales Coach critiques, which editions carry it, and the four features it needs first.

    7 min readRead →
    Sales Automation

    Chili Piper Pricing: The Seat Floors and What Each Extra Rep Costs

    Chili Piper publishes its floors, seat allowances and overage rates. What the $15,000 entry tier includes, and why the standalone calendar costs almost twice as much.

    7 min readRead →
    Sales Automation

    Salesforce Auto Dialers: Telephony Is Not on the Sales Cloud Rate Card

    Salesforce prices six Sales Cloud editions and no telephony. That absence makes a dialer a separate purchase whatever else you compare when you compare editions.

    7 min readRead →