In this article

The AI Adoption Ladder for Revenue Teams: From One-Off Prompting to Autonomous Multi-Agent Systems

Written by
Ishan Chhabra
Last Updated :
September 22, 2026
Skim in :
12
mins
The AI adoption ladder for revenue teams title card spanning one-off prompting to autonomous multi-agent systems
In this article
Video thumbnail

Revenue teams love Oliv

Here’s why:
All your deal data unified (from 30+ tools and tabs).
Insights are delivered to you directly, no digging.
AI agents automate tasks for you.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Meet Oliv’s AI Agents

Hi! I’m,
Deal Driver

I track deals, flag risks, send weekly pipeline updates and give sales managers full visibility into deal progress

Hi! I’m,
CRM Manager

I maintain CRM hygiene by updating core, custom and qualification fields all without your team lifting a finger

Hi! I’m,
Forecaster

I build accurate forecasts based on real deal movement and tell you which deals to pull in to hit your number

Hi! I’m,
Coach

I believe performance fuels revenue. I spot skill gaps, score calls and build coaching plans to help every rep level up

Hi! I’m,  
Prospector

I dig into target accounts to surface the right contacts, tailor and time outreach so you always strike when it counts

Hi! I’m, 
Pipeline tracker

I call reps to get deal updates, and deliver a real-time, CRM-synced roll-up view of deal progress

Illustration of a person in a blue hat and coat holding a magnifying glass, flanked by two blurred characters on either side.

Hi! I’m,
Analyst

I answer complex pipeline questions, uncover deal patterns, and build reports that guide strategic decisions

TL;DR

  • Most revenue orgs are not behind on AI tooling. They are stuck between individual capability and shared capability, and that gap is organisational rather than technical.
  • The ladder has four rungs: individual prompting, individual reusable skills, centralised shared skills, and autonomous agents running in production under supervision.
  • The jump from rung two to rung three is a management change, not a skill. Someone has to own a shared library, and usually nobody has been assigned it.
  • Score five dimensions honestly: workflow standardisation, data quality and trust, inspection, cross-team alignment, and AI readiness. Your rung is your lowest score, never your average.
  • Rung three is the right ceiling for most companies. Rung four only earns its place when repeated judgement exceeds what a team can supervise by hand.
  • Stop reporting AI adoption percentage. Measure shared skill share, library lag, week-one workflows a new hire runs unaided, and agent correction rate.

Q1. Why should you trust a maturity model at all? [toc=1. Trusting the Model]

Most AI maturity models are sales artefacts. Four rungs, and the author's product sits on the top one. The test that survives that bias is a cost comparison, not a capability list. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. Below that line, coordination is pure overhead. Above it, inconsistency is already costing you deals.

⚠️ The board question nobody can answer

A CRO I spoke with last quarter had approved AI spend across three teams. Her reps were enthusiastic. Her SDRs had built prompts they swore by. Then her board asked a simple question: where are we, and what comes next?

She could name every tool. She could not name a stage. That gap is the reason this genre of article exists, and also the reason most of it is useless. It is the same gap I see when a team can list its stack but cannot describe how AI actually runs inside revenue operations.

❌ Why the genre deserves your suspicion

Here is the problem, said plainly. Maturity models are usually written by vendors, and the top rung is usually a description of the vendor's product. An experienced operator spots that in about four seconds, and rightly stops reading.

Believing one of those models has a real cost. You end up approving a transformation programme you cannot staff. You buy for a rung you do not operate at. Gartner's 2026 Hype Cycle placed agentic AI at the Peak of Inflated Expectations, with only 17% of organisations having actually deployed agents while more than 60% plan to within two years. That gap between intent and deployment is where budgets go to die, and it is why the build versus buy decision on revenue AI deserves more scrutiny than a maturity chart.

✅ The test that does not involve any vendor

So use a test that no software can tilt. Compare two costs.

  • Coordination cost. What it takes to agree a standard, write it down, and keep it current.
  • Inconsistency cost. What you lose when every rep runs a different version of the same task.

You belong on the rung where the first number is smaller than the second. That is the whole diagnostic. It works if you never buy anything.

Diagram comparing coordination cost against inconsistency cost to decide the correct GTM AI maturity rung
The diagnostic that survives vendor bias: you belong on the rung where coordination costs less than inconsistency.

⭐ What I will do differently here

Every rung in this article gets an honest reason to stay, not just a reason to climb. I will say the thing vendors do not: Rung 3 is the right ceiling for most companies. Shared standards, no supervision burden, no new role to hire.

I am also going to be specific about where each move gets expensive, because the expense is rarely the licence. AI spend gets approved per tool. Value accrues per operating model. That mismatch is how a company ends up buying at Rung 4 while still operating at Rung 2, with nothing in the numbers to show it.

⏰ What you should leave with

Two things, in about ten minutes of reading. A rung you can defend to a board, and one next move you can start on Monday.

Not a programme. One move. For most readers it turns out to cost nothing at all, which is the part I did not expect when I started mapping this.

Q2. What are the four rungs of the GTM AI adoption ladder, and what blocks each move up? [toc=2. The Four Rungs]

Four rungs. One, everyone prompts individually, context copy-pasted each time, nothing reusable. Two, individuals build reusable skills and run them themselves, which is the most common rung. Three, skills are centralised, teams use the same ones daily, and someone owns them. Four, agents run in production on documented process, with GTM engineers monitoring and optimising. The jump from one to two is a skill. The jump from two to three is a management change. The jump from three to four is documented process plus supervision capacity. Rungs cannot be skipped.

Two definitions before the table, because these words get used loosely. A skill is a saved, reusable instruction set a person runs on demand. An agent is a process that runs without a person starting it. A GTM engineer is the person who builds, monitors, and maintains those agents. If the distinction between the last two is still fuzzy, the difference between AI agents built for sales teams and a saved prompt is exactly what this ladder is measuring.

📊 The ladder, with the constraint that blocks each move

The GTM AI Adoption Ladder: Four Rungs, Constraints, and Next Actions
RungWhat is observably trueWho does the workBlocking constraintSingle next action
01. ChatGPT for everyonePeople prompt ad hoc. Context is pasted in fresh each time. Nothing is saved.Individual reps, unevenlyNobody knows which prompts workAsk three reps to share their best prompt in one doc
02. Claude and DIY skillsIndividuals build reusable skills and run them personally. Quality varies by person.Capable individualsNo owner for a shared libraryName an owner for the shared skill library
03. Centralised skillsOne library, agreed standards, teams use the same skills daily, someone maintains them.The team, on a standardProcess is undocumented, supervision capacity is zeroWrite the playbook as it actually runs
04. Autonomous agents in productionAgents run on documented process. Humans monitor, correct, and optimise.Agents, supervised by a GTM engineerNobody owns agent monitoringAssign or buy the monitoring role

This ladder is Oliv AI's framework for organisational AI adoption, and it is deliberately built on what a company has arranged rather than what it has licensed. Two orgs on identical stacks routinely sit two rungs apart.

🔹 Rung 1, and why it is not nothing

Signature symptom: your best rep's workflow lives in their browser history.

What it is good at is discovery. People find out what AI is useful for, cheaply. What it cannot do is survive turnover. When that rep leaves, the capability leaves with them.

🔹 Rung 2, the most crowded rung in B2B

Signature symptom: three reps have three different research prompts, all decent, none identical.

The gains here are real, and they are personal. Salesforce's 2026 State of Sales, surveying more than 4,000 sellers, found research time down 34% and content creation down 36% for AI users. That is genuine leverage, and it is the same leverage most generative AI use in sales delivers at the individual level. What it cannot do is show up in team numbers, because nothing is shared.

🔹 Rung 3, where standards start to hold

Signature symptom: a new hire can run the standard account-research workflow in week one without shadowing anyone.

What it is good at is consistency without supervision cost. It is also where methodology compliance stops being a slide and starts being a system, which is the whole argument behind automating MEDDIC, BANT, and SPICED scoring from calls. What it cannot do is act unprompted. Every run still waits for a human to start it.

🔹 Rung 4, and what it actually demands

Signature symptom: work completes overnight and someone reviews it in the morning.

What it is good at is volume of repeated judgement. What it cannot do is compensate for undocumented process. An agent inherits your history, not your standards.

⭐ The shape of the gap, not its size

Notice what changes between rungs. One to two is a skill anyone can pick up in a week. Three to four is a capability question.

Two to three is neither. It is a management change, and that is why it stalls.

Q3. How do you tell which rung you are actually on? [toc=3. Self-Assessment]

Score five dimensions from one to five: workflow standardisation, data quality and trust, inspection and accountability, cross-team alignment, and AI readiness. Then apply the rule most vendor assessments bury. Your rung is your lowest score, not your average, because the weakest dimension caps what agents can reliably do. Four practical tests settle it: can a new hire run your best rep's workflow unaided, is there one place approved skills live, does anyone's job include maintaining them, and does any AI task finish without a human starting it.

📊 The five dimensions, scored honestly

Outreach publishes a comparable public scorecard built on these same five dimensions, which is a useful cross-check if you want a second opinion on your own scoring.

Five Scoring Dimensions: What a 2 and a 4 Look Like in Practice
DimensionA score of 2 looks likeA score of 4 looks like
Workflow standardisationEvery rep researches accounts their own wayOne documented way, followed by most of the team
Data quality and trustDuplicate accounts, activity on the wrong opportunityActivity maps reliably to the right account and deal
Inspection and accountabilityNobody reviews AI outputOutput is sampled and corrected on a cadence
Cross-team alignmentSales, CS, and RevOps use different tools and promptsShared context and shared standards across functions
AI readinessProcess lives in people's headsProcess is written down, exceptions included

The second row is the one that quietly decides everything else, which is why CRM data quality automation for RevOps is usually the first real project rather than the last.

✅ The four-question version, if you have two minutes

Answer yes or no. Be strict.

  1. Can a new hire run your best rep's AI workflow without asking that rep?
  2. Is there one place where approved prompts or skills live?
  3. Does anyone's job description include maintaining them?
  4. Does any AI task complete without a human starting it?

Four noes puts you on Rung 1. Yes to the first only is Rung 2. Yes to the first three is Rung 3. Yes to all four is Rung 4.

⚠️ Score the floor, not the ceiling

Here is the part that trips up senior readers, and I include myself in that. Executives self-assess on their best rep. The organisation runs at the level of its median.

So take your lowest dimension score and treat that as your rung. If data quality sits at 2 while everything else sits at 4, you are a 2. Agents acting on mis-mapped records do not fail loudly. They act confidently on the wrong deal, and every correction costs more than the run saved.

❌ The misread I see most often

A company buys at Rung 4 and operates at Rung 2. The contract says autonomous. The behaviour says individual prompting with better branding.

That is not a vendor failure. It is the operating model never having changed to match the purchase. Re-score these five dimensions each quarter, and that drift becomes visible before your renewal does. Teams running a formal revenue intelligence platform comparison should bring this score into the evaluation rather than after it.

Q4. Which rung is right for you, and when is staying put the correct call? [toc=4. Reasons to Stay]

Rung 2 is genuinely correct for a small team of capable operators, because coordination overhead would cost more than the inconsistency it removes. Rung 3 is the right ceiling for most companies: shared standards, no supervision burden, no new role to hire. Rung 4 earns its place only where the volume of repeated judgement exceeds what a team can supervise by hand. Moving up is an organisational cost, not a licence cost, which is why buying more tools never moves you.

⚠️ The assumption I want to take away from you

Every maturity model implies that higher is better. Read four of them and you will feel behind by lunchtime.

The pressure is real, and it is loud. Gartner projects AI agents will outnumber human sellers by ten to one by 2028, with 81% of B2B sales teams already using AI in some form during 2026. That is a forecast about the market, though. It is not an instruction about your org.

💸 What climbing too early actually costs

I have watched this go wrong in a specific, boring way. A team approves a Rung 4 project without a Rung 3 foundation. Six months later there are agents nobody monitors and a playbook nobody wrote.

The cost lands in three places.

  • Staffing. A monitoring role you have not hired, absorbed by RevOps on top of their day job.
  • Trust. Reps stop using output they have caught being wrong twice.
  • Time. The documentation work you skipped still has to happen, now under pressure.

That third cost is the one finance never sees coming, and it is a large part of why agentic AI implementation depends on data architecture rather than enthusiasm.

✅ Honest reasons to stay where you are

So let me give each rung a real defence.

Stay on Rung 2 if you have a handful of genuinely capable operators and a process that changes monthly. Standardising something that keeps moving is waste. ChatGPT and Claude are the correct tools here, and I would keep using both.

Stay on Rung 3 if your work is varied, your judgement calls are one-offs, and a human reviewing output is not your bottleneck. This is the right ceiling for most companies, and I do not say that as a throwaway concession. Most revenue orgs would get more value from a maintained shared library than from any agent deployment they could staff this year.

Move to Rung 4 only when the same judgement is being made hundreds of times a week and supervision, not capability, is your constraint. The CRO guide to agentic AI in revenue intelligence walks through what that supervision load looks like before you commit to it.

Three stacked layers showing when to stay on rung two, rung three, or move to rung four of AI adoption
Rung 3 is the right ceiling for most revenue teams, which is why every layer here carries a reason to stay put.

⭐ The criterion, restated as a decision

Put the two numbers side by side and pick. Coordination cost versus inconsistency cost.

If coordinating is cheaper than the inconsistency you are absorbing, climb. If not, stay, and spend the money on something else. Nobody has ever lost a quarter by refusing to climb a rung.

❌ What being wrong looks like in either direction

Climb too early and you get unsupervised agents plus an unwritten playbook. Stay too long and you get a team where your best rep's method never becomes anyone else's.

Of those two failures, I think the second is more common and the first is more expensive. I might be weighting that from where I sit, so test it against your own numbers rather than mine.

Q5. Why do most revenue teams get stuck after individual AI use? [toc=5. The Stall Point]

Because the move from Rung 2 to Rung 3 is not a skill, it is a management decision. Someone has to own a shared library, decide what becomes standard, retire what does not work, and maintain it as the process changes. That job has usually not been given to anyone. So individual capability compounds while shared capability stays at zero. The organisation gets individually efficient and organisationally unchanged. The unlock costs nothing. Name an owner for the shared skill library.

⚠️ The scene I keep walking into

A VP of Sales pulls up a prompt her top AE built. It writes a genuinely good pre-call brief. She is delighted, and she should be.

Then I ask how many other reps use it. The answer is one. Her team number has not moved in two quarters, and she cannot work out why. It is the same pattern I see when coaching at scale using AI stays trapped inside one manager's habits.

❌ The default response, and what it costs

The usual reaction is to buy something. Another pilot, another seat block, another vendor evaluation. I have sat on the other side of those calls, and I can tell you the purchase rarely fails on capability.

It fails because the tool arrives into the same operating model. Gartner's 2026 data puts 81% of B2B sales teams using AI already, while fewer than 40% of sellers report a measurable productivity gain. That spread is not a software problem. Buying again does not close it, which is why consolidating the revenue tech stack usually beats adding to it.

A senior revenue leader put the distinction to me better than I had managed to. His organisation had been in an AI experimentation phase for a year. They had gained individual efficiency, he said, and no organisational efficiency at all. The mandate had been personal productivity, which is fine, and it is also a ceiling.

🔍 What actually decays, and how fast

Here is the mechanical reason it stalls. A shared library without an owner is not a library, it is a folder.

  • Week one. Someone posts four good prompts. Everyone is pleased.
  • Week three. The qualification criteria change in a pipeline review. Nobody edits the folder.
  • Week six. Two reps have quietly forked their own versions, because the shared one is now wrong.
  • Week ten. People stop opening it, and you are back on Rung 2 with extra steps.
Four-stage timeline showing a shared AI prompt library decaying from adoption to abandonment without an owner
Nothing breaks and no decision gets made. The library simply decays, which is why one named owner is the whole fix.

Nothing broke. No decision was made. Ownership was simply never assigned.

"I found the AI tracker setup to be quite difficult, especially concerning the user interface when setting up keywords or smart trackers. Moreover, I cannot download all the data myself unless we upgrade the plan, which isn't ideal and results in me not fully utilizing Gong."
Verified user, Sales Professional Gong G2 Verified Review [3 Oct 2025]

That review is about configuration friction, and it points at the same thing. When setup is hard enough that only one person does it, the capability never becomes the team's. The same complaint shows up across how Gong smart trackers are configured, where the setup burden sits with one admin.

✅ The one move, and what the owner actually decides

Give one person the shared library. Not a committee, not a workstream, one name.

Their job has four decisions in it. What becomes standard. What gets retired. Who can edit. When it gets reviewed after a process change. Two hours a week covers it at most mid-market scales.

⭐ The tell that you have crossed over

You will know it happened without running a survey. A new hire in week one runs the standard account-research workflow, start to finish, without asking your best rep for anything.

That is Rung 3. It cost a name on a job description, and I still find it hard to convince people the fix is that small. If ramp speed is the thing you are actually measuring, new hire ramp time and onboarding is the cleanest place to see the difference.

Q6. What has to be true before agents can run in production? [toc=6. Production Prerequisites]

Four things. Your process is written down as it actually runs, exceptions included. Your data is trustworthy enough that activity maps to the right account and opportunity, because an agent acting on a mis-mapped record compounds the error with full confidence. Permissions are explicit, covering which actions an agent takes alone and which need approval. And disclosure is wired in, because EU rules now require an agent to identify itself and who it acts for. A company whose process is undocumented cannot reach Rung 3, never mind Rung 4.

Ascending four-step diagram of prerequisites for running AI agents in production in a revenue team
These four prerequisites stack in order, and the first one is management work that no vendor can sell you.

1️⃣ Documented process, as it actually runs

An agent inherits your history, not your standards. If your qualification rules, discount triggers, and legal handoff live in people's heads, the agent will invent a version of them.

Verify it in an afternoon. Ask three managers, separately, when a deal goes to legal. If you get three answers, you have found your first piece of work.

⚠️ Why this one has no software in it

Writing the playbook is a management project. There is no tool that does it for you, and saying otherwise would be the easiest lie in this category.

The reason this matters so much is failure rate. The widely repeated figure in agentic AI is that most projects fail for the same reason, which is that the processes were never documented. Documented is also not the same as alive. Playbooks drift within weeks because exceptions get added verbally and never written back.

2️⃣ Data you can actually act on

This is the hard ceiling, not a footnote. Salesforce's 2026 State of Sales, across more than 4,000 sellers, names data quality and admin friction as the top blockers for teams not seeing agent ROI.

The failure mode is specific. An account has five open opportunities, or three duplicate records, and the meeting gets attached to the wrong one. The agent then writes a confident update on a deal that does not exist. This is exactly why a CRM data strategy tied to revenue predictability has to precede any agent rollout.

"The conversation intelligence tool is lacking, and we don't have the context of the deals against the conversation intelligence findings. The CRM writeback is not good; we cannot send MEDDIC values back to Salesforce or update fields in Salesforce from the conversation intelligence."
Verified user, Revenue Operations Clari G2 Verified Review [13 Jul 2026]

MEDDIC is a qualification checklist covering metrics, buyer, decision process, and pain. That review is describing a broken write path, which is exactly the prerequisite failing. If you want the underlying framework rather than the tooling complaint, the MEDDIC sales methodology explains what should be landing in those fields.

3️⃣ Permissions, written as a list

Split every agent action into two buckets before anything runs.

Agent Permission Buckets: What Needs Approval and What Does Not
BucketRuleExamples
Ask before actingAnything a customer sees, or anything that changes deal economicsFirst outbound email, pricing changes, stage moves
Act without askingReversible internal work with a visible audit trailCRM field updates, meeting summaries, task creation

Forrester's October 2025 predictions expected 20% of sellers to use agent led negotiation during 2026. Once agents touch commercial terms, an unwritten permission list becomes a real exposure, and the governance side of that is covered properly in the RevOps evaluation of AI CRM trust and governance risk.

4️⃣ Disclosure, with dates you should know

Article 50 of the EU AI Act became enforceable on 2 August 2026, and the Commission's July 2026 guidelines confirm agents must disclose their AI nature and the entity they act for. Systems already in place have until 2 December 2026 to implement marking.

The higher risk obligations in Annex III were deferred to 2 December 2027 by Regulation (EU) 2026/1744. So the disclosure work is live now, and the heavier compliance lift is not. Wire the disclosure into the system, not into the copy, and log the event.

⏰ Honest sequencing

If you are missing the first two, Rung 4 is not a roadmap item this year. Start with the playbook and a duplicate record audit. Both are unglamorous, and both are cheaper than a failed agent rollout. Teams that want the audit as a repeatable routine should look at CRM data quality automation for RevOps rather than a one-off cleanup sprint.

Q7. Who does this work, and can you reach Rung 4 without a GTM engineer? [toc=7. Who Owns Rung 4]

No, not without someone doing that job. Rung 4 means agents run while a GTM engineer monitors and optimises them, deciding what runs unsupervised, watching where output drifts, and updating context when the process changes. There are two honest routes. Grow the role from RevOps, which is slower but keeps the knowledge in house. Or buy it as a service. Oliv AI publishes named in house forward deployed GTM engineers who implement directly with a customer's stack, with a first agent live on real deals by week two. The role is a prerequisite either way.

⚠️ The objection, in the reader's own words

"We do not have a GTM engineer and we cannot hire one." I hear that on most calls where this ladder comes up, and it is a fair blocker, not an excuse.

A GTM engineer, to define it plainly, is the person who turns revenue processes into reliable agents and then keeps them reliable. The title is new. The work is not, and much of it already sits inside the job described in building a revenue operations function.

❌ What happens when nobody owns it

Agents get deployed and then quietly go wrong. Not dramatically wrong, which would be easier. Slightly wrong, in ways that take a quarter to surface.

The cost lands on RevOps, on top of their existing job. I have watched a RevOps lead absorb agent monitoring informally for four months, then hand the whole thing back. That was not a tooling failure. It was an unassigned role.

✅ The two routes, and the fork underneath them

Route one is to grow the role from RevOps. Slower, and the institutional knowledge stays with you. Route two is to buy the capability as a service.

That is the same build versus buy fork that decides most Rung 4 projects, and it deserves its own analysis rather than a paragraph here. If you are working through the data and architecture side of it, the RevOps guide to implementing agentic AI covers the ground properly, and the commercial side sits in build versus buy for revenue AI.

⭐ How this works in practice

Oliv AI publishes in house, forward deployed GTM engineers who implement directly with your stack, with named engineer profiles on the site and no third party implementation partner. The published sequence is discovery on day one, connected systems in week one, a first agent live on real deals with real reps in week two, and measurable ROI on that agent by week four. We built it this way because the role is scarce, not because implementation is easy to sell. The caveat belongs here too, and an executive checking the source will find it anyway: Oliv is the least publicly proven vendor named in this piece, with no G2, Capterra, or TrustRadius presence, and case studies that sit behind an email form. On a framework article that matters less commercially than it does to your ability to verify what I am claiming, so I would rather state it than have you discover it. Oliv is not the only route to Rung 4, and any vendor offering this should be asked for one named production customer before a second meeting.

💰 The cost note nobody mentions

There is a practical first step that does not need a staffing decision at all. Oliv AI's published price ladder starts at Amplify $0, then Converse $19, Sell $49, and Grow $79 per user per month, with agent actions billed at $0.01 per credit.

Free seats matter more than they sound here. The executives, product people, and ops folk who read an article like this can sit on a shared platform before anyone approves a seat block, and shared access is the actual mechanic of Rung 3.

⏰ What I would do if the role feels impossible

Do not hire for it first. Pick one repeated judgement, assign one existing person two hours a week to supervise it, and see whether the correction rate falls.

If it does, you have your business case. If it does not, you have saved yourself a hire, and Rung 3 was the right ceiling after all.

Q8. How do you measure whether you are actually moving up? [toc=8. Measuring Movement]

Stop reporting AI adoption percentage. At 81% industry wide it distinguishes nothing. Measure four things instead: share of AI work running through shared skills rather than personal prompts, time from a process change to the shared library reflecting it, workflows a new hire can run in week one unaided, and at Rung 4 the ratio of agent actions completed to agent actions corrected. Correction rate is your observability measure. An agent you correct constantly is a Rung 3 workflow wearing a Rung 4 label.

📊 The four numbers worth reporting

Four Metrics That Track Movement Up the AI Adoption Ladder
MetricHow to collect itWhat a bad number means
Shared skill shareCount runs from the shared library against total AI runs, monthlyIndividual capability is compounding, shared capability is not
Library lagDays from a process change to the library reflecting itNobody owns maintenance, so you are drifting back to Rung 2
Week one workflowsAsk each new hire what they ran unaided in week oneYour standards live in people, not in the system
Agent correction rateCorrected actions divided by completed actionsSupervision cost is eating the automation gain

All four are countable this quarter, without new tooling. Each moves only when the operating model moves, which is the point. If your board wants these sitting next to the commercial numbers, revenue performance analytics is where they belong.

⭐ Why correction rate is the honest one

Every other agent metric can be gamed by narrowing scope. Run an agent on one trivial task and your success rate looks superb.

Correction rate resists that. If your team fixes four out of ten outputs, you have a supervised workflow, not an autonomous one. I would rather a client told me their correction rate was 40% than showed me a dashboard with a green tick.

⚠️ How to read the benchmark someone will hand you

Pavilion's 2026 GTM benchmark reports that AI powered teams generate roughly 40% more pipeline per rep. You will see that number in a deck within a month, probably in a vendor deck.

Treat it as a hypothesis, not a target. Benchmarks built on self selected respondents carry survivorship bias, which means the teams that struggled often did not report. Publish your own before and after per rep number instead. It will be less impressive and far more useful, and a revenue intelligence ROI calculator is a better starting point than someone else's average.

❌ Three numbers never to report

Some figures look rigorous and prove nothing. I would keep all three out of your board pack.

  1. AI adoption percentage. With 81% of B2B sales teams already using AI in 2026, this measures nothing about your operating model.
  2. Industry stage distributions. Any claim that a specific percentage of companies sit on a given rung is unfalsifiable. I do not have that number, and I have not seen anyone who does.
  3. Typical time between levels. The 2 to 3 move can take a week if someone is assigned, or never if nobody is.

✅ Make it a quarterly habit

Re score the five dimensions from the self assessment every quarter, on the same day you review pipeline coverage. It takes ten minutes.

If you want a shared vocabulary with your board, G2's public taxonomy runs from ad hoc and aware through developing, mature, leading, and transformative. Map your rungs to it once, then stop arguing about labels and go and fix your lowest score. Where that score is forecast discipline rather than tooling, running evidence based forecast commits is the fastest thing to fix next.

Where is your revenue org on AI right now? [toc=0. Introduction]

Where is your revenue org on AI right now? Not which tools you own. What stage you operate at.

Most executives I ask cannot answer that, and they are not being evasive. They can list every licence. They have seen good demos from their own reps. What they do not have is a way to think about sequence, which is the thing a board actually asks for.

So let me concede something before I hand you a framework. Most AI maturity models are sales tools. Four rungs, and the author's product sits on the top one, and any operator who has read two of them knows it.

I am going to do the opposite here, because the framework is only useful if it can tell you to stay put. Every rung in this article gets an honest reason to remain on it. Rung 3 is the right ceiling for most companies, and I would rather say that plainly than sell you a climb you cannot staff. If you want the longer version of that argument, the honest read on what AI agents can actually do for your team today covers the same ground for a VP of Sales.

The test I will use does not involve any vendor, including mine. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. That is the whole diagnostic.

One more thing worth naming early. The argument underneath all eight sections is that most revenue orgs are not behind on tooling at all. They are stuck between individual capability and shared capability, and that gap is organisational, which is exactly why buying more software never closes it. It is the same structural shift described in how revenue teams evolved from RevOps to intelligence to orchestration.

You will leave with two things. A rung you can defend, and one next move. For most readers that move costs nothing.

What is the one move you should make next? [toc=9. Your Next Move]

If you read only one line of this, make it this one: your rung is set by what your organisation has arranged, not by what it has licensed. Score the five dimensions honestly, take your lowest number, and accept that as your answer for this quarter. Then pick the single next move that sits directly above it, which for most teams reading this means giving one named person the shared skill library and two hours a week to maintain it. That is a management decision, it costs nothing, and it is the actual unlock between individual and shared capability. If your next move turns out to be the Rung 3 to Rung 4 jump, and you want to pressure test whether the supervision maths works at your volume, book a seven minute chat and bring your correction rate rather than your tool list.

Two adjacent reads, depending on which way your lowest score points. If the gap is ownership and structure, start with scaling revenue operations in a growth stage sales team. If the gap is whether agents belong in your motion at all, agentic AI for revenue execution works through the pipeline side of it, and the future of revenue intelligence sets the longer horizon.

✍🏼 About the author

Ishan Chhabra is the founder and CEO of Oliv AI, an AI native revenue intelligence and revenue orchestration platform for B2B revenue teams. He built Oliv's context graph, the infrastructure layer that resolves accounts, opportunities, and conversations across messy CRMs so agents can act on them safely.

He writes about what he sees working and failing inside revenue organisations adopting AI, including the parts that fail for reasons no vendor likes to publish. His work on AI agents for RevOps and the CRO view of ROI on a revenue platform covers the same territory in more depth.

Q1. Why should you trust a maturity model at all? [toc=1. Trusting the Model]

Most AI maturity models are sales artefacts. Four rungs, and the author's product sits on the top one. The test that survives that bias is a cost comparison, not a capability list. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. Below that line, coordination is pure overhead. Above it, inconsistency is already costing you deals.

⚠️ The board question nobody can answer

A CRO I spoke with last quarter had approved AI spend across three teams. Her reps were enthusiastic. Her SDRs had built prompts they swore by. Then her board asked a simple question: where are we, and what comes next?

She could name every tool. She could not name a stage. That gap is the reason this genre of article exists, and also the reason most of it is useless. It is the same gap I see when a team can list its stack but cannot describe how AI actually runs inside revenue operations.

❌ Why the genre deserves your suspicion

Here is the problem, said plainly. Maturity models are usually written by vendors, and the top rung is usually a description of the vendor's product. An experienced operator spots that in about four seconds, and rightly stops reading.

Believing one of those models has a real cost. You end up approving a transformation programme you cannot staff. You buy for a rung you do not operate at. Gartner's 2026 Hype Cycle placed agentic AI at the Peak of Inflated Expectations, with only 17% of organisations having actually deployed agents while more than 60% plan to within two years. That gap between intent and deployment is where budgets go to die, and it is why the build versus buy decision on revenue AI deserves more scrutiny than a maturity chart.

✅ The test that does not involve any vendor

So use a test that no software can tilt. Compare two costs.

  • Coordination cost. What it takes to agree a standard, write it down, and keep it current.
  • Inconsistency cost. What you lose when every rep runs a different version of the same task.

You belong on the rung where the first number is smaller than the second. That is the whole diagnostic. It works if you never buy anything.

Diagram comparing coordination cost against inconsistency cost to decide the correct GTM AI maturity rung
The diagnostic that survives vendor bias: you belong on the rung where coordination costs less than inconsistency.

⭐ What I will do differently here

Every rung in this article gets an honest reason to stay, not just a reason to climb. I will say the thing vendors do not: Rung 3 is the right ceiling for most companies. Shared standards, no supervision burden, no new role to hire.

I am also going to be specific about where each move gets expensive, because the expense is rarely the licence. AI spend gets approved per tool. Value accrues per operating model. That mismatch is how a company ends up buying at Rung 4 while still operating at Rung 2, with nothing in the numbers to show it.

⏰ What you should leave with

Two things, in about ten minutes of reading. A rung you can defend to a board, and one next move you can start on Monday.

Not a programme. One move. For most readers it turns out to cost nothing at all, which is the part I did not expect when I started mapping this.

Q2. What are the four rungs of the GTM AI adoption ladder, and what blocks each move up? [toc=2. The Four Rungs]

Four rungs. One, everyone prompts individually, context copy-pasted each time, nothing reusable. Two, individuals build reusable skills and run them themselves, which is the most common rung. Three, skills are centralised, teams use the same ones daily, and someone owns them. Four, agents run in production on documented process, with GTM engineers monitoring and optimising. The jump from one to two is a skill. The jump from two to three is a management change. The jump from three to four is documented process plus supervision capacity. Rungs cannot be skipped.

Two definitions before the table, because these words get used loosely. A skill is a saved, reusable instruction set a person runs on demand. An agent is a process that runs without a person starting it. A GTM engineer is the person who builds, monitors, and maintains those agents. If the distinction between the last two is still fuzzy, the difference between AI agents built for sales teams and a saved prompt is exactly what this ladder is measuring.

📊 The ladder, with the constraint that blocks each move

The GTM AI Adoption Ladder: Four Rungs, Constraints, and Next Actions
RungWhat is observably trueWho does the workBlocking constraintSingle next action
01. ChatGPT for everyonePeople prompt ad hoc. Context is pasted in fresh each time. Nothing is saved.Individual reps, unevenlyNobody knows which prompts workAsk three reps to share their best prompt in one doc
02. Claude and DIY skillsIndividuals build reusable skills and run them personally. Quality varies by person.Capable individualsNo owner for a shared libraryName an owner for the shared skill library
03. Centralised skillsOne library, agreed standards, teams use the same skills daily, someone maintains them.The team, on a standardProcess is undocumented, supervision capacity is zeroWrite the playbook as it actually runs
04. Autonomous agents in productionAgents run on documented process. Humans monitor, correct, and optimise.Agents, supervised by a GTM engineerNobody owns agent monitoringAssign or buy the monitoring role

This ladder is Oliv AI's framework for organisational AI adoption, and it is deliberately built on what a company has arranged rather than what it has licensed. Two orgs on identical stacks routinely sit two rungs apart.

🔹 Rung 1, and why it is not nothing

Signature symptom: your best rep's workflow lives in their browser history.

What it is good at is discovery. People find out what AI is useful for, cheaply. What it cannot do is survive turnover. When that rep leaves, the capability leaves with them.

🔹 Rung 2, the most crowded rung in B2B

Signature symptom: three reps have three different research prompts, all decent, none identical.

The gains here are real, and they are personal. Salesforce's 2026 State of Sales, surveying more than 4,000 sellers, found research time down 34% and content creation down 36% for AI users. That is genuine leverage, and it is the same leverage most generative AI use in sales delivers at the individual level. What it cannot do is show up in team numbers, because nothing is shared.

🔹 Rung 3, where standards start to hold

Signature symptom: a new hire can run the standard account-research workflow in week one without shadowing anyone.

What it is good at is consistency without supervision cost. It is also where methodology compliance stops being a slide and starts being a system, which is the whole argument behind automating MEDDIC, BANT, and SPICED scoring from calls. What it cannot do is act unprompted. Every run still waits for a human to start it.

🔹 Rung 4, and what it actually demands

Signature symptom: work completes overnight and someone reviews it in the morning.

What it is good at is volume of repeated judgement. What it cannot do is compensate for undocumented process. An agent inherits your history, not your standards.

⭐ The shape of the gap, not its size

Notice what changes between rungs. One to two is a skill anyone can pick up in a week. Three to four is a capability question.

Two to three is neither. It is a management change, and that is why it stalls.

Q3. How do you tell which rung you are actually on? [toc=3. Self-Assessment]

Score five dimensions from one to five: workflow standardisation, data quality and trust, inspection and accountability, cross-team alignment, and AI readiness. Then apply the rule most vendor assessments bury. Your rung is your lowest score, not your average, because the weakest dimension caps what agents can reliably do. Four practical tests settle it: can a new hire run your best rep's workflow unaided, is there one place approved skills live, does anyone's job include maintaining them, and does any AI task finish without a human starting it.

📊 The five dimensions, scored honestly

Outreach publishes a comparable public scorecard built on these same five dimensions, which is a useful cross-check if you want a second opinion on your own scoring.

Five Scoring Dimensions: What a 2 and a 4 Look Like in Practice
DimensionA score of 2 looks likeA score of 4 looks like
Workflow standardisationEvery rep researches accounts their own wayOne documented way, followed by most of the team
Data quality and trustDuplicate accounts, activity on the wrong opportunityActivity maps reliably to the right account and deal
Inspection and accountabilityNobody reviews AI outputOutput is sampled and corrected on a cadence
Cross-team alignmentSales, CS, and RevOps use different tools and promptsShared context and shared standards across functions
AI readinessProcess lives in people's headsProcess is written down, exceptions included

The second row is the one that quietly decides everything else, which is why CRM data quality automation for RevOps is usually the first real project rather than the last.

✅ The four-question version, if you have two minutes

Answer yes or no. Be strict.

  1. Can a new hire run your best rep's AI workflow without asking that rep?
  2. Is there one place where approved prompts or skills live?
  3. Does anyone's job description include maintaining them?
  4. Does any AI task complete without a human starting it?

Four noes puts you on Rung 1. Yes to the first only is Rung 2. Yes to the first three is Rung 3. Yes to all four is Rung 4.

⚠️ Score the floor, not the ceiling

Here is the part that trips up senior readers, and I include myself in that. Executives self-assess on their best rep. The organisation runs at the level of its median.

So take your lowest dimension score and treat that as your rung. If data quality sits at 2 while everything else sits at 4, you are a 2. Agents acting on mis-mapped records do not fail loudly. They act confidently on the wrong deal, and every correction costs more than the run saved.

❌ The misread I see most often

A company buys at Rung 4 and operates at Rung 2. The contract says autonomous. The behaviour says individual prompting with better branding.

That is not a vendor failure. It is the operating model never having changed to match the purchase. Re-score these five dimensions each quarter, and that drift becomes visible before your renewal does. Teams running a formal revenue intelligence platform comparison should bring this score into the evaluation rather than after it.

Q4. Which rung is right for you, and when is staying put the correct call? [toc=4. Reasons to Stay]

Rung 2 is genuinely correct for a small team of capable operators, because coordination overhead would cost more than the inconsistency it removes. Rung 3 is the right ceiling for most companies: shared standards, no supervision burden, no new role to hire. Rung 4 earns its place only where the volume of repeated judgement exceeds what a team can supervise by hand. Moving up is an organisational cost, not a licence cost, which is why buying more tools never moves you.

⚠️ The assumption I want to take away from you

Every maturity model implies that higher is better. Read four of them and you will feel behind by lunchtime.

The pressure is real, and it is loud. Gartner projects AI agents will outnumber human sellers by ten to one by 2028, with 81% of B2B sales teams already using AI in some form during 2026. That is a forecast about the market, though. It is not an instruction about your org.

💸 What climbing too early actually costs

I have watched this go wrong in a specific, boring way. A team approves a Rung 4 project without a Rung 3 foundation. Six months later there are agents nobody monitors and a playbook nobody wrote.

The cost lands in three places.

  • Staffing. A monitoring role you have not hired, absorbed by RevOps on top of their day job.
  • Trust. Reps stop using output they have caught being wrong twice.
  • Time. The documentation work you skipped still has to happen, now under pressure.

That third cost is the one finance never sees coming, and it is a large part of why agentic AI implementation depends on data architecture rather than enthusiasm.

✅ Honest reasons to stay where you are

So let me give each rung a real defence.

Stay on Rung 2 if you have a handful of genuinely capable operators and a process that changes monthly. Standardising something that keeps moving is waste. ChatGPT and Claude are the correct tools here, and I would keep using both.

Stay on Rung 3 if your work is varied, your judgement calls are one-offs, and a human reviewing output is not your bottleneck. This is the right ceiling for most companies, and I do not say that as a throwaway concession. Most revenue orgs would get more value from a maintained shared library than from any agent deployment they could staff this year.

Move to Rung 4 only when the same judgement is being made hundreds of times a week and supervision, not capability, is your constraint. The CRO guide to agentic AI in revenue intelligence walks through what that supervision load looks like before you commit to it.

Three stacked layers showing when to stay on rung two, rung three, or move to rung four of AI adoption
Rung 3 is the right ceiling for most revenue teams, which is why every layer here carries a reason to stay put.

⭐ The criterion, restated as a decision

Put the two numbers side by side and pick. Coordination cost versus inconsistency cost.

If coordinating is cheaper than the inconsistency you are absorbing, climb. If not, stay, and spend the money on something else. Nobody has ever lost a quarter by refusing to climb a rung.

❌ What being wrong looks like in either direction

Climb too early and you get unsupervised agents plus an unwritten playbook. Stay too long and you get a team where your best rep's method never becomes anyone else's.

Of those two failures, I think the second is more common and the first is more expensive. I might be weighting that from where I sit, so test it against your own numbers rather than mine.

Q5. Why do most revenue teams get stuck after individual AI use? [toc=5. The Stall Point]

Because the move from Rung 2 to Rung 3 is not a skill, it is a management decision. Someone has to own a shared library, decide what becomes standard, retire what does not work, and maintain it as the process changes. That job has usually not been given to anyone. So individual capability compounds while shared capability stays at zero. The organisation gets individually efficient and organisationally unchanged. The unlock costs nothing. Name an owner for the shared skill library.

⚠️ The scene I keep walking into

A VP of Sales pulls up a prompt her top AE built. It writes a genuinely good pre-call brief. She is delighted, and she should be.

Then I ask how many other reps use it. The answer is one. Her team number has not moved in two quarters, and she cannot work out why. It is the same pattern I see when coaching at scale using AI stays trapped inside one manager's habits.

❌ The default response, and what it costs

The usual reaction is to buy something. Another pilot, another seat block, another vendor evaluation. I have sat on the other side of those calls, and I can tell you the purchase rarely fails on capability.

It fails because the tool arrives into the same operating model. Gartner's 2026 data puts 81% of B2B sales teams using AI already, while fewer than 40% of sellers report a measurable productivity gain. That spread is not a software problem. Buying again does not close it, which is why consolidating the revenue tech stack usually beats adding to it.

A senior revenue leader put the distinction to me better than I had managed to. His organisation had been in an AI experimentation phase for a year. They had gained individual efficiency, he said, and no organisational efficiency at all. The mandate had been personal productivity, which is fine, and it is also a ceiling.

🔍 What actually decays, and how fast

Here is the mechanical reason it stalls. A shared library without an owner is not a library, it is a folder.

  • Week one. Someone posts four good prompts. Everyone is pleased.
  • Week three. The qualification criteria change in a pipeline review. Nobody edits the folder.
  • Week six. Two reps have quietly forked their own versions, because the shared one is now wrong.
  • Week ten. People stop opening it, and you are back on Rung 2 with extra steps.
Four-stage timeline showing a shared AI prompt library decaying from adoption to abandonment without an owner
Nothing breaks and no decision gets made. The library simply decays, which is why one named owner is the whole fix.

Nothing broke. No decision was made. Ownership was simply never assigned.

"I found the AI tracker setup to be quite difficult, especially concerning the user interface when setting up keywords or smart trackers. Moreover, I cannot download all the data myself unless we upgrade the plan, which isn't ideal and results in me not fully utilizing Gong."
Verified user, Sales Professional Gong G2 Verified Review [3 Oct 2025]

That review is about configuration friction, and it points at the same thing. When setup is hard enough that only one person does it, the capability never becomes the team's. The same complaint shows up across how Gong smart trackers are configured, where the setup burden sits with one admin.

✅ The one move, and what the owner actually decides

Give one person the shared library. Not a committee, not a workstream, one name.

Their job has four decisions in it. What becomes standard. What gets retired. Who can edit. When it gets reviewed after a process change. Two hours a week covers it at most mid-market scales.

⭐ The tell that you have crossed over

You will know it happened without running a survey. A new hire in week one runs the standard account-research workflow, start to finish, without asking your best rep for anything.

That is Rung 3. It cost a name on a job description, and I still find it hard to convince people the fix is that small. If ramp speed is the thing you are actually measuring, new hire ramp time and onboarding is the cleanest place to see the difference.

Q6. What has to be true before agents can run in production? [toc=6. Production Prerequisites]

Four things. Your process is written down as it actually runs, exceptions included. Your data is trustworthy enough that activity maps to the right account and opportunity, because an agent acting on a mis-mapped record compounds the error with full confidence. Permissions are explicit, covering which actions an agent takes alone and which need approval. And disclosure is wired in, because EU rules now require an agent to identify itself and who it acts for. A company whose process is undocumented cannot reach Rung 3, never mind Rung 4.

Ascending four-step diagram of prerequisites for running AI agents in production in a revenue team
These four prerequisites stack in order, and the first one is management work that no vendor can sell you.

1️⃣ Documented process, as it actually runs

An agent inherits your history, not your standards. If your qualification rules, discount triggers, and legal handoff live in people's heads, the agent will invent a version of them.

Verify it in an afternoon. Ask three managers, separately, when a deal goes to legal. If you get three answers, you have found your first piece of work.

⚠️ Why this one has no software in it

Writing the playbook is a management project. There is no tool that does it for you, and saying otherwise would be the easiest lie in this category.

The reason this matters so much is failure rate. The widely repeated figure in agentic AI is that most projects fail for the same reason, which is that the processes were never documented. Documented is also not the same as alive. Playbooks drift within weeks because exceptions get added verbally and never written back.

2️⃣ Data you can actually act on

This is the hard ceiling, not a footnote. Salesforce's 2026 State of Sales, across more than 4,000 sellers, names data quality and admin friction as the top blockers for teams not seeing agent ROI.

The failure mode is specific. An account has five open opportunities, or three duplicate records, and the meeting gets attached to the wrong one. The agent then writes a confident update on a deal that does not exist. This is exactly why a CRM data strategy tied to revenue predictability has to precede any agent rollout.

"The conversation intelligence tool is lacking, and we don't have the context of the deals against the conversation intelligence findings. The CRM writeback is not good; we cannot send MEDDIC values back to Salesforce or update fields in Salesforce from the conversation intelligence."
Verified user, Revenue Operations Clari G2 Verified Review [13 Jul 2026]

MEDDIC is a qualification checklist covering metrics, buyer, decision process, and pain. That review is describing a broken write path, which is exactly the prerequisite failing. If you want the underlying framework rather than the tooling complaint, the MEDDIC sales methodology explains what should be landing in those fields.

3️⃣ Permissions, written as a list

Split every agent action into two buckets before anything runs.

Agent Permission Buckets: What Needs Approval and What Does Not
BucketRuleExamples
Ask before actingAnything a customer sees, or anything that changes deal economicsFirst outbound email, pricing changes, stage moves
Act without askingReversible internal work with a visible audit trailCRM field updates, meeting summaries, task creation

Forrester's October 2025 predictions expected 20% of sellers to use agent led negotiation during 2026. Once agents touch commercial terms, an unwritten permission list becomes a real exposure, and the governance side of that is covered properly in the RevOps evaluation of AI CRM trust and governance risk.

4️⃣ Disclosure, with dates you should know

Article 50 of the EU AI Act became enforceable on 2 August 2026, and the Commission's July 2026 guidelines confirm agents must disclose their AI nature and the entity they act for. Systems already in place have until 2 December 2026 to implement marking.

The higher risk obligations in Annex III were deferred to 2 December 2027 by Regulation (EU) 2026/1744. So the disclosure work is live now, and the heavier compliance lift is not. Wire the disclosure into the system, not into the copy, and log the event.

⏰ Honest sequencing

If you are missing the first two, Rung 4 is not a roadmap item this year. Start with the playbook and a duplicate record audit. Both are unglamorous, and both are cheaper than a failed agent rollout. Teams that want the audit as a repeatable routine should look at CRM data quality automation for RevOps rather than a one-off cleanup sprint.

Q7. Who does this work, and can you reach Rung 4 without a GTM engineer? [toc=7. Who Owns Rung 4]

No, not without someone doing that job. Rung 4 means agents run while a GTM engineer monitors and optimises them, deciding what runs unsupervised, watching where output drifts, and updating context when the process changes. There are two honest routes. Grow the role from RevOps, which is slower but keeps the knowledge in house. Or buy it as a service. Oliv AI publishes named in house forward deployed GTM engineers who implement directly with a customer's stack, with a first agent live on real deals by week two. The role is a prerequisite either way.

⚠️ The objection, in the reader's own words

"We do not have a GTM engineer and we cannot hire one." I hear that on most calls where this ladder comes up, and it is a fair blocker, not an excuse.

A GTM engineer, to define it plainly, is the person who turns revenue processes into reliable agents and then keeps them reliable. The title is new. The work is not, and much of it already sits inside the job described in building a revenue operations function.

❌ What happens when nobody owns it

Agents get deployed and then quietly go wrong. Not dramatically wrong, which would be easier. Slightly wrong, in ways that take a quarter to surface.

The cost lands on RevOps, on top of their existing job. I have watched a RevOps lead absorb agent monitoring informally for four months, then hand the whole thing back. That was not a tooling failure. It was an unassigned role.

✅ The two routes, and the fork underneath them

Route one is to grow the role from RevOps. Slower, and the institutional knowledge stays with you. Route two is to buy the capability as a service.

That is the same build versus buy fork that decides most Rung 4 projects, and it deserves its own analysis rather than a paragraph here. If you are working through the data and architecture side of it, the RevOps guide to implementing agentic AI covers the ground properly, and the commercial side sits in build versus buy for revenue AI.

⭐ How this works in practice

Oliv AI publishes in house, forward deployed GTM engineers who implement directly with your stack, with named engineer profiles on the site and no third party implementation partner. The published sequence is discovery on day one, connected systems in week one, a first agent live on real deals with real reps in week two, and measurable ROI on that agent by week four. We built it this way because the role is scarce, not because implementation is easy to sell. The caveat belongs here too, and an executive checking the source will find it anyway: Oliv is the least publicly proven vendor named in this piece, with no G2, Capterra, or TrustRadius presence, and case studies that sit behind an email form. On a framework article that matters less commercially than it does to your ability to verify what I am claiming, so I would rather state it than have you discover it. Oliv is not the only route to Rung 4, and any vendor offering this should be asked for one named production customer before a second meeting.

💰 The cost note nobody mentions

There is a practical first step that does not need a staffing decision at all. Oliv AI's published price ladder starts at Amplify $0, then Converse $19, Sell $49, and Grow $79 per user per month, with agent actions billed at $0.01 per credit.

Free seats matter more than they sound here. The executives, product people, and ops folk who read an article like this can sit on a shared platform before anyone approves a seat block, and shared access is the actual mechanic of Rung 3.

⏰ What I would do if the role feels impossible

Do not hire for it first. Pick one repeated judgement, assign one existing person two hours a week to supervise it, and see whether the correction rate falls.

If it does, you have your business case. If it does not, you have saved yourself a hire, and Rung 3 was the right ceiling after all.

Q8. How do you measure whether you are actually moving up? [toc=8. Measuring Movement]

Stop reporting AI adoption percentage. At 81% industry wide it distinguishes nothing. Measure four things instead: share of AI work running through shared skills rather than personal prompts, time from a process change to the shared library reflecting it, workflows a new hire can run in week one unaided, and at Rung 4 the ratio of agent actions completed to agent actions corrected. Correction rate is your observability measure. An agent you correct constantly is a Rung 3 workflow wearing a Rung 4 label.

📊 The four numbers worth reporting

Four Metrics That Track Movement Up the AI Adoption Ladder
MetricHow to collect itWhat a bad number means
Shared skill shareCount runs from the shared library against total AI runs, monthlyIndividual capability is compounding, shared capability is not
Library lagDays from a process change to the library reflecting itNobody owns maintenance, so you are drifting back to Rung 2
Week one workflowsAsk each new hire what they ran unaided in week oneYour standards live in people, not in the system
Agent correction rateCorrected actions divided by completed actionsSupervision cost is eating the automation gain

All four are countable this quarter, without new tooling. Each moves only when the operating model moves, which is the point. If your board wants these sitting next to the commercial numbers, revenue performance analytics is where they belong.

⭐ Why correction rate is the honest one

Every other agent metric can be gamed by narrowing scope. Run an agent on one trivial task and your success rate looks superb.

Correction rate resists that. If your team fixes four out of ten outputs, you have a supervised workflow, not an autonomous one. I would rather a client told me their correction rate was 40% than showed me a dashboard with a green tick.

⚠️ How to read the benchmark someone will hand you

Pavilion's 2026 GTM benchmark reports that AI powered teams generate roughly 40% more pipeline per rep. You will see that number in a deck within a month, probably in a vendor deck.

Treat it as a hypothesis, not a target. Benchmarks built on self selected respondents carry survivorship bias, which means the teams that struggled often did not report. Publish your own before and after per rep number instead. It will be less impressive and far more useful, and a revenue intelligence ROI calculator is a better starting point than someone else's average.

❌ Three numbers never to report

Some figures look rigorous and prove nothing. I would keep all three out of your board pack.

  1. AI adoption percentage. With 81% of B2B sales teams already using AI in 2026, this measures nothing about your operating model.
  2. Industry stage distributions. Any claim that a specific percentage of companies sit on a given rung is unfalsifiable. I do not have that number, and I have not seen anyone who does.
  3. Typical time between levels. The 2 to 3 move can take a week if someone is assigned, or never if nobody is.

✅ Make it a quarterly habit

Re score the five dimensions from the self assessment every quarter, on the same day you review pipeline coverage. It takes ten minutes.

If you want a shared vocabulary with your board, G2's public taxonomy runs from ad hoc and aware through developing, mature, leading, and transformative. Map your rungs to it once, then stop arguing about labels and go and fix your lowest score. Where that score is forecast discipline rather than tooling, running evidence based forecast commits is the fastest thing to fix next.

Where is your revenue org on AI right now? [toc=0. Introduction]

Where is your revenue org on AI right now? Not which tools you own. What stage you operate at.

Most executives I ask cannot answer that, and they are not being evasive. They can list every licence. They have seen good demos from their own reps. What they do not have is a way to think about sequence, which is the thing a board actually asks for.

So let me concede something before I hand you a framework. Most AI maturity models are sales tools. Four rungs, and the author's product sits on the top one, and any operator who has read two of them knows it.

I am going to do the opposite here, because the framework is only useful if it can tell you to stay put. Every rung in this article gets an honest reason to remain on it. Rung 3 is the right ceiling for most companies, and I would rather say that plainly than sell you a climb you cannot staff. If you want the longer version of that argument, the honest read on what AI agents can actually do for your team today covers the same ground for a VP of Sales.

The test I will use does not involve any vendor, including mine. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. That is the whole diagnostic.

One more thing worth naming early. The argument underneath all eight sections is that most revenue orgs are not behind on tooling at all. They are stuck between individual capability and shared capability, and that gap is organisational, which is exactly why buying more software never closes it. It is the same structural shift described in how revenue teams evolved from RevOps to intelligence to orchestration.

You will leave with two things. A rung you can defend, and one next move. For most readers that move costs nothing.

What is the one move you should make next? [toc=9. Your Next Move]

If you read only one line of this, make it this one: your rung is set by what your organisation has arranged, not by what it has licensed. Score the five dimensions honestly, take your lowest number, and accept that as your answer for this quarter. Then pick the single next move that sits directly above it, which for most teams reading this means giving one named person the shared skill library and two hours a week to maintain it. That is a management decision, it costs nothing, and it is the actual unlock between individual and shared capability. If your next move turns out to be the Rung 3 to Rung 4 jump, and you want to pressure test whether the supervision maths works at your volume, book a seven minute chat and bring your correction rate rather than your tool list.

Two adjacent reads, depending on which way your lowest score points. If the gap is ownership and structure, start with scaling revenue operations in a growth stage sales team. If the gap is whether agents belong in your motion at all, agentic AI for revenue execution works through the pipeline side of it, and the future of revenue intelligence sets the longer horizon.

✍🏼 About the author

Ishan Chhabra is the founder and CEO of Oliv AI, an AI native revenue intelligence and revenue orchestration platform for B2B revenue teams. He built Oliv's context graph, the infrastructure layer that resolves accounts, opportunities, and conversations across messy CRMs so agents can act on them safely.

He writes about what he sees working and failing inside revenue organisations adopting AI, including the parts that fail for reasons no vendor likes to publish. His work on AI agents for RevOps and the CRO view of ROI on a revenue platform covers the same territory in more depth.

Q1. Why should you trust a maturity model at all? [toc=1. Trusting the Model]

Most AI maturity models are sales artefacts. Four rungs, and the author's product sits on the top one. The test that survives that bias is a cost comparison, not a capability list. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. Below that line, coordination is pure overhead. Above it, inconsistency is already costing you deals.

⚠️ The board question nobody can answer

A CRO I spoke with last quarter had approved AI spend across three teams. Her reps were enthusiastic. Her SDRs had built prompts they swore by. Then her board asked a simple question: where are we, and what comes next?

She could name every tool. She could not name a stage. That gap is the reason this genre of article exists, and also the reason most of it is useless. It is the same gap I see when a team can list its stack but cannot describe how AI actually runs inside revenue operations.

❌ Why the genre deserves your suspicion

Here is the problem, said plainly. Maturity models are usually written by vendors, and the top rung is usually a description of the vendor's product. An experienced operator spots that in about four seconds, and rightly stops reading.

Believing one of those models has a real cost. You end up approving a transformation programme you cannot staff. You buy for a rung you do not operate at. Gartner's 2026 Hype Cycle placed agentic AI at the Peak of Inflated Expectations, with only 17% of organisations having actually deployed agents while more than 60% plan to within two years. That gap between intent and deployment is where budgets go to die, and it is why the build versus buy decision on revenue AI deserves more scrutiny than a maturity chart.

✅ The test that does not involve any vendor

So use a test that no software can tilt. Compare two costs.

  • Coordination cost. What it takes to agree a standard, write it down, and keep it current.
  • Inconsistency cost. What you lose when every rep runs a different version of the same task.

You belong on the rung where the first number is smaller than the second. That is the whole diagnostic. It works if you never buy anything.

Diagram comparing coordination cost against inconsistency cost to decide the correct GTM AI maturity rung
The diagnostic that survives vendor bias: you belong on the rung where coordination costs less than inconsistency.

⭐ What I will do differently here

Every rung in this article gets an honest reason to stay, not just a reason to climb. I will say the thing vendors do not: Rung 3 is the right ceiling for most companies. Shared standards, no supervision burden, no new role to hire.

I am also going to be specific about where each move gets expensive, because the expense is rarely the licence. AI spend gets approved per tool. Value accrues per operating model. That mismatch is how a company ends up buying at Rung 4 while still operating at Rung 2, with nothing in the numbers to show it.

⏰ What you should leave with

Two things, in about ten minutes of reading. A rung you can defend to a board, and one next move you can start on Monday.

Not a programme. One move. For most readers it turns out to cost nothing at all, which is the part I did not expect when I started mapping this.

Q2. What are the four rungs of the GTM AI adoption ladder, and what blocks each move up? [toc=2. The Four Rungs]

Four rungs. One, everyone prompts individually, context copy-pasted each time, nothing reusable. Two, individuals build reusable skills and run them themselves, which is the most common rung. Three, skills are centralised, teams use the same ones daily, and someone owns them. Four, agents run in production on documented process, with GTM engineers monitoring and optimising. The jump from one to two is a skill. The jump from two to three is a management change. The jump from three to four is documented process plus supervision capacity. Rungs cannot be skipped.

Two definitions before the table, because these words get used loosely. A skill is a saved, reusable instruction set a person runs on demand. An agent is a process that runs without a person starting it. A GTM engineer is the person who builds, monitors, and maintains those agents. If the distinction between the last two is still fuzzy, the difference between AI agents built for sales teams and a saved prompt is exactly what this ladder is measuring.

📊 The ladder, with the constraint that blocks each move

The GTM AI Adoption Ladder: Four Rungs, Constraints, and Next Actions
RungWhat is observably trueWho does the workBlocking constraintSingle next action
01. ChatGPT for everyonePeople prompt ad hoc. Context is pasted in fresh each time. Nothing is saved.Individual reps, unevenlyNobody knows which prompts workAsk three reps to share their best prompt in one doc
02. Claude and DIY skillsIndividuals build reusable skills and run them personally. Quality varies by person.Capable individualsNo owner for a shared libraryName an owner for the shared skill library
03. Centralised skillsOne library, agreed standards, teams use the same skills daily, someone maintains them.The team, on a standardProcess is undocumented, supervision capacity is zeroWrite the playbook as it actually runs
04. Autonomous agents in productionAgents run on documented process. Humans monitor, correct, and optimise.Agents, supervised by a GTM engineerNobody owns agent monitoringAssign or buy the monitoring role

This ladder is Oliv AI's framework for organisational AI adoption, and it is deliberately built on what a company has arranged rather than what it has licensed. Two orgs on identical stacks routinely sit two rungs apart.

🔹 Rung 1, and why it is not nothing

Signature symptom: your best rep's workflow lives in their browser history.

What it is good at is discovery. People find out what AI is useful for, cheaply. What it cannot do is survive turnover. When that rep leaves, the capability leaves with them.

🔹 Rung 2, the most crowded rung in B2B

Signature symptom: three reps have three different research prompts, all decent, none identical.

The gains here are real, and they are personal. Salesforce's 2026 State of Sales, surveying more than 4,000 sellers, found research time down 34% and content creation down 36% for AI users. That is genuine leverage, and it is the same leverage most generative AI use in sales delivers at the individual level. What it cannot do is show up in team numbers, because nothing is shared.

🔹 Rung 3, where standards start to hold

Signature symptom: a new hire can run the standard account-research workflow in week one without shadowing anyone.

What it is good at is consistency without supervision cost. It is also where methodology compliance stops being a slide and starts being a system, which is the whole argument behind automating MEDDIC, BANT, and SPICED scoring from calls. What it cannot do is act unprompted. Every run still waits for a human to start it.

🔹 Rung 4, and what it actually demands

Signature symptom: work completes overnight and someone reviews it in the morning.

What it is good at is volume of repeated judgement. What it cannot do is compensate for undocumented process. An agent inherits your history, not your standards.

⭐ The shape of the gap, not its size

Notice what changes between rungs. One to two is a skill anyone can pick up in a week. Three to four is a capability question.

Two to three is neither. It is a management change, and that is why it stalls.

Q3. How do you tell which rung you are actually on? [toc=3. Self-Assessment]

Score five dimensions from one to five: workflow standardisation, data quality and trust, inspection and accountability, cross-team alignment, and AI readiness. Then apply the rule most vendor assessments bury. Your rung is your lowest score, not your average, because the weakest dimension caps what agents can reliably do. Four practical tests settle it: can a new hire run your best rep's workflow unaided, is there one place approved skills live, does anyone's job include maintaining them, and does any AI task finish without a human starting it.

📊 The five dimensions, scored honestly

Outreach publishes a comparable public scorecard built on these same five dimensions, which is a useful cross-check if you want a second opinion on your own scoring.

Five Scoring Dimensions: What a 2 and a 4 Look Like in Practice
DimensionA score of 2 looks likeA score of 4 looks like
Workflow standardisationEvery rep researches accounts their own wayOne documented way, followed by most of the team
Data quality and trustDuplicate accounts, activity on the wrong opportunityActivity maps reliably to the right account and deal
Inspection and accountabilityNobody reviews AI outputOutput is sampled and corrected on a cadence
Cross-team alignmentSales, CS, and RevOps use different tools and promptsShared context and shared standards across functions
AI readinessProcess lives in people's headsProcess is written down, exceptions included

The second row is the one that quietly decides everything else, which is why CRM data quality automation for RevOps is usually the first real project rather than the last.

✅ The four-question version, if you have two minutes

Answer yes or no. Be strict.

  1. Can a new hire run your best rep's AI workflow without asking that rep?
  2. Is there one place where approved prompts or skills live?
  3. Does anyone's job description include maintaining them?
  4. Does any AI task complete without a human starting it?

Four noes puts you on Rung 1. Yes to the first only is Rung 2. Yes to the first three is Rung 3. Yes to all four is Rung 4.

⚠️ Score the floor, not the ceiling

Here is the part that trips up senior readers, and I include myself in that. Executives self-assess on their best rep. The organisation runs at the level of its median.

So take your lowest dimension score and treat that as your rung. If data quality sits at 2 while everything else sits at 4, you are a 2. Agents acting on mis-mapped records do not fail loudly. They act confidently on the wrong deal, and every correction costs more than the run saved.

❌ The misread I see most often

A company buys at Rung 4 and operates at Rung 2. The contract says autonomous. The behaviour says individual prompting with better branding.

That is not a vendor failure. It is the operating model never having changed to match the purchase. Re-score these five dimensions each quarter, and that drift becomes visible before your renewal does. Teams running a formal revenue intelligence platform comparison should bring this score into the evaluation rather than after it.

Q4. Which rung is right for you, and when is staying put the correct call? [toc=4. Reasons to Stay]

Rung 2 is genuinely correct for a small team of capable operators, because coordination overhead would cost more than the inconsistency it removes. Rung 3 is the right ceiling for most companies: shared standards, no supervision burden, no new role to hire. Rung 4 earns its place only where the volume of repeated judgement exceeds what a team can supervise by hand. Moving up is an organisational cost, not a licence cost, which is why buying more tools never moves you.

⚠️ The assumption I want to take away from you

Every maturity model implies that higher is better. Read four of them and you will feel behind by lunchtime.

The pressure is real, and it is loud. Gartner projects AI agents will outnumber human sellers by ten to one by 2028, with 81% of B2B sales teams already using AI in some form during 2026. That is a forecast about the market, though. It is not an instruction about your org.

💸 What climbing too early actually costs

I have watched this go wrong in a specific, boring way. A team approves a Rung 4 project without a Rung 3 foundation. Six months later there are agents nobody monitors and a playbook nobody wrote.

The cost lands in three places.

  • Staffing. A monitoring role you have not hired, absorbed by RevOps on top of their day job.
  • Trust. Reps stop using output they have caught being wrong twice.
  • Time. The documentation work you skipped still has to happen, now under pressure.

That third cost is the one finance never sees coming, and it is a large part of why agentic AI implementation depends on data architecture rather than enthusiasm.

✅ Honest reasons to stay where you are

So let me give each rung a real defence.

Stay on Rung 2 if you have a handful of genuinely capable operators and a process that changes monthly. Standardising something that keeps moving is waste. ChatGPT and Claude are the correct tools here, and I would keep using both.

Stay on Rung 3 if your work is varied, your judgement calls are one-offs, and a human reviewing output is not your bottleneck. This is the right ceiling for most companies, and I do not say that as a throwaway concession. Most revenue orgs would get more value from a maintained shared library than from any agent deployment they could staff this year.

Move to Rung 4 only when the same judgement is being made hundreds of times a week and supervision, not capability, is your constraint. The CRO guide to agentic AI in revenue intelligence walks through what that supervision load looks like before you commit to it.

Three stacked layers showing when to stay on rung two, rung three, or move to rung four of AI adoption
Rung 3 is the right ceiling for most revenue teams, which is why every layer here carries a reason to stay put.

⭐ The criterion, restated as a decision

Put the two numbers side by side and pick. Coordination cost versus inconsistency cost.

If coordinating is cheaper than the inconsistency you are absorbing, climb. If not, stay, and spend the money on something else. Nobody has ever lost a quarter by refusing to climb a rung.

❌ What being wrong looks like in either direction

Climb too early and you get unsupervised agents plus an unwritten playbook. Stay too long and you get a team where your best rep's method never becomes anyone else's.

Of those two failures, I think the second is more common and the first is more expensive. I might be weighting that from where I sit, so test it against your own numbers rather than mine.

Q5. Why do most revenue teams get stuck after individual AI use? [toc=5. The Stall Point]

Because the move from Rung 2 to Rung 3 is not a skill, it is a management decision. Someone has to own a shared library, decide what becomes standard, retire what does not work, and maintain it as the process changes. That job has usually not been given to anyone. So individual capability compounds while shared capability stays at zero. The organisation gets individually efficient and organisationally unchanged. The unlock costs nothing. Name an owner for the shared skill library.

⚠️ The scene I keep walking into

A VP of Sales pulls up a prompt her top AE built. It writes a genuinely good pre-call brief. She is delighted, and she should be.

Then I ask how many other reps use it. The answer is one. Her team number has not moved in two quarters, and she cannot work out why. It is the same pattern I see when coaching at scale using AI stays trapped inside one manager's habits.

❌ The default response, and what it costs

The usual reaction is to buy something. Another pilot, another seat block, another vendor evaluation. I have sat on the other side of those calls, and I can tell you the purchase rarely fails on capability.

It fails because the tool arrives into the same operating model. Gartner's 2026 data puts 81% of B2B sales teams using AI already, while fewer than 40% of sellers report a measurable productivity gain. That spread is not a software problem. Buying again does not close it, which is why consolidating the revenue tech stack usually beats adding to it.

A senior revenue leader put the distinction to me better than I had managed to. His organisation had been in an AI experimentation phase for a year. They had gained individual efficiency, he said, and no organisational efficiency at all. The mandate had been personal productivity, which is fine, and it is also a ceiling.

🔍 What actually decays, and how fast

Here is the mechanical reason it stalls. A shared library without an owner is not a library, it is a folder.

  • Week one. Someone posts four good prompts. Everyone is pleased.
  • Week three. The qualification criteria change in a pipeline review. Nobody edits the folder.
  • Week six. Two reps have quietly forked their own versions, because the shared one is now wrong.
  • Week ten. People stop opening it, and you are back on Rung 2 with extra steps.
Four-stage timeline showing a shared AI prompt library decaying from adoption to abandonment without an owner
Nothing breaks and no decision gets made. The library simply decays, which is why one named owner is the whole fix.

Nothing broke. No decision was made. Ownership was simply never assigned.

"I found the AI tracker setup to be quite difficult, especially concerning the user interface when setting up keywords or smart trackers. Moreover, I cannot download all the data myself unless we upgrade the plan, which isn't ideal and results in me not fully utilizing Gong."
Verified user, Sales Professional Gong G2 Verified Review [3 Oct 2025]

That review is about configuration friction, and it points at the same thing. When setup is hard enough that only one person does it, the capability never becomes the team's. The same complaint shows up across how Gong smart trackers are configured, where the setup burden sits with one admin.

✅ The one move, and what the owner actually decides

Give one person the shared library. Not a committee, not a workstream, one name.

Their job has four decisions in it. What becomes standard. What gets retired. Who can edit. When it gets reviewed after a process change. Two hours a week covers it at most mid-market scales.

⭐ The tell that you have crossed over

You will know it happened without running a survey. A new hire in week one runs the standard account-research workflow, start to finish, without asking your best rep for anything.

That is Rung 3. It cost a name on a job description, and I still find it hard to convince people the fix is that small. If ramp speed is the thing you are actually measuring, new hire ramp time and onboarding is the cleanest place to see the difference.

Q6. What has to be true before agents can run in production? [toc=6. Production Prerequisites]

Four things. Your process is written down as it actually runs, exceptions included. Your data is trustworthy enough that activity maps to the right account and opportunity, because an agent acting on a mis-mapped record compounds the error with full confidence. Permissions are explicit, covering which actions an agent takes alone and which need approval. And disclosure is wired in, because EU rules now require an agent to identify itself and who it acts for. A company whose process is undocumented cannot reach Rung 3, never mind Rung 4.

Ascending four-step diagram of prerequisites for running AI agents in production in a revenue team
These four prerequisites stack in order, and the first one is management work that no vendor can sell you.

1️⃣ Documented process, as it actually runs

An agent inherits your history, not your standards. If your qualification rules, discount triggers, and legal handoff live in people's heads, the agent will invent a version of them.

Verify it in an afternoon. Ask three managers, separately, when a deal goes to legal. If you get three answers, you have found your first piece of work.

⚠️ Why this one has no software in it

Writing the playbook is a management project. There is no tool that does it for you, and saying otherwise would be the easiest lie in this category.

The reason this matters so much is failure rate. The widely repeated figure in agentic AI is that most projects fail for the same reason, which is that the processes were never documented. Documented is also not the same as alive. Playbooks drift within weeks because exceptions get added verbally and never written back.

2️⃣ Data you can actually act on

This is the hard ceiling, not a footnote. Salesforce's 2026 State of Sales, across more than 4,000 sellers, names data quality and admin friction as the top blockers for teams not seeing agent ROI.

The failure mode is specific. An account has five open opportunities, or three duplicate records, and the meeting gets attached to the wrong one. The agent then writes a confident update on a deal that does not exist. This is exactly why a CRM data strategy tied to revenue predictability has to precede any agent rollout.

"The conversation intelligence tool is lacking, and we don't have the context of the deals against the conversation intelligence findings. The CRM writeback is not good; we cannot send MEDDIC values back to Salesforce or update fields in Salesforce from the conversation intelligence."
Verified user, Revenue Operations Clari G2 Verified Review [13 Jul 2026]

MEDDIC is a qualification checklist covering metrics, buyer, decision process, and pain. That review is describing a broken write path, which is exactly the prerequisite failing. If you want the underlying framework rather than the tooling complaint, the MEDDIC sales methodology explains what should be landing in those fields.

3️⃣ Permissions, written as a list

Split every agent action into two buckets before anything runs.

Agent Permission Buckets: What Needs Approval and What Does Not
BucketRuleExamples
Ask before actingAnything a customer sees, or anything that changes deal economicsFirst outbound email, pricing changes, stage moves
Act without askingReversible internal work with a visible audit trailCRM field updates, meeting summaries, task creation

Forrester's October 2025 predictions expected 20% of sellers to use agent led negotiation during 2026. Once agents touch commercial terms, an unwritten permission list becomes a real exposure, and the governance side of that is covered properly in the RevOps evaluation of AI CRM trust and governance risk.

4️⃣ Disclosure, with dates you should know

Article 50 of the EU AI Act became enforceable on 2 August 2026, and the Commission's July 2026 guidelines confirm agents must disclose their AI nature and the entity they act for. Systems already in place have until 2 December 2026 to implement marking.

The higher risk obligations in Annex III were deferred to 2 December 2027 by Regulation (EU) 2026/1744. So the disclosure work is live now, and the heavier compliance lift is not. Wire the disclosure into the system, not into the copy, and log the event.

⏰ Honest sequencing

If you are missing the first two, Rung 4 is not a roadmap item this year. Start with the playbook and a duplicate record audit. Both are unglamorous, and both are cheaper than a failed agent rollout. Teams that want the audit as a repeatable routine should look at CRM data quality automation for RevOps rather than a one-off cleanup sprint.

Q7. Who does this work, and can you reach Rung 4 without a GTM engineer? [toc=7. Who Owns Rung 4]

No, not without someone doing that job. Rung 4 means agents run while a GTM engineer monitors and optimises them, deciding what runs unsupervised, watching where output drifts, and updating context when the process changes. There are two honest routes. Grow the role from RevOps, which is slower but keeps the knowledge in house. Or buy it as a service. Oliv AI publishes named in house forward deployed GTM engineers who implement directly with a customer's stack, with a first agent live on real deals by week two. The role is a prerequisite either way.

⚠️ The objection, in the reader's own words

"We do not have a GTM engineer and we cannot hire one." I hear that on most calls where this ladder comes up, and it is a fair blocker, not an excuse.

A GTM engineer, to define it plainly, is the person who turns revenue processes into reliable agents and then keeps them reliable. The title is new. The work is not, and much of it already sits inside the job described in building a revenue operations function.

❌ What happens when nobody owns it

Agents get deployed and then quietly go wrong. Not dramatically wrong, which would be easier. Slightly wrong, in ways that take a quarter to surface.

The cost lands on RevOps, on top of their existing job. I have watched a RevOps lead absorb agent monitoring informally for four months, then hand the whole thing back. That was not a tooling failure. It was an unassigned role.

✅ The two routes, and the fork underneath them

Route one is to grow the role from RevOps. Slower, and the institutional knowledge stays with you. Route two is to buy the capability as a service.

That is the same build versus buy fork that decides most Rung 4 projects, and it deserves its own analysis rather than a paragraph here. If you are working through the data and architecture side of it, the RevOps guide to implementing agentic AI covers the ground properly, and the commercial side sits in build versus buy for revenue AI.

⭐ How this works in practice

Oliv AI publishes in house, forward deployed GTM engineers who implement directly with your stack, with named engineer profiles on the site and no third party implementation partner. The published sequence is discovery on day one, connected systems in week one, a first agent live on real deals with real reps in week two, and measurable ROI on that agent by week four. We built it this way because the role is scarce, not because implementation is easy to sell. The caveat belongs here too, and an executive checking the source will find it anyway: Oliv is the least publicly proven vendor named in this piece, with no G2, Capterra, or TrustRadius presence, and case studies that sit behind an email form. On a framework article that matters less commercially than it does to your ability to verify what I am claiming, so I would rather state it than have you discover it. Oliv is not the only route to Rung 4, and any vendor offering this should be asked for one named production customer before a second meeting.

💰 The cost note nobody mentions

There is a practical first step that does not need a staffing decision at all. Oliv AI's published price ladder starts at Amplify $0, then Converse $19, Sell $49, and Grow $79 per user per month, with agent actions billed at $0.01 per credit.

Free seats matter more than they sound here. The executives, product people, and ops folk who read an article like this can sit on a shared platform before anyone approves a seat block, and shared access is the actual mechanic of Rung 3.

⏰ What I would do if the role feels impossible

Do not hire for it first. Pick one repeated judgement, assign one existing person two hours a week to supervise it, and see whether the correction rate falls.

If it does, you have your business case. If it does not, you have saved yourself a hire, and Rung 3 was the right ceiling after all.

Q8. How do you measure whether you are actually moving up? [toc=8. Measuring Movement]

Stop reporting AI adoption percentage. At 81% industry wide it distinguishes nothing. Measure four things instead: share of AI work running through shared skills rather than personal prompts, time from a process change to the shared library reflecting it, workflows a new hire can run in week one unaided, and at Rung 4 the ratio of agent actions completed to agent actions corrected. Correction rate is your observability measure. An agent you correct constantly is a Rung 3 workflow wearing a Rung 4 label.

📊 The four numbers worth reporting

Four Metrics That Track Movement Up the AI Adoption Ladder
MetricHow to collect itWhat a bad number means
Shared skill shareCount runs from the shared library against total AI runs, monthlyIndividual capability is compounding, shared capability is not
Library lagDays from a process change to the library reflecting itNobody owns maintenance, so you are drifting back to Rung 2
Week one workflowsAsk each new hire what they ran unaided in week oneYour standards live in people, not in the system
Agent correction rateCorrected actions divided by completed actionsSupervision cost is eating the automation gain

All four are countable this quarter, without new tooling. Each moves only when the operating model moves, which is the point. If your board wants these sitting next to the commercial numbers, revenue performance analytics is where they belong.

⭐ Why correction rate is the honest one

Every other agent metric can be gamed by narrowing scope. Run an agent on one trivial task and your success rate looks superb.

Correction rate resists that. If your team fixes four out of ten outputs, you have a supervised workflow, not an autonomous one. I would rather a client told me their correction rate was 40% than showed me a dashboard with a green tick.

⚠️ How to read the benchmark someone will hand you

Pavilion's 2026 GTM benchmark reports that AI powered teams generate roughly 40% more pipeline per rep. You will see that number in a deck within a month, probably in a vendor deck.

Treat it as a hypothesis, not a target. Benchmarks built on self selected respondents carry survivorship bias, which means the teams that struggled often did not report. Publish your own before and after per rep number instead. It will be less impressive and far more useful, and a revenue intelligence ROI calculator is a better starting point than someone else's average.

❌ Three numbers never to report

Some figures look rigorous and prove nothing. I would keep all three out of your board pack.

  1. AI adoption percentage. With 81% of B2B sales teams already using AI in 2026, this measures nothing about your operating model.
  2. Industry stage distributions. Any claim that a specific percentage of companies sit on a given rung is unfalsifiable. I do not have that number, and I have not seen anyone who does.
  3. Typical time between levels. The 2 to 3 move can take a week if someone is assigned, or never if nobody is.

✅ Make it a quarterly habit

Re score the five dimensions from the self assessment every quarter, on the same day you review pipeline coverage. It takes ten minutes.

If you want a shared vocabulary with your board, G2's public taxonomy runs from ad hoc and aware through developing, mature, leading, and transformative. Map your rungs to it once, then stop arguing about labels and go and fix your lowest score. Where that score is forecast discipline rather than tooling, running evidence based forecast commits is the fastest thing to fix next.

Where is your revenue org on AI right now? [toc=0. Introduction]

Where is your revenue org on AI right now? Not which tools you own. What stage you operate at.

Most executives I ask cannot answer that, and they are not being evasive. They can list every licence. They have seen good demos from their own reps. What they do not have is a way to think about sequence, which is the thing a board actually asks for.

So let me concede something before I hand you a framework. Most AI maturity models are sales tools. Four rungs, and the author's product sits on the top one, and any operator who has read two of them knows it.

I am going to do the opposite here, because the framework is only useful if it can tell you to stay put. Every rung in this article gets an honest reason to remain on it. Rung 3 is the right ceiling for most companies, and I would rather say that plainly than sell you a climb you cannot staff. If you want the longer version of that argument, the honest read on what AI agents can actually do for your team today covers the same ground for a VP of Sales.

The test I will use does not involve any vendor, including mine. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. That is the whole diagnostic.

One more thing worth naming early. The argument underneath all eight sections is that most revenue orgs are not behind on tooling at all. They are stuck between individual capability and shared capability, and that gap is organisational, which is exactly why buying more software never closes it. It is the same structural shift described in how revenue teams evolved from RevOps to intelligence to orchestration.

You will leave with two things. A rung you can defend, and one next move. For most readers that move costs nothing.

What is the one move you should make next? [toc=9. Your Next Move]

If you read only one line of this, make it this one: your rung is set by what your organisation has arranged, not by what it has licensed. Score the five dimensions honestly, take your lowest number, and accept that as your answer for this quarter. Then pick the single next move that sits directly above it, which for most teams reading this means giving one named person the shared skill library and two hours a week to maintain it. That is a management decision, it costs nothing, and it is the actual unlock between individual and shared capability. If your next move turns out to be the Rung 3 to Rung 4 jump, and you want to pressure test whether the supervision maths works at your volume, book a seven minute chat and bring your correction rate rather than your tool list.

Two adjacent reads, depending on which way your lowest score points. If the gap is ownership and structure, start with scaling revenue operations in a growth stage sales team. If the gap is whether agents belong in your motion at all, agentic AI for revenue execution works through the pipeline side of it, and the future of revenue intelligence sets the longer horizon.

✍🏼 About the author

Ishan Chhabra is the founder and CEO of Oliv AI, an AI native revenue intelligence and revenue orchestration platform for B2B revenue teams. He built Oliv's context graph, the infrastructure layer that resolves accounts, opportunities, and conversations across messy CRMs so agents can act on them safely.

He writes about what he sees working and failing inside revenue organisations adopting AI, including the parts that fail for reasons no vendor likes to publish. His work on AI agents for RevOps and the CRO view of ROI on a revenue platform covers the same territory in more depth.

Q1. Why should you trust a maturity model at all? [toc=1. Trusting the Model]

Most AI maturity models are sales artefacts. Four rungs, and the author's product sits on the top one. The test that survives that bias is a cost comparison, not a capability list. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. Below that line, coordination is pure overhead. Above it, inconsistency is already costing you deals.

⚠️ The board question nobody can answer

A CRO I spoke with last quarter had approved AI spend across three teams. Her reps were enthusiastic. Her SDRs had built prompts they swore by. Then her board asked a simple question: where are we, and what comes next?

She could name every tool. She could not name a stage. That gap is the reason this genre of article exists, and also the reason most of it is useless. It is the same gap I see when a team can list its stack but cannot describe how AI actually runs inside revenue operations.

❌ Why the genre deserves your suspicion

Here is the problem, said plainly. Maturity models are usually written by vendors, and the top rung is usually a description of the vendor's product. An experienced operator spots that in about four seconds, and rightly stops reading.

Believing one of those models has a real cost. You end up approving a transformation programme you cannot staff. You buy for a rung you do not operate at. Gartner's 2026 Hype Cycle placed agentic AI at the Peak of Inflated Expectations, with only 17% of organisations having actually deployed agents while more than 60% plan to within two years. That gap between intent and deployment is where budgets go to die, and it is why the build versus buy decision on revenue AI deserves more scrutiny than a maturity chart.

✅ The test that does not involve any vendor

So use a test that no software can tilt. Compare two costs.

  • Coordination cost. What it takes to agree a standard, write it down, and keep it current.
  • Inconsistency cost. What you lose when every rep runs a different version of the same task.

You belong on the rung where the first number is smaller than the second. That is the whole diagnostic. It works if you never buy anything.

Diagram comparing coordination cost against inconsistency cost to decide the correct GTM AI maturity rung
The diagnostic that survives vendor bias: you belong on the rung where coordination costs less than inconsistency.

⭐ What I will do differently here

Every rung in this article gets an honest reason to stay, not just a reason to climb. I will say the thing vendors do not: Rung 3 is the right ceiling for most companies. Shared standards, no supervision burden, no new role to hire.

I am also going to be specific about where each move gets expensive, because the expense is rarely the licence. AI spend gets approved per tool. Value accrues per operating model. That mismatch is how a company ends up buying at Rung 4 while still operating at Rung 2, with nothing in the numbers to show it.

⏰ What you should leave with

Two things, in about ten minutes of reading. A rung you can defend to a board, and one next move you can start on Monday.

Not a programme. One move. For most readers it turns out to cost nothing at all, which is the part I did not expect when I started mapping this.

Q2. What are the four rungs of the GTM AI adoption ladder, and what blocks each move up? [toc=2. The Four Rungs]

Four rungs. One, everyone prompts individually, context copy-pasted each time, nothing reusable. Two, individuals build reusable skills and run them themselves, which is the most common rung. Three, skills are centralised, teams use the same ones daily, and someone owns them. Four, agents run in production on documented process, with GTM engineers monitoring and optimising. The jump from one to two is a skill. The jump from two to three is a management change. The jump from three to four is documented process plus supervision capacity. Rungs cannot be skipped.

Two definitions before the table, because these words get used loosely. A skill is a saved, reusable instruction set a person runs on demand. An agent is a process that runs without a person starting it. A GTM engineer is the person who builds, monitors, and maintains those agents. If the distinction between the last two is still fuzzy, the difference between AI agents built for sales teams and a saved prompt is exactly what this ladder is measuring.

📊 The ladder, with the constraint that blocks each move

The GTM AI Adoption Ladder: Four Rungs, Constraints, and Next Actions
RungWhat is observably trueWho does the workBlocking constraintSingle next action
01. ChatGPT for everyonePeople prompt ad hoc. Context is pasted in fresh each time. Nothing is saved.Individual reps, unevenlyNobody knows which prompts workAsk three reps to share their best prompt in one doc
02. Claude and DIY skillsIndividuals build reusable skills and run them personally. Quality varies by person.Capable individualsNo owner for a shared libraryName an owner for the shared skill library
03. Centralised skillsOne library, agreed standards, teams use the same skills daily, someone maintains them.The team, on a standardProcess is undocumented, supervision capacity is zeroWrite the playbook as it actually runs
04. Autonomous agents in productionAgents run on documented process. Humans monitor, correct, and optimise.Agents, supervised by a GTM engineerNobody owns agent monitoringAssign or buy the monitoring role

This ladder is Oliv AI's framework for organisational AI adoption, and it is deliberately built on what a company has arranged rather than what it has licensed. Two orgs on identical stacks routinely sit two rungs apart.

🔹 Rung 1, and why it is not nothing

Signature symptom: your best rep's workflow lives in their browser history.

What it is good at is discovery. People find out what AI is useful for, cheaply. What it cannot do is survive turnover. When that rep leaves, the capability leaves with them.

🔹 Rung 2, the most crowded rung in B2B

Signature symptom: three reps have three different research prompts, all decent, none identical.

The gains here are real, and they are personal. Salesforce's 2026 State of Sales, surveying more than 4,000 sellers, found research time down 34% and content creation down 36% for AI users. That is genuine leverage, and it is the same leverage most generative AI use in sales delivers at the individual level. What it cannot do is show up in team numbers, because nothing is shared.

🔹 Rung 3, where standards start to hold

Signature symptom: a new hire can run the standard account-research workflow in week one without shadowing anyone.

What it is good at is consistency without supervision cost. It is also where methodology compliance stops being a slide and starts being a system, which is the whole argument behind automating MEDDIC, BANT, and SPICED scoring from calls. What it cannot do is act unprompted. Every run still waits for a human to start it.

🔹 Rung 4, and what it actually demands

Signature symptom: work completes overnight and someone reviews it in the morning.

What it is good at is volume of repeated judgement. What it cannot do is compensate for undocumented process. An agent inherits your history, not your standards.

⭐ The shape of the gap, not its size

Notice what changes between rungs. One to two is a skill anyone can pick up in a week. Three to four is a capability question.

Two to three is neither. It is a management change, and that is why it stalls.

Q3. How do you tell which rung you are actually on? [toc=3. Self-Assessment]

Score five dimensions from one to five: workflow standardisation, data quality and trust, inspection and accountability, cross-team alignment, and AI readiness. Then apply the rule most vendor assessments bury. Your rung is your lowest score, not your average, because the weakest dimension caps what agents can reliably do. Four practical tests settle it: can a new hire run your best rep's workflow unaided, is there one place approved skills live, does anyone's job include maintaining them, and does any AI task finish without a human starting it.

📊 The five dimensions, scored honestly

Outreach publishes a comparable public scorecard built on these same five dimensions, which is a useful cross-check if you want a second opinion on your own scoring.

Five Scoring Dimensions: What a 2 and a 4 Look Like in Practice
DimensionA score of 2 looks likeA score of 4 looks like
Workflow standardisationEvery rep researches accounts their own wayOne documented way, followed by most of the team
Data quality and trustDuplicate accounts, activity on the wrong opportunityActivity maps reliably to the right account and deal
Inspection and accountabilityNobody reviews AI outputOutput is sampled and corrected on a cadence
Cross-team alignmentSales, CS, and RevOps use different tools and promptsShared context and shared standards across functions
AI readinessProcess lives in people's headsProcess is written down, exceptions included

The second row is the one that quietly decides everything else, which is why CRM data quality automation for RevOps is usually the first real project rather than the last.

✅ The four-question version, if you have two minutes

Answer yes or no. Be strict.

  1. Can a new hire run your best rep's AI workflow without asking that rep?
  2. Is there one place where approved prompts or skills live?
  3. Does anyone's job description include maintaining them?
  4. Does any AI task complete without a human starting it?

Four noes puts you on Rung 1. Yes to the first only is Rung 2. Yes to the first three is Rung 3. Yes to all four is Rung 4.

⚠️ Score the floor, not the ceiling

Here is the part that trips up senior readers, and I include myself in that. Executives self-assess on their best rep. The organisation runs at the level of its median.

So take your lowest dimension score and treat that as your rung. If data quality sits at 2 while everything else sits at 4, you are a 2. Agents acting on mis-mapped records do not fail loudly. They act confidently on the wrong deal, and every correction costs more than the run saved.

❌ The misread I see most often

A company buys at Rung 4 and operates at Rung 2. The contract says autonomous. The behaviour says individual prompting with better branding.

That is not a vendor failure. It is the operating model never having changed to match the purchase. Re-score these five dimensions each quarter, and that drift becomes visible before your renewal does. Teams running a formal revenue intelligence platform comparison should bring this score into the evaluation rather than after it.

Q4. Which rung is right for you, and when is staying put the correct call? [toc=4. Reasons to Stay]

Rung 2 is genuinely correct for a small team of capable operators, because coordination overhead would cost more than the inconsistency it removes. Rung 3 is the right ceiling for most companies: shared standards, no supervision burden, no new role to hire. Rung 4 earns its place only where the volume of repeated judgement exceeds what a team can supervise by hand. Moving up is an organisational cost, not a licence cost, which is why buying more tools never moves you.

⚠️ The assumption I want to take away from you

Every maturity model implies that higher is better. Read four of them and you will feel behind by lunchtime.

The pressure is real, and it is loud. Gartner projects AI agents will outnumber human sellers by ten to one by 2028, with 81% of B2B sales teams already using AI in some form during 2026. That is a forecast about the market, though. It is not an instruction about your org.

💸 What climbing too early actually costs

I have watched this go wrong in a specific, boring way. A team approves a Rung 4 project without a Rung 3 foundation. Six months later there are agents nobody monitors and a playbook nobody wrote.

The cost lands in three places.

  • Staffing. A monitoring role you have not hired, absorbed by RevOps on top of their day job.
  • Trust. Reps stop using output they have caught being wrong twice.
  • Time. The documentation work you skipped still has to happen, now under pressure.

That third cost is the one finance never sees coming, and it is a large part of why agentic AI implementation depends on data architecture rather than enthusiasm.

✅ Honest reasons to stay where you are

So let me give each rung a real defence.

Stay on Rung 2 if you have a handful of genuinely capable operators and a process that changes monthly. Standardising something that keeps moving is waste. ChatGPT and Claude are the correct tools here, and I would keep using both.

Stay on Rung 3 if your work is varied, your judgement calls are one-offs, and a human reviewing output is not your bottleneck. This is the right ceiling for most companies, and I do not say that as a throwaway concession. Most revenue orgs would get more value from a maintained shared library than from any agent deployment they could staff this year.

Move to Rung 4 only when the same judgement is being made hundreds of times a week and supervision, not capability, is your constraint. The CRO guide to agentic AI in revenue intelligence walks through what that supervision load looks like before you commit to it.

Three stacked layers showing when to stay on rung two, rung three, or move to rung four of AI adoption
Rung 3 is the right ceiling for most revenue teams, which is why every layer here carries a reason to stay put.

⭐ The criterion, restated as a decision

Put the two numbers side by side and pick. Coordination cost versus inconsistency cost.

If coordinating is cheaper than the inconsistency you are absorbing, climb. If not, stay, and spend the money on something else. Nobody has ever lost a quarter by refusing to climb a rung.

❌ What being wrong looks like in either direction

Climb too early and you get unsupervised agents plus an unwritten playbook. Stay too long and you get a team where your best rep's method never becomes anyone else's.

Of those two failures, I think the second is more common and the first is more expensive. I might be weighting that from where I sit, so test it against your own numbers rather than mine.

Q5. Why do most revenue teams get stuck after individual AI use? [toc=5. The Stall Point]

Because the move from Rung 2 to Rung 3 is not a skill, it is a management decision. Someone has to own a shared library, decide what becomes standard, retire what does not work, and maintain it as the process changes. That job has usually not been given to anyone. So individual capability compounds while shared capability stays at zero. The organisation gets individually efficient and organisationally unchanged. The unlock costs nothing. Name an owner for the shared skill library.

⚠️ The scene I keep walking into

A VP of Sales pulls up a prompt her top AE built. It writes a genuinely good pre-call brief. She is delighted, and she should be.

Then I ask how many other reps use it. The answer is one. Her team number has not moved in two quarters, and she cannot work out why. It is the same pattern I see when coaching at scale using AI stays trapped inside one manager's habits.

❌ The default response, and what it costs

The usual reaction is to buy something. Another pilot, another seat block, another vendor evaluation. I have sat on the other side of those calls, and I can tell you the purchase rarely fails on capability.

It fails because the tool arrives into the same operating model. Gartner's 2026 data puts 81% of B2B sales teams using AI already, while fewer than 40% of sellers report a measurable productivity gain. That spread is not a software problem. Buying again does not close it, which is why consolidating the revenue tech stack usually beats adding to it.

A senior revenue leader put the distinction to me better than I had managed to. His organisation had been in an AI experimentation phase for a year. They had gained individual efficiency, he said, and no organisational efficiency at all. The mandate had been personal productivity, which is fine, and it is also a ceiling.

🔍 What actually decays, and how fast

Here is the mechanical reason it stalls. A shared library without an owner is not a library, it is a folder.

  • Week one. Someone posts four good prompts. Everyone is pleased.
  • Week three. The qualification criteria change in a pipeline review. Nobody edits the folder.
  • Week six. Two reps have quietly forked their own versions, because the shared one is now wrong.
  • Week ten. People stop opening it, and you are back on Rung 2 with extra steps.
Four-stage timeline showing a shared AI prompt library decaying from adoption to abandonment without an owner
Nothing breaks and no decision gets made. The library simply decays, which is why one named owner is the whole fix.

Nothing broke. No decision was made. Ownership was simply never assigned.

"I found the AI tracker setup to be quite difficult, especially concerning the user interface when setting up keywords or smart trackers. Moreover, I cannot download all the data myself unless we upgrade the plan, which isn't ideal and results in me not fully utilizing Gong."
Verified user, Sales Professional Gong G2 Verified Review [3 Oct 2025]

That review is about configuration friction, and it points at the same thing. When setup is hard enough that only one person does it, the capability never becomes the team's. The same complaint shows up across how Gong smart trackers are configured, where the setup burden sits with one admin.

✅ The one move, and what the owner actually decides

Give one person the shared library. Not a committee, not a workstream, one name.

Their job has four decisions in it. What becomes standard. What gets retired. Who can edit. When it gets reviewed after a process change. Two hours a week covers it at most mid-market scales.

⭐ The tell that you have crossed over

You will know it happened without running a survey. A new hire in week one runs the standard account-research workflow, start to finish, without asking your best rep for anything.

That is Rung 3. It cost a name on a job description, and I still find it hard to convince people the fix is that small. If ramp speed is the thing you are actually measuring, new hire ramp time and onboarding is the cleanest place to see the difference.

Q6. What has to be true before agents can run in production? [toc=6. Production Prerequisites]

Four things. Your process is written down as it actually runs, exceptions included. Your data is trustworthy enough that activity maps to the right account and opportunity, because an agent acting on a mis-mapped record compounds the error with full confidence. Permissions are explicit, covering which actions an agent takes alone and which need approval. And disclosure is wired in, because EU rules now require an agent to identify itself and who it acts for. A company whose process is undocumented cannot reach Rung 3, never mind Rung 4.

Ascending four-step diagram of prerequisites for running AI agents in production in a revenue team
These four prerequisites stack in order, and the first one is management work that no vendor can sell you.

1️⃣ Documented process, as it actually runs

An agent inherits your history, not your standards. If your qualification rules, discount triggers, and legal handoff live in people's heads, the agent will invent a version of them.

Verify it in an afternoon. Ask three managers, separately, when a deal goes to legal. If you get three answers, you have found your first piece of work.

⚠️ Why this one has no software in it

Writing the playbook is a management project. There is no tool that does it for you, and saying otherwise would be the easiest lie in this category.

The reason this matters so much is failure rate. The widely repeated figure in agentic AI is that most projects fail for the same reason, which is that the processes were never documented. Documented is also not the same as alive. Playbooks drift within weeks because exceptions get added verbally and never written back.

2️⃣ Data you can actually act on

This is the hard ceiling, not a footnote. Salesforce's 2026 State of Sales, across more than 4,000 sellers, names data quality and admin friction as the top blockers for teams not seeing agent ROI.

The failure mode is specific. An account has five open opportunities, or three duplicate records, and the meeting gets attached to the wrong one. The agent then writes a confident update on a deal that does not exist. This is exactly why a CRM data strategy tied to revenue predictability has to precede any agent rollout.

"The conversation intelligence tool is lacking, and we don't have the context of the deals against the conversation intelligence findings. The CRM writeback is not good; we cannot send MEDDIC values back to Salesforce or update fields in Salesforce from the conversation intelligence."
Verified user, Revenue Operations Clari G2 Verified Review [13 Jul 2026]

MEDDIC is a qualification checklist covering metrics, buyer, decision process, and pain. That review is describing a broken write path, which is exactly the prerequisite failing. If you want the underlying framework rather than the tooling complaint, the MEDDIC sales methodology explains what should be landing in those fields.

3️⃣ Permissions, written as a list

Split every agent action into two buckets before anything runs.

Agent Permission Buckets: What Needs Approval and What Does Not
BucketRuleExamples
Ask before actingAnything a customer sees, or anything that changes deal economicsFirst outbound email, pricing changes, stage moves
Act without askingReversible internal work with a visible audit trailCRM field updates, meeting summaries, task creation

Forrester's October 2025 predictions expected 20% of sellers to use agent led negotiation during 2026. Once agents touch commercial terms, an unwritten permission list becomes a real exposure, and the governance side of that is covered properly in the RevOps evaluation of AI CRM trust and governance risk.

4️⃣ Disclosure, with dates you should know

Article 50 of the EU AI Act became enforceable on 2 August 2026, and the Commission's July 2026 guidelines confirm agents must disclose their AI nature and the entity they act for. Systems already in place have until 2 December 2026 to implement marking.

The higher risk obligations in Annex III were deferred to 2 December 2027 by Regulation (EU) 2026/1744. So the disclosure work is live now, and the heavier compliance lift is not. Wire the disclosure into the system, not into the copy, and log the event.

⏰ Honest sequencing

If you are missing the first two, Rung 4 is not a roadmap item this year. Start with the playbook and a duplicate record audit. Both are unglamorous, and both are cheaper than a failed agent rollout. Teams that want the audit as a repeatable routine should look at CRM data quality automation for RevOps rather than a one-off cleanup sprint.

Q7. Who does this work, and can you reach Rung 4 without a GTM engineer? [toc=7. Who Owns Rung 4]

No, not without someone doing that job. Rung 4 means agents run while a GTM engineer monitors and optimises them, deciding what runs unsupervised, watching where output drifts, and updating context when the process changes. There are two honest routes. Grow the role from RevOps, which is slower but keeps the knowledge in house. Or buy it as a service. Oliv AI publishes named in house forward deployed GTM engineers who implement directly with a customer's stack, with a first agent live on real deals by week two. The role is a prerequisite either way.

⚠️ The objection, in the reader's own words

"We do not have a GTM engineer and we cannot hire one." I hear that on most calls where this ladder comes up, and it is a fair blocker, not an excuse.

A GTM engineer, to define it plainly, is the person who turns revenue processes into reliable agents and then keeps them reliable. The title is new. The work is not, and much of it already sits inside the job described in building a revenue operations function.

❌ What happens when nobody owns it

Agents get deployed and then quietly go wrong. Not dramatically wrong, which would be easier. Slightly wrong, in ways that take a quarter to surface.

The cost lands on RevOps, on top of their existing job. I have watched a RevOps lead absorb agent monitoring informally for four months, then hand the whole thing back. That was not a tooling failure. It was an unassigned role.

✅ The two routes, and the fork underneath them

Route one is to grow the role from RevOps. Slower, and the institutional knowledge stays with you. Route two is to buy the capability as a service.

That is the same build versus buy fork that decides most Rung 4 projects, and it deserves its own analysis rather than a paragraph here. If you are working through the data and architecture side of it, the RevOps guide to implementing agentic AI covers the ground properly, and the commercial side sits in build versus buy for revenue AI.

⭐ How this works in practice

Oliv AI publishes in house, forward deployed GTM engineers who implement directly with your stack, with named engineer profiles on the site and no third party implementation partner. The published sequence is discovery on day one, connected systems in week one, a first agent live on real deals with real reps in week two, and measurable ROI on that agent by week four. We built it this way because the role is scarce, not because implementation is easy to sell. The caveat belongs here too, and an executive checking the source will find it anyway: Oliv is the least publicly proven vendor named in this piece, with no G2, Capterra, or TrustRadius presence, and case studies that sit behind an email form. On a framework article that matters less commercially than it does to your ability to verify what I am claiming, so I would rather state it than have you discover it. Oliv is not the only route to Rung 4, and any vendor offering this should be asked for one named production customer before a second meeting.

💰 The cost note nobody mentions

There is a practical first step that does not need a staffing decision at all. Oliv AI's published price ladder starts at Amplify $0, then Converse $19, Sell $49, and Grow $79 per user per month, with agent actions billed at $0.01 per credit.

Free seats matter more than they sound here. The executives, product people, and ops folk who read an article like this can sit on a shared platform before anyone approves a seat block, and shared access is the actual mechanic of Rung 3.

⏰ What I would do if the role feels impossible

Do not hire for it first. Pick one repeated judgement, assign one existing person two hours a week to supervise it, and see whether the correction rate falls.

If it does, you have your business case. If it does not, you have saved yourself a hire, and Rung 3 was the right ceiling after all.

Q8. How do you measure whether you are actually moving up? [toc=8. Measuring Movement]

Stop reporting AI adoption percentage. At 81% industry wide it distinguishes nothing. Measure four things instead: share of AI work running through shared skills rather than personal prompts, time from a process change to the shared library reflecting it, workflows a new hire can run in week one unaided, and at Rung 4 the ratio of agent actions completed to agent actions corrected. Correction rate is your observability measure. An agent you correct constantly is a Rung 3 workflow wearing a Rung 4 label.

📊 The four numbers worth reporting

Four Metrics That Track Movement Up the AI Adoption Ladder
MetricHow to collect itWhat a bad number means
Shared skill shareCount runs from the shared library against total AI runs, monthlyIndividual capability is compounding, shared capability is not
Library lagDays from a process change to the library reflecting itNobody owns maintenance, so you are drifting back to Rung 2
Week one workflowsAsk each new hire what they ran unaided in week oneYour standards live in people, not in the system
Agent correction rateCorrected actions divided by completed actionsSupervision cost is eating the automation gain

All four are countable this quarter, without new tooling. Each moves only when the operating model moves, which is the point. If your board wants these sitting next to the commercial numbers, revenue performance analytics is where they belong.

⭐ Why correction rate is the honest one

Every other agent metric can be gamed by narrowing scope. Run an agent on one trivial task and your success rate looks superb.

Correction rate resists that. If your team fixes four out of ten outputs, you have a supervised workflow, not an autonomous one. I would rather a client told me their correction rate was 40% than showed me a dashboard with a green tick.

⚠️ How to read the benchmark someone will hand you

Pavilion's 2026 GTM benchmark reports that AI powered teams generate roughly 40% more pipeline per rep. You will see that number in a deck within a month, probably in a vendor deck.

Treat it as a hypothesis, not a target. Benchmarks built on self selected respondents carry survivorship bias, which means the teams that struggled often did not report. Publish your own before and after per rep number instead. It will be less impressive and far more useful, and a revenue intelligence ROI calculator is a better starting point than someone else's average.

❌ Three numbers never to report

Some figures look rigorous and prove nothing. I would keep all three out of your board pack.

  1. AI adoption percentage. With 81% of B2B sales teams already using AI in 2026, this measures nothing about your operating model.
  2. Industry stage distributions. Any claim that a specific percentage of companies sit on a given rung is unfalsifiable. I do not have that number, and I have not seen anyone who does.
  3. Typical time between levels. The 2 to 3 move can take a week if someone is assigned, or never if nobody is.

✅ Make it a quarterly habit

Re score the five dimensions from the self assessment every quarter, on the same day you review pipeline coverage. It takes ten minutes.

If you want a shared vocabulary with your board, G2's public taxonomy runs from ad hoc and aware through developing, mature, leading, and transformative. Map your rungs to it once, then stop arguing about labels and go and fix your lowest score. Where that score is forecast discipline rather than tooling, running evidence based forecast commits is the fastest thing to fix next.

Where is your revenue org on AI right now? [toc=0. Introduction]

Where is your revenue org on AI right now? Not which tools you own. What stage you operate at.

Most executives I ask cannot answer that, and they are not being evasive. They can list every licence. They have seen good demos from their own reps. What they do not have is a way to think about sequence, which is the thing a board actually asks for.

So let me concede something before I hand you a framework. Most AI maturity models are sales tools. Four rungs, and the author's product sits on the top one, and any operator who has read two of them knows it.

I am going to do the opposite here, because the framework is only useful if it can tell you to stay put. Every rung in this article gets an honest reason to remain on it. Rung 3 is the right ceiling for most companies, and I would rather say that plainly than sell you a climb you cannot staff. If you want the longer version of that argument, the honest read on what AI agents can actually do for your team today covers the same ground for a VP of Sales.

The test I will use does not involve any vendor, including mine. You are on the right rung when the cost of coordinating AI use is lower than the cost of inconsistent AI use. That is the whole diagnostic.

One more thing worth naming early. The argument underneath all eight sections is that most revenue orgs are not behind on tooling at all. They are stuck between individual capability and shared capability, and that gap is organisational, which is exactly why buying more software never closes it. It is the same structural shift described in how revenue teams evolved from RevOps to intelligence to orchestration.

You will leave with two things. A rung you can defend, and one next move. For most readers that move costs nothing.

What is the one move you should make next? [toc=9. Your Next Move]

If you read only one line of this, make it this one: your rung is set by what your organisation has arranged, not by what it has licensed. Score the five dimensions honestly, take your lowest number, and accept that as your answer for this quarter. Then pick the single next move that sits directly above it, which for most teams reading this means giving one named person the shared skill library and two hours a week to maintain it. That is a management decision, it costs nothing, and it is the actual unlock between individual and shared capability. If your next move turns out to be the Rung 3 to Rung 4 jump, and you want to pressure test whether the supervision maths works at your volume, book a seven minute chat and bring your correction rate rather than your tool list.

Two adjacent reads, depending on which way your lowest score points. If the gap is ownership and structure, start with scaling revenue operations in a growth stage sales team. If the gap is whether agents belong in your motion at all, agentic AI for revenue execution works through the pipeline side of it, and the future of revenue intelligence sets the longer horizon.

✍🏼 About the author

Ishan Chhabra is the founder and CEO of Oliv AI, an AI native revenue intelligence and revenue orchestration platform for B2B revenue teams. He built Oliv's context graph, the infrastructure layer that resolves accounts, opportunities, and conversations across messy CRMs so agents can act on them safely.

He writes about what he sees working and failing inside revenue organisations adopting AI, including the parts that fail for reasons no vendor likes to publish. His work on AI agents for RevOps and the CRO view of ROI on a revenue platform covers the same territory in more depth.

FAQ's

What is an AI maturity model for revenue teams?

An AI maturity model for revenue teams places a sales, CS, or RevOps organisation on a staged ladder that runs from one-off individual prompting through to autonomous agents running in production. It measures the operating model built around AI, not the number of licences purchased.

Most credible versions score the same underlying dimensions:

  • Workflow standardisation. Does everyone do the same task the same way?
  • Data quality and trust. Does activity map to the right account and opportunity?
  • Inspection and accountability. Is AI output reviewed on a cadence, by someone named?
  • Cross-team alignment. Do sales, CS, and RevOps share standards or run separate stacks?
  • AI readiness. Is the process written down, exceptions included?

The important structural point is that rungs cannot be skipped. A company can buy software built for the top rung and still operate two rungs below it, because purchase changes the stack while the rung is set by what the organisation has arranged.

Used well, the model answers one question for a board: where are we, and what is the single next move. Used badly, it becomes a reason to approve a transformation programme nobody can staff. If you want the wider shift this sits inside, read how revenue teams evolved from RevOps to intelligence to orchestration.

How do I know what stage of AI adoption we are at?

Answer four yes or no questions, and be strict about each one.

  • Can a new hire run your best rep's AI workflow without asking that rep?
  • Is there one place where approved prompts or skills live?
  • Does anyone's job description include maintaining them?
  • Does any AI task complete without a human starting it?

Four noes puts you on rung one. Yes to the first only is rung two. Yes to the first three is rung three. Yes to all four is rung four.

Then apply the rule most vendor assessments bury. Score the five dimensions from one to five, and take your lowest score as your rung, not your average. The weakest dimension caps what agents can reliably do, so a team scoring four on standardisation and two on data quality is a two.

The common misread is self-assessing on your best rep. Organisations run at the level of their median, not their star. Re-score every quarter on the same day you review pipeline coverage, and drift shows up before your renewal does. For the RevOps-side view of what that scoring exposes, see our revenue intelligence platform comparison for RevOps.

Why do most teams get stuck after individual AI use?

Because moving from individual reusable skills to shared, centralised skills is a management decision rather than a technical one. Someone has to own a shared library, decide what becomes standard, retire what stops working, and update it when the process changes. That job is rarely assigned to anyone.

The decay is predictable and fast:

  • Week one. Someone posts four good prompts and the team is pleased.
  • Week three. Qualification criteria change in a pipeline review. Nobody edits the folder.
  • Week six. Two reps fork private versions, because the shared one is now wrong.
  • Week ten. People stop opening it, and you are back where you started with extra steps.

Nothing broke and no decision was made. Ownership was simply never given. That is why individual capability compounds while team numbers stay flat, and why buying another tool does not close the gap. The purchase lands into the same operating model.

The unlock costs nothing: name one person, not a committee, and give them roughly two hours a week. Oliv AI sees the same pattern across mid-market deployments, where the teams that progress are the ones that assigned an owner before they bought anything. For the structural version of this fix, read building a revenue operations function.

What is the difference between AI skills and AI agents?

The difference is who starts the work.

  • A skill is a saved, reusable instruction set that a person runs on demand. It waits for a human. It is repeatable, cheap, and auditable, and it scales only as far as people remember to use it.
  • An agent is a process that runs without a person starting it. It works on a trigger or a schedule, acts on systems of record, and needs supervision because nobody is watching each run.

That single distinction changes what the organisation must provide. Skills need a shared library and an owner. Agents need documented process, trustworthy data, an explicit permission list separating actions that need approval from actions that do not, and someone monitoring for drift.

It also changes the honest measure of success. A skill is working if people use it. An agent is working only if its correction rate is low. An agent whose output your team fixes four times out of ten is a supervised workflow wearing an autonomous label.

Treating the two as interchangeable is how teams end up buying agent software to solve a shared-library problem. For a grounded read on what the agent layer genuinely does today, see AI agents for sales teams.

Who should own AI standards in a revenue org?

One named person, and in most mid-market teams that person sits in RevOps or enablement. Not a committee, not a working group, and not the whole team by default, because shared ownership of a library is functionally the same as no ownership.

The role makes four decisions:

  • What becomes the standard version of a task.
  • What gets retired when it stops working.
  • Who is allowed to edit the library.
  • When it gets reviewed after a process change.

Two hours a week covers it at most mid-market scales, which is why this is the cheapest genuine progression available. It needs no budget approval and no procurement cycle.

Two cautions from practice. First, do not hand it to your best rep as an extra duty, because their incentive is their own quota, not the median. Second, tie the review cadence to an existing meeting, since standalone maintenance rituals quietly die.

You will know ownership is real when a new hire runs the standard workflow in week one without asking anyone. Oliv AI treats that same ownership question as a precondition in its implementations rather than an afterthought. For the operating detail, see scaling revenue operations in a growth-stage sales team.

What is a GTM engineer, and do we need to hire one?

A GTM engineer is the person who turns revenue processes into reliable agents and then keeps them reliable. The work involves deciding which actions run unsupervised, watching where output drifts, and updating context when the process changes. The title is new. The work is not, and much of it already sits inside modern RevOps.

You need the function before you reach the top rung, but you do not necessarily need a new hire. There are two honest routes:

  • Grow it from RevOps. Slower, and the institutional knowledge stays in house.
  • Buy it as a service. Faster, and you depend on the vendor keeping the context current.

Oliv AI publishes named in-house forward-deployed GTM engineers who implement directly with a customer's stack rather than through a third-party partner, with a first agent live on real deals by week two. That is one route, not the only one.

If hiring feels impossible right now, do not hire. Pick one repeated judgement, give one existing person two hours a week to supervise it, and watch whether the correction rate falls. If it does, you have a business case. If it does not, rung three was the right ceiling. The architecture side is covered in our RevOps guide to implementing agentic AI.

What has to be true before AI agents can run in production?

Four things, and the first two are hard ceilings rather than nice-to-haves.

  1. Documented process, as it actually runs. Agents inherit your history, not your standards. Test it by asking three managers separately when a deal goes to legal. Three answers means you have found your first project.
  2. Data you can act on. Sellers consistently name data quality and admin friction as the top blockers to agent returns. An agent that attaches a meeting to the wrong opportunity writes a confident update on a deal that does not exist.
  3. Explicit permissions. Split actions into ask-before-acting (anything a customer sees, anything that changes deal economics) and act-without-asking (reversible internal work with an audit trail).
  4. Disclosure wired in. Under EU rules effective from August 2026, an agent must identify its AI nature and the entity it acts for. Build that into the system and log the event rather than handling it in copy.

The uncomfortable part is that the first item has no software in it. Writing the playbook is a management project, and any vendor implying otherwise is selling you a shortcut that does not exist. Start with the playbook and a duplicate-record audit, both unglamorous and both cheaper than a failed rollout. See CRM data quality automation for RevOps.

Enjoyed the read? Join our founder for a quick 7-minute chat — no pitch, just a real conversation on how we’re rethinking RevOps with AI.

Video thumbnail

Revenue teams love Oliv

Here’s why:
All your deal data unified (from 30+ tools and tabs).
Insights are delivered to you directly, no digging.
AI agents automate tasks for you.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.