Skip to main content

Two Ladders, One Climb: Reading Boris Cherny's AI Adoption Steps Alongside Ours

Post by John Espey
Two Ladders, One Climb: Reading Boris Cherny's AI Adoption Steps Alongside Ours

I came across Boris Cherny's Steps of AI Adoption last week and grinned. Cherny runs Claude Code at Anthropic, and his five-step ladder for how much work you can safely hand an AI agent is the cleanest short version I have seen of the capability side of this whole journey.

We have been drawing a five-level ladder of our own for clients since last spring, and reading his felt like finding the other half of a map we had been sketching from one side. His tracks what the agent can do. Ours tracks what the organization is ready to operate. Together they cover a customer from both directions at once, and that is exactly how we have started using them. 

The Ladder Boris Drew

Cherny's steps track agent capability: how much planning and execution you can safely hand the agent, and what tooling and guardrails each step needs. At the bottom sits Manual and Isolated, one person and one AI tool with no shared context. Assisted is the Copilot stage, where AI drafts and a person reviews. Then an agent starts planning and executing multi-step work under supervision, what he calls the Agentic Loop, before that context gets shared across tools and systems as an Integrated Ecosystem. At the top is Autonomous Operations, standing work that runs without a human touching every step.

His throughline gives the ladder a spine: increasing agency, decreasing cognitive load. Each rung comes with the Claude tooling behind it and the guardrails that keep it from running away from you. It is a genuinely useful map of where agent capability is headed, drawn by someone who spends his days at that frontier. Reading it clarified our own thinking, which is the highest thing I can say about anyone else's framework.

The Ladder We Draw for Clients

Cherny's model got me looking at ours from a fresh angle. Where his ladder charts the agent, ours charts the organization around it: how far a company can absorb, operate, and eventually compound the capability the agent already has. Five levels, and I will admit we borrowed the grid format he uses, because it is simply the right way to lay a ladder out. His work made ours sharper before I had written a sentence of this post.

Level

Definition

What it looks like

Reference-arch components

Guardrails

L1 Experimental

Individuals discover AI on their own

Personal ChatGPT or Copilot accounts, no oversight

None yet

The risk is unmanaged data exposure

L2 Adopted

Sanctioned tools, basic policy

Enterprise licenses, an acceptable-use policy

Vendor SaaS, no custom code

Acceptable-use sign-off, procurement review

L3 Engineered

AI workloads built like production software

Named workflows, golden-set evals, code review on prompts

Model APIs, vector store, eval suite, tracing

CI gates on eval regression

L4 Operated

The feedback loop runs continuously

Multi-vendor routing, replay of any past decision

Tiered model fallback, policy gates in the deploy path

Enforced PII and risk gates, live in the path

L5 Compounding

Years of tightened signal turn into leverage

Eval sets built from years of production edge cases

Same architecture as L4, aged

Discipline against complacency and vendor lock-in

 

We score this per capability. As an example, a bank can run L4 fraud detection and L1 HR in the same building. The reference architecture stays vendor-neutral. Those components are categories, a vector store, an eval suite, a routing layer, assembled from whatever fits the client rather than any one product line. We walk through the full model, with the market calibration and engagement sizing behind it, when we scope a client engagement. 

Why You Read Both at Once

Here is the simple version: for any one capability, picture two dials. The first is how far the agent can go, and that comes straight off Cherny's ladder. The second is how far your organization can actually run it day to day, and that comes off ours. When a program is healthy, the two dials climb together. When the agent dial races out ahead of what the organization can operate, you have a gap, and that gap is where the risk lives. The two ladders were clearly drawn by people watching the same thing from different seats, so they rhyme more often than not, but you do not need to line them up rung by rung to use them. You just read both and watch the space between.

His guardrails and our reference architecture are the same concern seen from two seats. He flags the guardrails each step needs, which is exactly the right instinct. We spend our days on the plumbing that makes those guardrails real at enterprise scale: the routing, the eval gates, the audit trail. The work we get hired for is closing whichever gap is holding a company back.

That is where most of the value hides. I once watched a team push an agent well past what the operations group around it could actually run, while that group was still hand-checking every output. The demo was gorgeous. The program was shelved by the next quarter, not because the agent fell short but because the organization had not climbed with it. Reading both dials at once is how you catch that early, from either side.

What Our Ladder Adds: Compounding

Our top level, Compounding, lives on the organizational side, which is precisely why it pairs so well with Cherny's capability view rather than crowding it. It is a data flywheel. Years of eval signal, tightened against your specific operation, turn into judgment no competitor can buy off a price list. A company can stand up a fully autonomous agent tomorrow and still not have this, because it takes years of production traffic, corrected errors, and replayed decisions to build. Autonomy sets the pace. The flywheel is what keeps you ahead once everyone has the same tools.

Held side by side, the two ladders answer different halves of the same question. His tells you how far the agent can go. Ours tells you how ready the organization is to go there with it, and what starts compounding once it does. Two years from now the tooling on his ladder will look different again, probably several rungs further along, and that is a good thing. The companies that come out ahead will be the ones who spent those two years building the eval discipline and the replay habits that make the flywheel spin, the part of the advantage no one shortcuts by copying an architecture diagram or hiring the same vendor.

So look at both when you plan your next move. How far can the agent go on this capability, and how far can your organization actually operate it. Reading them together is what turns a promising pilot into something that lasts.

If you want to walk through where your organization sits on both ladders, grab time with me: https://meetings.hubspot.com/john-espey



 

Post by John Espey