Codos
All posts

The AI transformation playbook: four pillars of AI-native companies

Ninety-five percent of AI programs never reach the P&L. The companies that compound share one kind of CEO and four pillars. From inside dozens of engagements, here is what they do differently.

Thirty people from every function of a company, sitting in a loft, watching a demo on a large screenDima Khanarin talking to visitors at the Codos stand at the a16z AI Faire
Left: ten days in a hacker house, the top 30 people of one client rebuilding their own functions. Right: the Codos stand at the a16z AI Faire, San Francisco, August 2026.

The most common question I heard at the recent AI Faire: how to apply AI successfully?

Most companies look at AI the way they looked at every technology before it: a lever to cut some cost and squeeze out some margin. But surprisingly, most execs still don't see P&L impact! My personal chats with 100+ people and MIT study confirm that.

The line from that same report that nobody quotes is the important one: the core barrier is not infrastructure, regulation or talent. It is learning. Most AI systems in companies do not retain feedback, adapt to context, or improve over time.

The companies that get results look at it completely differently. Our best clients treat AI transformation as existential. Not innovation theater, not a line in the IT budget. A survival question. When leaders operate with that kind of urgency, they get impact in months, not years.

But urgency alone is not the playbook. The best leaders I work with also understand the endgame.

The endgame: a self-improving organization

The end state of AI transformation is not AI-assisted employees. It is an organization that improves itself. Agents that watch how work happens, propose better ways, test them, and ship the improvements, while humans set the goals and review the results.

The same week GPT-6 shipped, I got to ask Sam Altman a question at a Speedrun Q&A, and what stayed with me was not the answer but how matter-of-fact the exponential is to the people building it. Just as the frontier labs are racing toward the autonomous AI researcher, companies should be racing toward the self-improving organization.

1. Models improve exponentially. The length of a task AI can finish on its own keeps doubling, roughly every four months by METR's data. In 2023 models could handle tasks that take a person a few minutes. Today they complete tasks that take a full workday.

Chart: the length of tasks AI can complete autonomously, by model, doubling roughly every four months
The length of tasks AI can complete autonomously keeps doubling. Source: METR.

OpenAI's internal numbers show the same curve from the inside. Month by month through 2026, its coding agents complete a larger share of researcher tasks with zero human intervention, and the tasks they succeed at keep getting longer: by July, tasks that would take a person a full day or two are succeeding more than half the time.

Chart: success rate of OpenAI's internal agents with zero interventions, by estimated human task time, one line per month from January to July 2026; later months succeed more often on longer tasks
Agents are solving longer tasks for researchers, month by month. Source: OpenAI, September 2026.

2. The labs already live in this future. In June, Anthropic reported that Claude wrote more than 80% of the code merged into its own production codebase in May 2026, and that a typical engineer ships about 8x more code than in 2024. Anthropic itself calls the 8x number overstated, because lines of code is a vanity metric. Even a heavily discounted 8x is not a productivity gain. It is a different way of working. OpenAI's chart of the same metric looks identical: lines changed per active engineer flat at 1x for four years, then 7x by mid-2026.

Bar chart: lines changed per active OpenAI contributor, indexed to the pre-2025 average, flat near 1x from 2021 to 2024 and rising to about 7x in 2026
Across OpenAI, engineers ship code faster: lines changed per active contributor, pre-2025 average = 1x. Source: OpenAI, September 2026.

3. Research is next. On September 6, OpenAI published a view inside its own research organization. It had hit its goal of an automated research intern: a system that carries out well-defined research tasks that would take a skilled researcher a few days. The organization now gets 3.1 agent-workdays for every human workday. The median researcher spends over $600 a day on agent inference. The top 10% spend over $7,000 a day. The next stated goal is an automated AI researcher by March 2028.

Line chart: daily agent spend for a 90th-percentile OpenAI researcher, near zero in January 2026 and above $7,000 a day by August
Daily agent spend of a 90th-percentile OpenAI researcher, January to August 2026. Source: OpenAI.

4. Agentic developers are already 10-100x on output. A strong human developer does around 5,000 GitHub contributions a year. Developers running fleets of agents post 170,000+. You can argue about what a "contribution" measures. You cannot argue with the shape of the curve.

Comparison: a strong human developer at 4,875 GitHub contributions a year versus an agentic developer at 173,809
A strong human developer versus a developer running fleets of agents, contributions per year.

5. The labs say it out loud. Anthropic, in its own words: recursive self-improvement, AI autonomously improving AI, "could come sooner than most institutions are prepared for." Sergey Brin is reportedly pushing Google's resources specifically toward recursive self-improvement, "the point at which the technology can improve without human intervention."

Anthropic's own economists have now put numbers on this. In their September scenarios, the "extreme" case is defined by exactly two conditions, recursively self-improving AI plus rapid adoption, and it takes US GDP 32% higher by 2030. The "substantial" case, where AI handles half of knowledge work mostly autonomously, adds 8%. The models are the same in both. The difference is how fast companies adopt them.

If AI systems are learning to improve AI systems, the organizational analogue follows directly. Companies whose operations improve themselves will compound away from companies whose operations only improve when a human notices a problem, schedules a meeting and writes a ticket.

That is the destination. The rest of this playbook is about getting there. After working closely with dozens of companies, and having intelligence into hundreds more, the difference between the ones that compound and the ones that stall comes down to one person and four pillars.

Pillar zero: an AGI-pilled CEO

Before the pillars, a portrait. You will recognize whether this is you.

The CEO of a 500-person payments company came to us last winter. He was not "curious about AI." He had personally built 42 Claude skills, a CEO-level assistant, a Telegram bot that pulled the company's pulse every morning, and assistants for four of his top managers. And he still could not get his company to move. Five hundred people, an eight-figure annual loss, a founder who worked AI-natively and an organization that did not.

He pushed back on our price. Compared us to McKinsey. Then came back a week later, because the DIY path was not working at the scale of 500 people, and he knew it.

What happened next is the whole playbook in miniature. He set one goal for the company for 2026 that could not be reached without AI. He put himself in the room for the transformation cadence every two weeks. He let a four-person team rebuild the core product with agents while the legacy organization kept running. And he approved a headcount target that most CEOs would not say out loud.

Every company I have seen compound has a leader like this at the top. Not a technical one, necessarily. A convinced one. Someone who has personally felt what the models can do and refuses to run a company as if they cannot. If your CEO has not had that experience yet, that is the first project. Everything below assumes it.

Pillar 1: Set unachievable goals

Goals that cannot be hit without AI. Set authoritatively, from the top.

Time "freed up" does not appear on the P&L, no budget shrinks, nobody's number changes, and a year from now the company will have a few hundred chat licenses and the same headcount. Goals like this are the default because they threaten no one.

Here is what the annual goals of a mature 1,000-person organization usually look like, next to what AI-native goals look like:

The usual annual goals
  • Grow revenue by 25%
  • Improve margin by 3 points
AI-native goals
  • 2-3x revenue per employee
  • 5x faster feature lead time
  • Cut 50% of a cost base by end of year

Every one of the goals on the left can be achieved the old way. More headcount, more budget, more grind. So that is exactly how the organization will achieve them. AI stays optional, and optional means ignored.

Nobody hits those numbers by working harder. The only path runs through AI, which means the goal itself forces the transformation. The payments company above set a single goal for all of 2026: 2x revenue per employee. One number the whole company pulls toward. Every function then cascaded it into their own OKRs. Sales cut the cost of onboarding a partner from $92 to $20. Engineering took feature lead time from 110 days to 21.

These goals are not extreme. Anthropic's middle scenario, the one closest to what executives themselves expect, has AI handling half of all knowledge work by 2030, mostly without a human in the loop. Two times revenue per employee is that scenario applied to one company.

Two cards: Sales, cost to onboard a partner $92 to $20, 4.5x cheaper; Engineering, feature lead time 110 to 21 days, 5x faster
One company-wide goal cascades into functional OKRs.

Two details matter here, and both are counterintuitive.

First, AI transformation goals are business goals. Not "80% of employees active on ChatGPT weekly." Adoption metrics are vanity metrics. I have seen a company with 70% activation and almost no real usage. The goal has to be a number the CFO already cares about. The only constraint is that it is unreachable without AI.

Second, the best leaders set these goals authoritatively. Not through months of consensus-building workshops. The CEO declares the number, explains why survival depends on it, and lets every team figure out their own path. Consensus produces goals everyone can live with. Which are precisely the goals that force nothing.

One more pattern from this summer. A company I work with looked at its half-year numbers, saw the plan was 40% off, and did something I rarely see: it cancelled its OKR process mid-year. It replaced it with two working groups, one owning efficiency with a hard operating-expense ceiling and one owning growth with the revenue gap as its only target, with daily standups until the portfolio of projects was set. That is what a goal reset looks like when leadership means it.

The most common failure mode

Declare an AI transformation, but change no goals and no incentives.

I see this constantly. The town hall happens, the "AI-first" memo goes out, licenses get purchased. And then people do tiny bits, fake AI usage in demos, or quietly ignore the whole thing. Because everything they are actually measured on still rewards the old way of working. If your goals can survive without AI, so will your org chart, your processes and your habits.

No goals, no transformation. This is Pillar 1 because everything else depends on it.

Pillar 2: It's about people, not technology

Vibecoding is easy. Changing behavior is hard.

Here is the uncomfortable truth about the technical side of AI transformation in 2026: it is the easy part. The tools are mature, the models are extraordinary, a capable team can stand up working agents in weeks. What is hard is getting a 45-year-old department head, successful for 15 years doing things their way, to change how they work every single day.

The economists agree. In Anthropic's scenario model the constraints that slow everything down are not compute or model quality. They are people: workers who do not want to change how they work, skills that take time to build, and how hard it is to move into a new role.

That is a behavior change problem, and behavior change has a known playbook. We use the four levers of the McKinsey Influence Model: people change behavior when they see leaders model it, when they believe the why, when they have the skills, and when consequences reinforce it. Miss a lever and the transformation leaks.

Here is what each lever looks like in practice, from real engagements.

Role modeling: ten days in a hacker house

One of our clients took the top 30 people in the organization, every function represented, co-founders and CEO included, and put them in a hacker house for ten days. The mandate: rebuild your function's core processes AI-natively by demo day. Every day ended with an evening demo. The engineering team enabled the non-technical vibecoders and built the shared infrastructure. Operations kept running the whole time, at lower intensity, while the company's leadership publicly rebuilt their own jobs.

Slide: ten days in a hacker house. The top 30 people in the org, AI-native processes across every function by demo day.
Every function represented, co-founders and CEO leading by example, an evening demo every day.

You cannot delegate this. When the CEO ships their own agent and demos it to the company, the message lands in a way no memo ever will: this is how we work now, and I am doing it too.

Formal mechanisms: an AI office with full transparency

Same discipline you would apply to any P&L-critical program. A bi-weekly C-level cadence: CEO, CPO, CTO, COO, plus the AI team and function champions. AI pulls the data from every source, metrics, P&L, each initiative's status, and surfaces the cost centers and functions nobody has touched with AI yet. Leadership sees every initiative in one waterfall and reviews it together.

Slide: an AI office for full transparency, with a waterfall of initiatives and their verified impact
An AI office: one transparent view of every initiative in front of the leadership team.

At the payments company this cadence produced several $M in realized savings in five months, ahead of a stretch goal that had been set for twelve. What gets measured in front of the CEO gets done. What lives in a working group's Notion page dies there. I have watched both happen at the same company.

Developing skills: mandatory training and a new performance bar

Hope is not a training strategy. One client made AI-nativeness an explicit criterion in performance reviews and put everyone through mandatory weekly two-hour training, with levels assigned by a five-step ladder: (1) chat assistants, (2) CLI tools and vibecoding, (3) a personal AI OS, (4) agents, (5) orchestrating fleets of agents. Real adoption went from 15% to 60% in two months.

The five-level AI-native ladder: chat assistants, CLI and vibecoding, personal OS, agents, orchestration; adoption 15% to 60% in two months
The five-level ladder behind mandatory weekly training.

And across the market the bar for staying employed is visibly rising. At one 1,000-person company there is now a mandate called "My First Pull Request": every PM, designer and non-engineer must ship code using AI tools. At the same company, a quarter of product meetings open with working prototypes instead of slide decks. Another large company is asking every employee to reinterview for their own role, by building an app that makes them better at their job.

Slide: the bar for staying is rising, with examples of companies raising the AI-native requirement for every role
The bar for staying employed is rising across the market.

Reinforcing: hire only AI-native people

The fastest way to change a company's average behavior is to change who you hire. We give candidates a task that is impossible to finish in the time allotted without AI, then watch how they use it. Not "do you know the tools" but "is this how you naturally work." Every AI-native hire raises the bar, every legacy hire lowers it.

One rule of thumb captures the whole mindset shift, and it is honestly the single best diagnostic I know for whether a company is serious:

If your API bill doesn't make you uncomfortable, you're not spending enough.

Remember the OpenAI number: the top 10% of their researchers spend $7,000 a day on inference. Companies still treating tokens as a cost to minimize have not understood what they are buying. Tokens are labor. The cheapest labor in history, on a price curve falling by orders of magnitude every year for equivalent capability. Spend accordingly. In Anthropic's extreme scenario, labor's share of output falls from about 60% to 45% by 2030 even as wages rise, and the gains go to whoever owns the AI capital. Inside a company, the token bill is that capital line.

Pillar 3: Context is the moat. Aim for 100% token coverage

It's either 100% token coverage or nothing.

He had discovered the thing most AI transformations silently die of, and I want to put a name for it into the water supply.

Token coverage
The share of your organization's working context that AI systems can actually see and reason over. Conversations, documents, decisions, metrics, code, meetings, customer interactions. A company at 30% token coverage has AI that guesses. A company approaching 100% has AI that knows.

Why does this matter so much? Because an agent's usefulness collapses non-linearly with missing context. Take the simplest possible workflow: a sales call transcript comes in, an agent should update your CRM. Easy, right? But the transcript mentions "Dima K." Should the agent create a new profile? Update the existing Dima K? Or is this a different Dima entirely, and it is about to corrupt the profile of your biggest account?

Diagram: one uploaded transcript, and the agent must decide whether to create a new profile, update an existing one, or wrongly update another
One transcript, one simple task, and the agent needs the whole organization's context to get it right.

A human assistant answers that instantly because they hold your company's context in their head. An agent without that context is confidently wrong. And after the third wrong CRM update your team stops trusting all of it. Partial context does not produce partial value. It produces distrust, and distrust produces zero value. That is why it is 100% or nothing.

The problem: your organization's context is scattered across 10+ platforms. Slack, email, calendars, Notion, CRM, call recordings, GitHub, spreadsheets. One person's work alone spans a dozen tools. No model, however smart, can reason over what it cannot see.

This is why the serious players are converging on the same answer: an org brain. A living, continuously updated memory of the organization that every agent reads from and writes to. Not a data lake (storage without meaning). Not RAG over a document dump (retrieval without structure). Not enterprise search, which is what that AI lead had, and which answers "where is the document" without answering "what should we do." An ontology: people, projects, decisions, processes and relationships, extracted from raw activity and kept current automatically.

Pipeline diagram: raw data flows through extraction, observer agents, a merge judge and memory updaters into a living company vault
From raw data to living memory: the pipeline of agents that keeps an org brain current.

We are not the only ones saying this. Jack Dorsey's Block is one of the most watched examples, profiled by Sequoia under the title "From Hierarchy to Intelligence": companies aiming for exactly this, all tokens stored in the org brain, AI as the fabric of how the organization coordinates rather than a productivity add-on.

Slide: leading tech companies aiming for 100% token coverage, with all tokens stored in the org brain
Leading tech companies are aiming for 100% token coverage.

A note for the COO who told me last week that a previous vendor had leaked her firm's data and that nothing could leave her servers again. She is right, and it does not change the pillar. The brain is your company's memory. It should live where your security team can see it, on infrastructure you control. What you buy is the machinery that builds and maintains it, not a place to send your context.

What an agent with context actually does

At the payments company, merchant onboarding used to take a team of twenty people collecting the same documents twice. With the brain in place, an agent now drafts the onboarding from the application itself. Eighty-two percent of merchants complete it self-serve, 42 fields are filled automatically from the documents they upload, and the median application takes 23 minutes. Seventeen of the twenty roles were returned to the business from that one workflow. That is not a chatbot. That is what context looks like on the P&L.

Context is what makes self-improvement possible

Here is the connection back to the endgame, and it is the real reason this pillar exists: you cannot have a self-improving organization without near-total token coverage. Self-improvement requires a loop. Observe how work happens, form a hypothesis, test it, measure, keep what works. Every step of that loop runs on context. An agent cannot improve a process it cannot see.

Once the context layer exists you can run improvement loops with evals, where every output gets scored against a rubric and iterated until it passes. We run this on our own systems: an autonomous research loop that ran 83 experiments on our org brain and shipped the 15 that measurably improved it, spending $1-1.5k a day on tokens (now regularly spending $10k/day). See the rule of thumb above.

Slide: autoresearch improves the brain on its own; 83 experiments, 15 kept improvements, running at $1–1.5k a day
A self-improving loop in miniature: 83 experiments, 15 kept improvements.

That is the pattern in miniature. The organization improving itself, humans reviewing the scoreboard instead of doing the iterations.

Pillar 4: Disrupt yourself

Launch a small AI-native team to disrupt your own legacy business. And accept what that means for the old one.

The first three pillars transform the company you have. The fourth one admits a hard truth: sometimes the company you have is the constraint.

Legacy organizations carry legacy physics. Five management layers, coordination meetings about coordination meetings, processes designed for a world where information moved at human speed. You can improve that system with AI, and you should. But you will hit a ceiling, because the org structure was designed around human limitations that no longer exist.

So the boldest move our clients make: spin up a small AI-native team and point it at your own business.

A global fintech we work with did exactly this with one regional unit. Before: 85 people, five management layers (global CEO, regional CEO, execs, team leads, ICs) across six functions. Endless meetings, strategy offsites, "alignment." The unit was unprofitable and growing slowly. After: 6 people in three layers. A CTO/CPO plus two engineers and two product people. Everyone ships, everyone holds full context. The new unit tests hypotheses in days instead of quarters, and it found a growing customer segment the 85-person version had missed for years.

Then the same company went further. A four-person team rebuilt the entire core product, roughly half a million lines of code, in three months. In that time they merged 848 pull requests at a median of about an hour each and shipped 646 releases, around fourteen a day. About 30% of those pull requests ran end to end with no human involved: a signal from monitoring or a ticket, a plan, code, review, deploy. Agents reviewed every change and blocked a quarter of them. Not one line of product code was typed by a person. By the team's own estimate, that is about two years of a traditional engineering organization's work.

That is a 14x reduction in headcount with better output, and a 10x reduction in time to ship. Not by asking 85 people to work harder, but by rebuilding the unit around what small AI-native teams can now do.

The structural patterns of these teams are becoming visible across the industry:

Two-person engineering teams: a pirate and an architect. The pirate moves as fast as possible, shipping product by vibecoding. The architect turns what the pirate discovers into a reliable, structured machine. Two people covering what used to take a squad.

Slide: a new engineering team, one pirate who ships fast and one architect who makes it reliable
A new engineering team: one pirate, one architect.

No middle management. Y Combinator now explicitly promotes this model: everyone is a builder-operator. Engineering, ops, support, sales all build, and everyone comes to meetings with prototypes, not decks. Every outcome has exactly one DRI. No layer whose job is aggregating status.

Slide: no middle management, everyone is a builder-operator, quoting Y Combinator
No middle management: every outcome has exactly one owner, and everyone builds.

The part nobody puts in the deck

Disrupting yourself means the legacy organization gets smaller. The payments company's target for the end of this year is to cut more than half of its headquarters headcount while growing revenue in its newest market. The companies that compound are not afraid of that sentence. They say it out loud, they do it once and decisively rather than in quarterly slices that demoralize everyone, and they hire AI-native people into the smaller company that remains. The companies that stall keep 85 people in the unit and ask them to "use AI more."

Why does the disruption have to come from a separate team? Because inside the legacy structure every AI-native experiment gets negotiated down by the immune system. The layers, the incentives, the "that's not how we do things here." A separate small team playing by new rules gives you a clean read on what is actually possible. Then you have options: grow the new unit, migrate the old one toward it, or let the comparison speak for itself. What you cannot do is unsee the result.

Yes, this is threatening. It is meant to be. The alternative is waiting for a startup with six people and 100% token coverage to run the same experiment on your market, from the outside.

Which pillar is your bottleneck

The pillars are a sequence, not a menu, and most companies are stuck on exactly one. You can usually tell which from what leadership says.

If you hear yourself sayingYour missing pillar
"Our AI goal is to free up 30% of everyone\'s time"Goals. Nothing cascades from one unavoidable number.
"The pilots worked, nobody uses them"People. No role modeling, no cadence, no consequences.
"Everyone is activated, the assistant can't answer anything useful"Context. Token coverage is too low for agents to be trusted.
"We can't grow without headcount growing with it"Disruption. The structure, not the tools, is the ceiling.
"This is interesting, let's revisit next quarter"Pillar zero. Nobody at the top is convinced yet.

Where to start

  1. Goals first. Set the number that cannot be hit without AI. Without it everything downstream is theater. If you want help finding it, that is what our diagnostic does.
  2. People second. Leaders role-model, an AI office creates accountability, training and hiring raise the floor. Behavior change is the bottleneck, resource it like one.
  3. Context third. Start immediately, expect it to compound. The org brain is the longest lead-time asset and the deepest moat. Every month of captured context makes every future agent smarter.
  4. Disrupt yourself once the first three are moving. The small AI-native team is your window into the endgame, and your insurance against someone else building it first.

And keep the endgame in view. The point was never "employees who use AI." The point is an organization that improves itself, where agents watch how work happens, propose better ways, prove them with evals and ship. The companies that get there first will not be incrementally better than their competitors. They will be compounding while everyone else is holding meetings.

I'm Dima Khanarin, CEO of Codos. We work hands-on with a small number of companies on exactly this: setting AI-native goals with leadership, building the org brain, standing up self-improving loops. If this resonates and you want to talk about your organization, get in touch. I write from the trenches, not from the sidelines.

Questions I get after this talk

A playbook is the sequence of decisions that turns AI spend into P&L impact, as opposed to a roadmap of tools to deploy. This one has four pillars: goals that cannot be hit without AI, behavior change run like a P&L program, near-total context coverage for agents, and a small AI-native team that disrupts the legacy business. The order matters. Skipping goals is why most roadmaps fail.

First measurable business impact: 60-90 days, if leadership sets real goals from day one. The payments company in the article realized first $M in four months. Meaningful token coverage and working improvement loops: 6-12 months. Companies that skip the goals step spend those same months on pilots that go nowhere.

The ontology itself must be yours. It is literally your company's memory and it should live where your security team can see it, on infrastructure you control. The machinery for building and maintaining it (ingestion, extraction, agents, evals) is what you can buy or adopt from open source. Be careful with anything that ships your entire company's context to a third party's cloud by default. Security teams are right to block that.

Three things nobody else can. Set the unachievable goal and attach real incentives to it. Personally use the tools and demo their own AI-native workflow to the company. And chair the AI office cadence. A CEO who delegates all three has delegated the transformation's failure. (This is also the job a fractional Chief AI Officer takes on when there is nobody to hire yet.)

More than feels comfortable. That is the point of the rule of thumb. Tokens are labor at a price falling by orders of magnitude every year for equivalent capability, and the frontier labs' own researchers now run on hundreds to thousands of dollars of inference a day each. A useful sanity check: if your AI spend is under 1% of the payroll it is supposed to leverage, you are experimenting, not transforming.

The pillars are function-agnostic. We have applied them in payments, cloud infrastructure and services businesses, and this year's conversations include multi-location retail, e-commerce rollups and trading firms. The mix shifts: heavily regulated companies weight security and on-prem context infrastructure higher, and the disrupt-yourself unit usually starts in a market-facing function rather than core operations. The endgame is the same.

Have an AI impact goal?

Let's move months ahead of schedule.

Book a demo