← Blog

Karpathy's Three-Layer AI Method: What He Actually Said

The three-layer method moves most of the work of using AI to before and after you type a request. Before, you write a detailed specification, or spec, of the real goal. After, a verifier checks the output against standards set in advance. Around both sits an environment of standing instructions, reference files and hard limits that carries over from one chat to the next. Andrej Karpathy never called it a three-layer method. A YouTube video posted on June 9, 2026 organized his April 2026 Sequoia talk that way, and most of it holds up.

Add us as a preferred source on Google

A free Google setting to see more of our relevant articles in your own search experience. You can change it any time.

Andrej Karpathy’s three-layer method is all over LinkedIn and agency blogs. It is usually sold as the right way to prompt Claude. Every version we traced cites, or closely follows, a 13-minute YouTube video that the creator Austin Marchese posted on June 9, 2026. In it, Marchese says Karpathy’s method “can be broken down into three simple layers.” The breakdown is Marchese’s.

Karpathy co-founded OpenAI, was head of AI at Tesla and coined the phrase “vibe coding” for building software by describing it to an AI. Marchese drew on Karpathy’s fireside chat with Sequoia Capital partner Stephanie Zhan at the firm’s AI Ascent 2026 event. Karpathy posted a summary and an edited transcript on his blog on April 30. Specs, checking AI output and investing in your own setup all come up in the conversation. He never groups them into layers.

Most of the method survives a check against the recording, and the parts Karpathy did not say are worth using anyway. The popular version also skips his best example, and it only hints at the idea that explains why the middle layer works.

Each layer, traced to its source

LayerWhat Karpathy said at AI AscentWhat Marchese’s video adds
SpecPeople stay in charge of the spec and the plan. Work out a very detailed spec with the AI, beyond what plan mode produces.Have the AI interview you about the goal. Keep specs small, with checkpoints. Make the AI ask you to confirm key decisions.
VerifierAI automates what can be verified, which makes it uneven. Yelling at a model changes nothing. Stay suspicious and stay in the loop.Write pass criteria first. Use a second model as a critic. Check the work against real signals, such as a live deployment.
EnvironmentInvest in your own setup. Keep an AI-maintained wiki of what you read.A standing instruction file, saved instructions for repeat jobs, limits the tool enforces, and an always do, ask first, never do split.

Three stages for working with AI: write the spec before you ask, check the answer afterward, and keep standing instructions ready for every later chat

The request sits between a spec and a check. Sources: Karpathy’s AI Ascent write-up, April 30, 2026, and Marchese’s breakdown, June 9, 2026.

The spec carries context the AI can’t see

Karpathy raised specs when Zhan asked which human skills gain value as AI does more of the work. His answer started with agents, the AI tools that carry out multi-step jobs on their own, such as writing and testing software. He said agents are like interns right now, and that “people have to be in charge of this spec, this plan.” Then he added:

“I actually don’t even like the plan mode… obviously it’s very useful, but I think there’s something more general here where you have to work with your agent to design a spec that is very detailed.”

Plan mode is a setting in Claude Code and similar coding tools where the AI drafts a plan for your approval before it changes anything. Karpathy wants more detail than a plan carries. He suggested the spec might be “basically the docs,” with the person in charge of oversight and the top-level categories while agents fill in the blanks.

His best example came from MenuGen, a small app he built that turns a photo of a restaurant menu into pictures of the dishes. Users sign in with Google and buy credits through Stripe, the payment service. His agent matched each purchase to an account by comparing the two email addresses. Customers can pay with a different email from the one they signed in with, and anyone who did would pay for credits and never get them. Karpathy had to insist that everything tie to a permanent user ID.

Karpathy knew that one customer can have two email addresses. The agent did not. A spec is where that kind of knowledge goes before an agent starts guessing.

Marchese’s steps for writing a spec are practical, and none of them come from Karpathy:

  • Ask the AI to interview you about the goal before it drafts anything. “Create the end-of-month report” is a task. The goal is the decision that report is meant to drive.
  • Keep each spec small, with a checkpoint where you review the output before the next piece starts. Handing over the whole job at once and waiting for a finished product is what Marchese calls working waterfall.
  • Read the spec the AI drafts line by line, and tell it to make you confirm each key decision.

The verifier: AI is strongest where work can be checked

This layer has the deepest roots in Karpathy’s own writing. In a November 2025 blog post called Verifiability, he argued that traditional software automates what you can specify, while AI automates what you can verify. A task counts as verifiable when a model can attempt it many times and have every attempt scored automatically. That is why AI has improved fastest at math and code, where answers can be checked, and lags on work that depends on context and common sense.

At AI Ascent he gave his newest example of that unevenness, which he calls jaggedness:

“I want to go to a car wash to wash my car, and it’s 50 meters away. Should I drive or should I walk? And state-of-the-art models today will tell you to walk because it’s so close.”

The car, of course, has to be at the car wash. “How is it possible that a state-of-the-art model can refactor a 100,000-line codebase or find zero-day vulnerabilities, yet tells me to walk to the car wash?” he asked. His answer was that people need to stay in the loop and treat these systems as tools.

The same talk explains why pleading with a model does not help. Karpathy describes these models as ghosts rather than animals, statistical simulations with none of the drives that push a person to try harder under pressure. “If you yell at them, they’re not going to work better or worse,” he said. “It’s more just being suspicious of it and figuring it out over time.” Marchese swaps the ghost for a robot librarian that answers from the books it has and confidently makes something up when the right book is missing. It is a good picture for explaining the idea to a team, and it is his, not Karpathy’s.

Marchese’s three checks turn that suspicion into a routine:

  • Write the pass criteria before the work starts. “Make this report look good” gives the model nothing to check. “Three sections, each ending with a recommendation” does.
  • Have a second AI model review the first one’s output. Karpathy suggested the same for writing, where there is no single right answer: “you can imagine having a council of LLM judges.” In plain terms, several AI models grade the same piece of work.
  • Connect the AI to a real-world signal. An agent that deploys a website can confirm the site is live, and a monthly report can be checked against last quarter’s.

The strongest outside support comes from Boris Cherny, who created Claude Code. In a January 2026 thread, his advice was to “give Claude a way to verify its work,” which he called “probably the most important thing to get great results out of Claude Code.” He added that “if Claude has that feedback loop, it will 2-3x the quality of the final result.” The figure is his estimate, not a published study. The same test applies when you choose an AI system in the first place. A small set of your own labelled cases tells you more than a benchmark score, as we explain in what an AI evaluation actually measures.

The environment carries your setup into every chat

Asked what separates a mediocre programmer using these tools from a fully AI-native one, Karpathy said it is about “getting the most out of the tools available, using their features, and investing in your own setup.” He compared it to the way programmers have always tuned their code editors. Marchese turns that into a workshop you build once and keep improving, so each new chat starts with your rules and reference material already in place.

The part of this layer Karpathy has written about in detail is his personal knowledge base. “When I read an article, I have my wiki being built up from those articles,” he said at AI Ascent. “I love asking questions about it.” He published the idea as a note on April 4, 2026, written to be pasted straight into an AI tool. In it, an AI reads the documents you collect and builds linked summary pages from them. He calls the pattern an LLM Wiki. A large language model (LLM) is the technology behind ChatGPT and Claude.

The pattern has three layers of its own. Raw sources go in untouched, and the AI reads them but never edits them. From those, the AI writes and maintains the wiki, a set of linked pages. The third layer, the schema, is an instruction file that tells the AI how the wiki is organized. Of the wiki, Karpathy writes, “You read it; the LLM writes it.” If you have seen “Karpathy’s three layers” described as sources, wiki and schema, the writer was describing this note.

That schema is the same kind of file that anchors Marchese’s version of this layer, which has four parts:

  • A standing instruction file. Claude Code reads a file called CLAUDE.md at the start of every chat, and OpenAI’s coding tool, Codex, reads one called AGENTS.md. Put the defaults you would otherwise repeat in it, such as “Before building anything multi-step, include a verification plan.”
  • A reference folder organized so the AI can find what it needs, his version of Karpathy’s knowledge base.
  • Saved instructions for recurring jobs, which Claude calls skills. His rule of thumb is to write one for anything you plan to do repeatedly and to fix it each time it falls short.
  • Hard limits for anything expensive to get wrong.

Limits need more than a sentence in that file. A line telling the AI to leave a folder alone is a request, and the model can ignore it. Claude Code also supports hooks, small checks the tool runs before the AI edits a file, and a hook can block the edit outright. Marchese sorts work into three bins: what the AI always does on its own, what it must ask about first and what it never does. The ask-first bin is the same idea as the human-in-the-loop designs in our breakdown of what AI agents cost, where the agent does the busywork and a person approves anything irreversible.

Understanding stays with you

Most write-ups end on “You can outsource your thinking, but you can’t outsource your understanding,” and so did Karpathy’s talk. He introduced the line as a tweet that “blew my mind,” so the words are someone else’s. He explained why they stuck with him: “I am becoming the bottleneck of even knowing what we are trying to build, why it is worth doing, and how to direct my agents.”

Each layer depends on that understanding. You cannot write the spec without knowing the real goal, judge the output without knowing what good looks like, or set limits without knowing which mistakes your business cannot absorb.

Karpathy’s blog write-up of the talk shows the method on a small job. He gave an AI model his recent blog posts and tweets, then had it turn the transcript into a summary and a cleaned-up version. He read both before posting and reported that the result “reads ok without glaring mistakes.” His own writing was the environment, the two outputs he asked for were the spec, and his read-through was the verifier.

To try it this week, pick one recurring task you already hand to an AI tool, such as a weekly sales summary or a client report. Before the next run, write down the decision the output feeds and the checks a good version must pass. Save both in the tool’s standing instructions, then compare the next few outputs with the last few.

Sources

Talk to us

Talk to the team that would run it

Tell us where to look and we reply within 24 hours with where we would start. No pitch until you see the value.

Teegan will follow up by email about your request. Unsubscribe anytime. Privacy

Prefer to talk first? Book a 30-minute call instead.

Add us as a preferred source on Google

A free Google setting to see more of our relevant articles in your own search experience. You can change it any time.

Want AI doing this for your growth?

We build AI-driven acquisition, content, and automation systems for operators across North America. See your levers in 30 minutes.

Book a growth consult → Rated 5.0 on Clutch

Explore AI automation services →