August 14, 2026 · 13 min read
The seven levels of AI for CPG: how to move up one, whichever level you are on
What each rung looks like when it is working, what you actually get for climbing it, and what it costs you if you climb it in the wrong place.
Almost every piece of AI advice aimed at CPG founders amounts to "start using AI." That is not an instruction, and it is why so many brands have a subscription, a vague sense of falling behind, and nothing they could point at.
There are seven levels of AI for a CPG company, and you are not on one of them. Score your company as a whole and almost every operation comes out between two and three, but that average covers areas that are nothing like each other. Your customer service replies might be at three while your reporting is still at one. Ad creative might be at four because somebody wired up a scheduler last year and it still runs. In one narrow area, usually the one you personally got interested in, you might already be running at six.
So each part of your operation sits at its own level, and the tools being sold to you are pitched as though the whole company moves up together. It moves one area at a time.
Take five areas, reporting, customer service, creative, inventory and wholesale, and give each its own number. The distance between your deepest area and your shallowest one is what a single company-wide score hides, and that distance is where your next move is.
The seven levels of AI for CPG are chat, research, projects, automation, dashboards, agents and intelligence, running from asking a question in a chat window at level 1 to a system that applies your standards to daily decisions at level 7. Find the area you scored lowest and read its section.
Level 1: chat
You ask a question in a chat window and get an answer back. What separates a useless answer from a good one is entirely what you put in. Somebody types "write me a reply to an unhappy customer", gets something bland, and decides the tool does not work. Paste in the actual email, your actual returns policy and two replies you were happy with, and the third draft is usually sendable.
You move up by handing it a source instead of asking it to answer from memory. Export your last two hundred reviews and ask for the three most common complaints, with quotes.
That single change gets you answers about your own business rather than answers about businesses in general, and it improves the hit rate more than anything else on this list. It is also the cheapest place in the company to be wrong, which is why anyone nervous about all of this should start here. The worst outcome is a wasted afternoon.
Level 2: research
You point it at a question and it goes and does the digging. Ask the things you would have paid an analyst for and never got round to. Which claims appear on the top ten products in your category? What do the one-star reviews of your closest competitor actually say? What changed on their PDP since spring?
Ask for a source beside every finding, because confidently wrong research arrives in clean paragraphs and reads exactly like the right kind. You only need to check the two findings you are about to make a decision on.
You move up by giving the research somewhere to live. Research that sits in a chat window you closed is research you will pay for twice, and the fix is a workspace that keeps its own files, its own standing instructions and its own memory.
Level 3: projects
At levels 1 and 2 the unit of work is a conversation, and every conversation starts from nothing. Level 3 changes the unit to a workspace that remembers. Claude's Cowork projects are the clearest example: each project holds its own files, its own standing instructions and its own memory, kept separate from every other project.
You stop re-explaining your own company. A project for your Amazon listings holds your brand guidelines, your claims rules, the twenty listings you already approved and the competitors you care about. You open it, ask for the next listing, and all of that is already in the room.
Set one up properly and it pays out for a year. Write the standing instructions the way you would brief a new hire: what the brand sounds like, what it never says, which claims legal will not allow, who the customer is. Add your best previous work as files rather than describing it, because three approved listings teach it more than a paragraph of adjectives about your tone. Get that wrong and you have something worse than no project, because it will apply your bad instructions confidently for months.
Level 3 still depends on you for everything. It has memory and no initiative, and nothing in that project happens until you sit down and start it, which keeps it on the same side of the ladder as chat. You move up by finding the job you repeat on a rhythm, every week or every launch or every new SKU, and writing down its steps exactly as you perform them, including where the inputs come from and who reads the output before it goes out. That written-down version is the automation, and writing it down is most of the work.
Level 4: automation
The easiest way in is recurring tasks in Cowork, which is the project you already built at level 3 with a schedule attached. Same files, same standing instructions, same memory, except now it runs at seven on Monday morning without you opening it. They also fire on events rather than clocks, so a file landing in a folder can be the trigger, which covers most of the "when the 3PL export arrives" work in a CPG operation.
This rung gets skipped more than any other, and for the wrong reason. Founders assume automation means engineering and a budget, so they sit at level 3 for a year. At this rung it means putting a schedule on something you already have working.
Run one job on a schedule and read its output every time. Twenty half-finished automations that each work most of the time leave you with twenty things to check and no idea which one broke, so one job you trust beats twenty you have to audit. The work then happens whether or not anyone remembered it, and fewer things live in one person's head.
An automation stops running, nothing announces it, and you find out weeks later from the gap it left. Anything you automate needs an answer to how you would know if it stopped.
You move up when two automations need the same number and each goes and gets it from a different place. That moment always arrives before you feel ready for it, and when a figure has two homes the next build is not another automation.
Level 5: dashboards
Level 5 is where information that lives in separate places gets presented as one thing. Sales sit in one system, spend in another, stock in a third, and somebody's spreadsheet holds the part that never fit anywhere. Put them together and you see the business more clearly than you ever have, because the picture is no longer limited to whichever system you happen to have open.
Almost every system you log into can hand its numbers to another system automatically, on a schedule, instead of somebody exporting them by hand. Those numbers land in one database, and something reads that database and draws the picture. My demo dashboard is a Next.js app on Vercel reading from Postgres, a deliberately unremarkable stack, because the hard part of level 5 was never the technology.
There is now an agreed way for an AI assistant to read a company's own database directly, called the Model Context Protocol, and most of the tools you would use already support it. Once your numbers live in one place you can ask a question in plain English rather than reading a chart, which is a large part of why level 5 is the rung everything above it depends on.
A good dashboard labels every figure with where it came from and when it last updated. Most skip that, and skipping it is what makes them ornamental: if you cannot tell at a glance whether a number is current you will go and check it in the source system anyway, and now you own both a dashboard and the habit of not trusting it. Build one on figures nobody agreed on and you have something worse than five browser tabs, because it launders a disagreement into a picture that looks settled.
When it works, the weekly meeting starts with a decision instead of a reconciliation, and the argument about whose number is right stops happening. The long version is here.
You move up by finding one decision you make off those numbers the same way every time, where you could write the rule down without hedging. Reorder when cover drops below X weeks. Pause the ad set when CAC crosses Y for three days running. If you cannot write the rule in one sentence, that decision is not ready to be handed over, and knowing that is worth the exercise on its own.
Level 6: agents, and the org chart
At level 6 you have things doing jobs rather than helping with them, on their own, on a schedule or whenever something sets them off, with outcomes that land in the business whether or not anyone looked.
Your org chart now has non-human entries on it, and every question you would ask about a new hire applies. What is its job, in one sentence? Who owns its output? What is it allowed to touch? Who notices when it is doing badly? What happens the day you need to switch it off? Most brands arrive here without ever drawing that box, which is how something consequential ends up running in a company with nobody owning it.
A level 6 agent is usually a scheduled job allowed to finish without stopping to ask, pointed at work that used to need a person's judgment. Routines for Claude Code, which Anthropic shipped in April, package exactly that: work that runs on a schedule, or whenever something else in your business happens.
Mine handles findability for both brands, the work of staying visible in search and in AI answers. It picks topics, drafts, revises and checks its own output against rules I set once, and those rules are strict enough that it frequently produces nothing at all. An empty log is the system working correctly, because a system allowed to publish something mediocre on Thursday just because the schedule said Thursday is worse than no system.
Permissions stop being paperwork here and start keeping you in business. Nothing acts on its own until it has all five of these.
- A backup, taken before it runs.
- A test on a single record, which you check yourself.
- Explicit sign-off before it is allowed to repeat that action.
- A named blast radius. Write down the worst thing it could do if every assumption it makes is wrong, and if you cannot live with that answer, narrow what it can reach until you can.
- A kill switch somebody other than you can reach without needing you to explain it first.
I learned the middle of that list the expensive way. An agent I had given a delete function read a wrong success message, concluded it was working, and looped until several hundred records were gone. It never thought it was failing, which is why the sign-off sits between the test and the repeat, and why the blast radius gets written down before anything runs rather than worked out afterwards from what is missing.
You add capacity without adding headcount in the areas where you have written the rule down, and this is the first level where work happens with you outside the loop entirely. It is also the first one where a mistake is not always recoverable. An agent acting on bad numbers does not hesitate, does not flag it, and does not stop at one.
You move up by adding something whose only job is to have an opinion about the first thing's work, and the authority to stop it. I run this on my email program: drafts get written, a separate check blocks anything that trips a rule or scores below the bar, and I find out from the log.
Level 7: intelligence, or the second brain
Level 6 gives you things that do jobs. Level 7 is what happens when those things share one memory of how your company decides, so a rule you set once gets applied everywhere without you carrying it between rooms.
People call that a second brain. It is the accumulated judgment of the business written down where the systems can read it, so the standard you set in March is still being applied in November by something you have not thought about since.
Block, Jack Dorsey's company, is the clearest public version of this. They built the software their agents run on, called goose, put named agents on top of it, and in July released Buzz, a workspace where agents and employees work in the same channels. The plumbing is what makes it worth studying. Every agent has its own account, its own credentials and its own permissions, and every message, code change and task it performs is logged under its own identity. Their head of AI capabilities put the thesis plainly: every company is going to need a place where humans and agents work together.
Those agents are not floating free with a mandate. They have identities, scoped permissions and a human who owns the outcome, which is the level 6 discipline carried up a rung rather than abandoned once the system got clever.
Your work at this level is setting standards and revising them rather than reviewing output. You decide what good looks like, the system applies it to the daily decisions, and you read the record of what it did instead of approving each thing before it happens. When you find yourself checking individual outputs again, either the standard was never written down precisely enough or it has stopped being the right standard.
Your attention moves from whether today's work met the bar to where the bar should be, which is the only part of the job that was ever really yours. The failure mode is quiet: a standard nobody revisits becomes a policy nobody remembers setting. Reading the record has to actually happen on a schedule, because a second brain you never check is a first brain making decisions without you.
Where this ended up for me
All of my own AI work now happens in Claude Code, and it is one of two things: building something, or research deep enough that I would previously have hired it out. I do not sit in a chat window any more.
The more useful half is what happened to NØRSE CØDE. Most of the work that used to run through me talking to Claude has moved into custom automations, and my job on those is approving what ships. I am not producing the output. I set the standard and say yes or no.
Read that as a destination, not as a starting instruction. I began in a chat window pasting things in, like everyone reading this, and every rung between there and here got climbed one at a time in the order described above. If you are at level two in most of the company, your move is level three, not a developer tool.
Should you start with the financials?
The instinct is to point the first serious build at the numbers, because that is where the pain is loudest and where your own hours are going. Start somewhere else. Financial reporting is the least forgiving place to learn, and you will be learning. It has the most edge cases, the most systems that disagree, the worst consequences when a figure is wrong, and the least tolerance for a week of a system being half-built. Your first three attempts should happen where a bad output costs an afternoon.
Pick customer service replies, review mining, listing copy, or the weekly summary somebody assembles by hand. Get it wrong twice. Learn what this is genuinely good at and what it fabricates without blinking, and set the permissions while the stakes are small enough that a mistake teaches you something.
Then do the financials, with three builds behind you and a clear idea of where it lies to you.
Your next move
Take the area you scored lowest, find its section above, and do the one move it describes. That works from wherever you are standing, which is the point of a ladder.
The full guide
What AI actually does inside a CPG operation, and what it costsThe three kinds of work worth pointing AI at, what each one costs to build, and the work I will tell you to leave alone at your size.
CPG AI Assessment
Fix the CPG workflow that keeps landing back on your desk.
I map the recurring reporting or operating workflow first, then build only when the first useful fix is clear.