Four Mornings of a Retainer Team: The Four Levels of Enterprise AI-Native Maturity
AI-native is not a scale of how much of the old org you have automated — it is a redesign of how people, agents, information, and decisions relate. From four mornings to why a digital employee can charge $20K–200K a year.

First, a word of vocabulary. A retainer is the advertising industry's long-term service model: a brand pays a fixed monthly or yearly fee and hands ongoing work — running its social media accounts, producing content — to an agency, as opposed to one-off, per-project campaigns. A typical retainer engagement has the agency permanently responsible for the brand's Instagram, TikTok, and X accounts: content planning, copy and design, publishing, community interaction, and performance reviews.
Now picture such a retainer team. All of them "use AI" — yet their mornings can look like four entirely different things.
Level 1 morning: everyone clicks away by hand
The company has no system. AI is a chat window each person opens on their own. Copywriters draft with AI, designers generate images with AI, the PM summarizes meetings with AI — but every task is still remembered, created, and pushed forward by a human.
The PM recalls from memory what is due today
HumanDigs through Slack, email, and meeting notes for client requests
HumanOpens a general-purpose LLM and pastes the brand background in — again
HumanThe model produces a first draft
AIRewrites heavily, sources images, manages versions alone
HumanPublishes by hand; the prompts and know-how stay in the PM's head
HumanAI appears only at step 4, and even the brand background has to be spoon-fed every time — prompts and experience walk out the door with each employee. AI capability never becomes an organizational asset.
Level 2 morning: the PM decides what to produce, then works the system
The company now runs a workflow system: brand knowledge base, content calendar, generation tools — all in place. And yet:
The PM remembers, unprompted, that work needs doing today
HumanThe PM opens the system and checks the schedule
HumanThe PM decides how many posts and image sets are needed
HumanThe PM creates or picks the tasks
HumanThe system generates candidates
AIThe PM selects, splices, and edits
HumanThe PM pushes everything to the next stage
HumanSix of the seven steps are driven by the PM; AI appears only at step 5.
The human is the engine of the project; AI is a production tool and a workbench. Execution labor has moved — management attention has not.
Level 3 morning: the agent has already decided and produced — the PM reviews
What the PM sees in the morning is not a piece of software waiting to be operated, but a report on work already advanced:
Today needs two posts and three image sets. I have produced all of them. One involves celebrity likeness rights and needs your judgment; the rest passed internal QA — please review them in one batch.
The human no longer needs to know in advance how much to produce today. Behind this stands a task loop the agent runs on its own (all 11 steps in 1.2 below).
Level 4 morning: the agent doesn't just run the accounts — it lands a hit
The level 4 morning goes one step further: the PM picks up the phone and finds that a piece the agent made has gone viral.
Last night I noticed one strand of user sentiment in the brand's comment sections warming up for the third day straight. I judged the angle had breakout potential, produced and published a piece within my authorized scope. It is now the brand's most-engaged post in six months and is lifting brand search and discussion. I have distilled the success into a reusable content pattern and recommend developing it into a recurring series — plan and data attached, your call.
The viral hit itself is not the point. The point is that this piece was on no agreed schedule: the system found the opportunity, validated it, and then distilled the success into a reusable pattern — turning virality from a lucky accident into a manageable capability (what it takes, in 1.6 below).
The dividing line of this whole essay
The difference between these four mornings is the dividing line this essay is about:
At level 2, the human remembers every day what to ask AI to do. At level 3, the AI remembers what needs doing, and tells the human how far it has gotten. At level 4, the AI doesn't just complete expected work — it keeps delivering results that surprise you.
One thing needs saying up front: an AI-native enterprise is not a rating scale for how much of the old organization you have automated — it is a redesign of the relationship between people, agents, information, and decisions. You cannot measure a company's AI-native level by how many models it has integrated, how many agents it has shipped, or how many employees use AI tools. What actually determines the level are six things — context reach, result closure, attention ownership, continuous learning, permissions to act, and whether processes and value get redefined (each unpacked in Part III · the six dimensions).
On that basis, this essay divides enterprise AI-native maturity into four levels, with software products mapping level for level:
Employees use AI tools on their own; capability never becomes an organizational asset
AI speeds up existing processes; each department runs its own agent (where leading companies are today)
AI redefines how work should happen: digital employees proactively, continuously own responsibility and deliver
AI redefines what work is worth doing, and proactively discovers new value
Figure 1 · The four levels of enterprise AI-native maturity, and what the corresponding product sells at each level.
| Level | Enterprise AI-native maturity | Corresponding product | Typical pricing |
|---|---|---|---|
| Level 1 | Employees use AI tools on their own | Any AI image / text / video generation tool | Per tool or per call |
| Level 2 | AI speeds up existing processes; each department has its own agent | SaaS | $20–200 a month |
| Level 3 | AI redefines how work should happen: digital employees proactively, continuously own responsibility and deliver | Digital employee | Can sell for $20K–200K a year |
| Level 4 | AI redefines what work is worth doing and proactively discovers new value | Value-discovery-and-realization engine | A share of the sales lift it creates |
The essay starts with the concrete retainer case, generalizes to the value levels of software products, and ends with the full enterprise-side framework.
Part I · A close look at the Retainer agent
1.1 What makes retainer work hard
A typical retainer engagement staffs one PM, half a copywriter, and one to two designers, plus shared hours from strategists and managers.
The pain is not just slow asset production. It includes:
- The PM holds a large set of open items in memory every day
- Client requests scatter across Slack, email, and meetings
- Copy and design go through endless revision rounds
- Constant context switching
- Client feedback resists being structured
- Project knowledge lives in individual memory
- Strong people get consumed by low-value work
- One team can only carry so many brands
In other words: the cost of a retainer is not just writing copy and making images — it is the cognitive burden of a PM continuously remembering, scheduling, coordinating, chasing, and closing open items.
1.2 The four product levels of a Retainer agent
Level 1: content generation tools
PMs, designers, and copywriters use scattered AI tools: writing headlines, generating copy, generating images, summarizing trends, drafting report outlines. Humans remain responsible for: remembering tasks, supplying context, organizing the process, editing outputs, managing versions, owning the outcome. What gets delivered: individual content assets.
Level 2: Retainer copilot / workflow SaaS
The PM enters the brand, the month, and the product focus, picks topics; the agent generates a few directions of copy and design; the PM selects, splices, edits, and reviews. The system covers trend-watching, topic selection, copy, artwork, publishing, and retrospectives — product capabilities include a brand knowledge base, monthly roadmaps, topic and content calendars, copy generation, visual briefs, review and version management, data rollups, report generation — but the PM is still the engine of the whole project.
The typical way of working is exactly the seven steps of the level 2 morning above.
Level 3: the Retainer digital employee
The real change at level 3 is not adding more generation modules. It is:
The digital employee takes over the working attention and the open loops of running the accounts.
The agent proactively and continuously owns the operating responsibility — for example, a standing duty like:
Keep the next seven days stocked with enough content that meets the internal first-review bar, at all times.
Every day, or whenever an event fires, it checks content inventory, publishing schedule, client feedback, approval status, trends, and performance data; decides how many posts and image sets today needs; and completes production, QA, and process advancement on its own. One full round of work is a closed loop (the dashed line means: when done, return to step 1 and wait for the next wake-up):
Wakes daily or on events
AgentChecks content inventory, schedule, and approvals
AgentChecks new client requests, trends, and competitor moves
AgentDetermines the current responsibility gap
AgentCreates and prioritizes tasks on its own
AgentCalls topic, copy, design, and QA tools
AgentCompletes production and a first review pass
AgentPushes only the items that need judgment to the PM for review
Agent → humanRevises according to the approvals
AgentUpdates task state and experience
AgentStops once safety stock is reached
AgentThe whole loop is driven by the agent; the human appears only at step 8 — the PM does not need to know in advance what today requires, and simply receives: "This week is short two pieces. I have finished them; one needs your judgment on celebrity likeness rights, the rest can be reviewed in one batch."
This is the complete transfer of attention:
| Level 2 | Level 3 |
|---|---|
| The human remembers that work needs doing | The agent remembers that work needs doing |
| The human enters the system to find tasks | The agent pushes what matters to the human |
| The human maintains the to-do queue | The agent maintains the to-do queue |
| The human decides production volume | The agent decides from the responsibility gap |
| The human pushes the process step by step | The agent advances autonomously to the review point |
| AI offers candidate material | The agent delivers acceptance-ready results |
| The human tracks open items | The agent closes open items |
What level 3 delivers: a portfolio of social accounts that runs steadily with minimal human involvement. (What level 4 looks like: see 1.6.)
1.3 Is level 3 just level 2 plus a scheduler
Structurally, yes — level 3 can be built on top of level 2. It reduces to:
Trend tools, topic tools, copy tools, image tools, publishing tools, data tools.
What is my responsibility right now; what is the current state; how far to the target; should I act; what tasks should I create; which capabilities should I call; when should I ask a human; when should I stop.
If the system merely "generates three posts at nine o'clock every day", it is still scheduled automation. A digital employee is:
At nine o'clock, inspect the responsibility domain; based on current inventory, schedule, trends, approvals, and targets, decide whether to produce, what to produce, how much, and where to stop.
So the timer only wakes the agent up. The real capability of a digital employee is:
Waking up knowing why it works, what to do, and when to stand down.
1.4 Does a Retainer agent need to run nonstop
No. The right concept is not "continuous execution" but:
Continuous responsibility, intermittent execution.
A retainer agent can wake on:
- Daily content-inventory checks
- Weekly prep for next week
- Monthly roadmap generation
- Monthly retrospectives
- Quarterly strategy reviews
- New client feedback arrives
- New assets are uploaded
- Content gets rejected
- A publish fails
- A new trend emerges
- A competitor makes a major move
- Data goes anomalous
- A sensitive comment appears
- Inventory falls below the safety line
- A task nears its deadline
- An approval waits too long
- A content type keeps failing
- An experiment has enough data
- Monthly metrics drift off target
On waking, the agent can reach one of three conclusions: it must act now; it must wait for a condition; or there is nothing worth doing.
A mature digital employee must be able to judge that "now is not the time to keep working." It should never manufacture content and tasks just to look busy.
1.5 How a Retainer agent keeps learning
Start with concrete things worth learning, then sort them into layers. The system should remember, for instance: this client dislikes loud emoji; "premium" means more restraint, not more complexity; a certain product-exposure ratio always gets bounced; the client wants two complete directions per review — that is client-preference learning. Layered, a level 3 retainer digital employee needs at least four kinds:
Remember where the work stands, what is waiting, what needs resuming, what comes next
The examples above: emoji, what "premium" means, exposure ratios, two directions per review
Turn experience into skills: how to handle vague feedback, how to rewrite for each platform, how to spot authenticity risks, how to judge whether a trend is worth chasing, how to keep moving when assets are missing
Let results change: topic weights, content mix, publishing cadence, creative direction, platform roles, experiment priorities
Level 3 requires at least the first three; level 4 demands the fourth.
1.6 Level 4: the Retainer value-discovery engine
What the brand originally asked for may be nothing more than: keep up the social presence, publish on time, support campaigns, avoid PR incidents, send a monthly report.
Getting the agreed account operations done, reliably.
Discovering and realizing value in the account portfolio the brand never knew was there.
The agent doesn't just maintain the brand's social presence — it keeps finding and realizing more value across the account portfolio. It proactively spots content directions, user sentiments, and brand narratives with breakout potential; grows one high-performing piece into a recurring series, a content IP, and a stable brand persona; turns virality from a lucky outcome into a capability that can be learned, reused, and managed; and ultimately produces verifiable lift in brand awareness, search, user preference, and sales.
- Which content has genuine breakout potential
- Which user sentiment is worth owning long-term
- What distinctive social persona the brand can form
- Which comments can grow into recurring segments
- Which success can be replicated into a series
- What role each platform should play
- Which content drives search and purchase, not just engagement
- Which routine content has no value and should stop
- Which content can become a long-term brand asset
- A rising share of high-performing content
- A markedly higher hit rate
- Users following and discussing unprompted
- A stable brand social persona
- Long-running segments and content IP
- The accounts reinforcing one another as a matrix
- Growth in brand search, preference, and discussion
- Verifiable incremental impact on sales
"Make every post go viral" works as an internal ambition, not as a per-post contractual promise. The more accurate definition of level 4 is:
The system continuously discovers, validates, and reuses high-value content patterns, gradually turning virality and brand lift from chance events into a manageable capability.
1.7 What justifies $20K–200K a year
Level 2 enterprise SaaS can also reach $20K a year, especially bundled with customization, deployment, data integration, and services. But to charge on the logic of a "digital employee" rather than "expensive SaaS," you must prove three things:
- The PM no longer enters the system to find work
- The agent creates the necessary tasks
- The agent pushes items for review
- Open items no longer depend on human memory
- Hours per brand per month drop visibly
- One PM supervises more brands
- Copy and design lose their repetitive load
- Project knowledge no longer lives in one person
- The same team serves more brands
- On-time delivery rate
- Unedited-into-review rate
- Rejection / rework rate
- Human-takeover rate
- High-risk error rate
- Client satisfaction
- Project gross margin
A suggested core dashboard for a digital employee:
| Metric | What it measures |
|---|---|
| Share of tasks created by humans | Whether work still relies on someone remembering it |
| Share of agent-created tasks that prove useful | Whether initiative creates real value |
| Human management minutes per brand per week | Whether the attention burden is falling |
| Times humans enter the system unprompted | Whether the product is pull or push |
| Auto-closure rate of open items | Whether the agent pushes work through to the end |
| Average autonomous steps before review | How long a responsibility chain the agent carries |
| Unedited-into-review rate | How close output is to deliverable |
| Rejection / rework rate | Quality stability |
| Repeat-error rate | Whether it truly keeps learning |
| Brands supervised per PM | Whether organizational capacity changed |
| Gross margin per brand | Whether financial value materialized |
The pricing logic in one sentence:
No client pays $20K–200K a year because you have seven agent modules. They pay because they no longer need to schedule, remember, and push this work every day.
1.8 Team and resources at each stage
Roles: one AI-native full-stack engineer; one advertising-business or content product lead; design support as needed.
Goals: validate copy and image quality; find high-frequency point tasks; get a usable prototype fast.
Roles: one product lead who deeply knows retainer work; one senior full-stack or tech lead; one to two AI application engineers; one knowledge-and-evals engineer; one product designer or agent-operations person; PMs, copywriters, and designers as a standing co-creation group.
Goals: build the brand knowledge base; connect the content workflow; cut generation and collaboration costs; measure true labor-hour changes precisely; find the responsibility loop best suited for a digital employee to take over.
Roles: one AI product architect; one agent systems architect; two to three AI-native full-stack engineers; one state/memory/platform engineer; one evals-and-learning engineer; one to two agent-operations people; one senior retainer business lead (role duties and cost assumptions in Part III · the talent ladder).
Key capabilities: event-driven systems, durable state machines, durable execution, timer and event triggers, task queues, dynamic workflows, forward-looking tasks, idempotency/retries/rollback, human-in-the-loop approval, agent observability, memory and skill updates, production-grade evals.
The minimum responsibility loop: do not open by promising to "fully take over the brand's accounts." A more realistic first duty is — keep the next seven days stocked with enough content that meets the internal first-review bar, at all times. It naturally contains: daily inspection, task discovery, content production, QA, approval, restocking, and a stop condition.
Added on top of stage 3: one brand strategy lead; one senior creative lead; one to two marketing data scientists; one data engineer; one experimentation/growth engineer; one multimodal content-understanding engineer; more content-quality and agent-operations support.
- Brand strategy lead: brand equity, user culture, platform context, content motifs, social persona, long-term brand value.
- Senior creative lead: turns the signals the system finds into segments, series, content IP, creative motifs, and cross-platform matrices.
- Marketing data scientists: content-to-sales attribution, incrementality experiments, causal analysis, content and audience clustering, portfolio optimization, separating vanity metrics from real value.
The scarcest person at level 4 is not the agent engineer, but:
someone who can read a data anomaly as a brand opportunity, and turn that opportunity into a long-term content asset.
Part II · The value levels of software products
The four retainer levels are not specific to that industry. Generalize "generation tool → copilot → digital employee → value-discovery engine" and you get the value levels of software products at large:
Level 1 sells features. Level 2 sells process efficiency. Level 3 sells role capacity and attention transfer. Level 4 sells growth and value discovery.
| Product level | Form | Division of labor | Typical pricing | What it sells |
|---|---|---|---|---|
| Level 1 · Point AI tools | Any AI image / text / video tool, summarizers, simple chatbots, one-shot automation plugins | The product delivers one result; the user holds all context, tasks, and responsibility | Priced like software tools, model calls, convenience | Point features |
| Level 2 · Ordinary SaaS / copilot | Centralized context, standardized process, AI generation, approval and analytics | Humans still remember tasks, create tasks, push the process, own the outcome | $20–200 a month; complex enterprise editions higher | Software capability and process efficiency |
| Level 3 · Digital employee | A software entity responsible for a defined, ongoing role | The agent creates and advances tasks on its own; humans review the key judgments | $2K–10K a month, benchmarked against role cost, capacity, and management attention | Role capacity, closed-loop responsibility, attention transfer |
| Level 4 · Value-discovery engine | Proactively finds better goals and new opportunities | The agent proposes and validates value hypotheses; the company reallocates resources | Benchmarked against incremental revenue and margin; can take a share of the sales lift | The sustained ability to find and realize growth |
The four levels also differ in their basic unit:
The unit of a tool is the feature. The unit of a workflow is the process. The unit of an agent is the task. The unit of a digital employee is the responsibility of a role.
Level 1: point AI tools
Typical products: copy generators, image generators, video tools, summarizers, simple chatbots, one-shot automation plugins. The product delivers a result, but the user still holds all context, tasks, and responsibility. Priced against software tools, model calls, and convenience.
Level 2: ordinary SaaS / copilot
The internal workflow products of leading companies sit roughly here — in retainer terms, the copilot form of 1.2.
- Centralize context
- Standardize the process
- Generate with AI
- Provide approval and analytics
- Reduce tool switching
- Raise the productivity of existing staff
- Remember the tasks
- Create the tasks
- Enter the system
- Push the process
- Track open items
- Decide the next step
- Own the final outcome
Typical pricing: $20–200 a month, more for complex enterprise editions. What this level sells: software capability and process efficiency.
Level 3: the digital employee
The basic definition of a digital employee:
A software entity that holds a defined, ongoing role responsibility.
Long-lived working state, a work calendar, open items, forward-looking tasks, priorities, durable context, approval rules, stop conditions.
Be woken by timers, business events, or state changes; judge whether a responsibility gap exists; create tasks autonomously; orchestrate workflows dynamically; call tools to execute; handle routine exceptions; push only what matters to humans; keep learning from feedback and outcomes; stand down once the responsibility is met.
A digital employee is not required to be busy all the time. It should be: continuously responsible, intermittently executing, waking on demand, sleeping when done.
A digital employee is not an agent that completes tasks. It is a software entity that permanently occupies a role, answers for long-term goals, and keeps finding its own next piece of work. Its core capability is not automated execution but taking over working attention. What is truly expensive is not the labor of generation — it is having someone keep the job on their mind.
Typical pricing: $2K–10K a month, benchmarked against role cost, capacity, and management attention (how to prove you deserve it: the three proofs in 1.7). What this level sells: sustained role capacity, closed-loop responsibility, attention transfer, changed staffing, delivery reliability.
Level 4: the value-discovery engine
A value-discovery engine carries the original responsibility and, beyond it, proactively finds: which goals are more worth pursuing, which work has no value, which opportunities the company has not seen, which experiments could create new revenue and profit.
Its pricing no longer benchmarks only software and labor. It can benchmark: incremental revenue, margin improvement, cost reduction, new business opportunities, avoided bad investments. Business models can include: base service fees, deployment fees, ongoing platform fees, revenue share on increments, profit-improvement share.
What this level sells: the sustained ability to discover and realize growth opportunities.
Part III · Enterprise AI-native maturity
Finally, the enterprise side: where does a whole company sit? First watch two departments across the four levels, then the full definitions of the four levels and the six dimensions that decide the answer.
3.1 Four levels through two departments
HR, level by level
HR uses scattered AI tools — writing JDs, summarizing interviews, generating interview questions — but the capability depends on individuals and never accrues to the organization.
An HR agent matches people to roles, screens résumés, builds the talent pool, assists interviews and funnel analytics — but the definition of talent still comes from business units' gut feel. The business says "we need a senior PM," and the agent helps HR find a senior PM faster.
AI connects HR with sales, operations, delivery, and finance across the company, and redefines how talent demand is generated. It no longer waits for departments to file hiring requests; it continuously analyzes project margins, client reviews, delivery quality, employee performance, and workload to surface the real capability bottlenecks. It might conclude that growth is limited not by headcount but by senior PMs' scarce capacity being consumed by low-value coordination — and proactively propose a new structure of "senior PM + digital employees + junior staff," then keep driving hiring, training, role redesign, and agent substitution.
AI redefines HR's value from first principles: HR is not hiring and managing people, but keeping the organization's capability supply matched to the demands of value creation. It judges which capabilities to hire, which to train, which to outsource, and which to hand to agents, continuously reallocating organizational resources so scarce talent always works on the most valuable things.
Marketing ops, level by level
Ops uses scattered AI tools — writing copy, generating images, cutting video, summarizing data — but everyone still manages their own tasks and context.
An ops agent maintains the prompt pool, mass-produces video and ad creative, feeds data back, analyzes and iterates automatically — making the existing production and delivery pipeline more efficient.
AI no longer waits for operators to type in topics and requirements; it continuously owns the operating responsibility for an account, a channel, or a growth target. It watches inventory, schedules, trends, competitors, and performance; creates tasks, produces content, initiates reviews, tracks publishing and retrospectives on its own; and pushes only strategic questions and high-risk items to the operators. Humans stop reminding themselves what to produce — the agent tells them what has been advanced and what still needs their judgment.
AI redefines the value of operations: not producing content continuously, but continuously discovering and occupying value gaps. It analyzes user sentiment, competitors' blind spots, platform shifts, and business results; finds new content motifs, audience segments, distribution scenarios, and growth opportunities; and validates them through experiments. Content production is merely the last step after a value judgment lands.
3.2 The four maturity levels defined
Level 1: the personal-tools company
Employees use general-purpose LLMs, AI coding tools, generation tools, and automation scripts on their own. AI capability is not yet an organizational asset: results depend on individual skill; prompts and experience leave with employees; data never flows back; workflows do not change; the financial contribution cannot be measured. At bottom it is still a traditional company in which some employees happen to use AI.
Level 2: the department-workflow company
Leading companies sit roughly here. The company can speed up each department's existing workflows with agents, producing cost, capacity, and financial improvements visible at group level — the level 2 of the HR and ops examples above. But the premise of this level is:
Do not redefine the department's value; only optimize what the department already does.
- One or two senior developers, or AI-native young developers
- One PM able to model the department's workflows
- A department head willing to share process and data
- A degree of CEO will
- Basic data access and permissions
Task time, manual steps removed, adoption rate of generated content, department throughput, per-department labor cost, unit output cost, financial impact on revenue and expenses.
What this level sells: doing existing things faster and cheaper.
Level 3: the process-rebuilt AI-native company
Level 3 stops treating existing departments and processes as immovable premises. It redefines from scratch: how information flows, how work gets created, how departments collaborate, how humans and agents divide labor, which processes stay, which processes go — the level where, in the examples above, HR discovers "the growth ceiling is senior-PM scarcity" and ops proactively owns account responsibility.
Break the department walls; build enterprise-level context. The reason those level 3 examples are impossible at level 2 is that a traditional company's information is scattered across nine domains: sales, operations, product, R&D, production and delivery, support, finance, HR, management. Each department sees only local facts, so it can only optimize locally. The first thing level 3 builds is not more agents but a unified enterprise context, so the system can continuously understand:
- Why customers buy or churn
- Which projects earn revenue but lose money
- Which people and capabilities actually drive delivery
- Which problems keep recurring across departments
- Which internal goals conflict
- Which link constrains the whole company's growth
- Which business assumptions have expired
This is fundamentally a question of CEO resolve and CTO field of view. The CTO and the AI product lead must: ① observe every department's real workflows; ② map the information and efficiency gaps; ③ redefine the data structures; ④ work with the CEO to break departmental isolation; ⑤ turn information assets into durable context agents can use.
Continuous responsibility and the autonomous task loop. A level 3 system no longer depends on people entering software to create tasks. One full round of its work (isomorphic to the retainer employee's 11-step loop):
Continuously maintains goals and working state
AgentWakes on schedule or on business events
AgentJudges the gap between current state and the responsibility target
AgentCreates and prioritizes tasks autonomously
AgentOrchestrates execution dynamically
AgentCalls tools to complete the work
AgentAsks humans to handle high-risk exceptions
Agent → humanAccepts and verifies results
AgentUpdates state and experience
AgentStops when the responsibility is met, until the next wake-up
AgentThe human appears only at step 7, as the handler of high-risk exceptions; the dashed line means each round returns to step 1.
So the core change of level 3 is:
Working attention transfers from humans to the system.
Continuous, autonomous learning. A level 3 agent must keep learning from: human edits, content acceptance and rejection, project wins and failures, client satisfaction, business results, execution anomalies, and its own errors. What it learns must be written into: long-term memory, business rules, prompts, eval sets, skills, model routing, dynamic workflows, decision policies. Early on the model need not modify its own weights — but one thing must hold:
The same mistake is never repeated forever; the same success can be reproduced reliably.
The talent ladder level 3 requires:
Observes workflows across the company, defines responsibility boundaries, builds the business state model, sets the human–agent division of labor, defines what "closed loop" and "deliverable" mean, aligns the CXOs. This role cannot be a collector of feature requests.
Scope: agent orchestration, memory systems, tool calling, RAG, long-task state, dynamic planning, human-in-the-loop, failure recovery, quality control, the learning loop. Needs to build three layers: ① grounded reasoning; ② quality checks; ③ experience learning.
The goal: a production-grade, 95-point closed loop that finishes the job itself — not a thin wrapper over a model API. One of the hardest roles to hire, and one that most determines the product's moat.
Scope: prompt engineering, model selection and routing, business scoring rubrics, offline evals, online quality monitoring, error ledgers, regression eval sets, failure attribution, training and fine-tuning.
The metrics that matter most: loop-closure rate, unedited-into-review rate, unedited-direct-publish rate (only where contract and risk allow), rejection/rework rate, human-takeover rate, cross-version regression rate.
Evals remain a scarce, dedicated craft in today's market — and a digital employee's price rests precisely on proving these numbers.
Scope: data-ingestion pipelines, structuring the company's information assets, backend, frontend product engineering, permissions, auditing, infrastructure, enterprise integrations, production stability. The value is shipping the complete product from data to systems to user experience — and data ingestion and structuring is usually the single largest block of work.
Scope: prototype exploration, tool integration, testing, data labeling, routine pipeline runs, eval-sample preparation, new-model experiments, edge automation. They can carry high-frequency experiments and the grunt work, but should not own evals, and should not own production architecture alone.
Scope: catching false and missed wake-ups, catching junk tasks, analyzing failure traces, cleaning contaminated memory, maintaining client-specific rules, updating skills and eval sets, and watching whether the agent truly reduces human attention.
At the current state of the art, "continuous learning" does not complete itself. Agent operations is the training, management, and performance department for digital employees.
Organizational preconditions: strong CEO backing; a CTO with a cross-department field of view; accessible information assets; walls that can actually be broken; solid permissions and auditing; a real business as a continuous training ground; a management team willing to redesign roles and processes.
Level 4: the value-discovering AI-native company
How does the company keep running in a better way?
What value should the company create — and what value has no one discovered yet?
Level 4 does not merely hand tasks to digital employees — it redefines, from first principles, why each department exists: the level where, in the examples above, HR becomes a "capability allocation system" and ops becomes a "market-cognition, value-validation, opportunity-occupation system".
The level 4 enterprise loop:
Continuously ingest signals from inside and outside the company
Discover value gaps
Form business hypotheses
Design low-cost experiments
Organize execution
Judge the true increment
Scale the successful patterns
Write the experience back into the company's cognition system
Seek the next, higher-value target
The dashed line means each round returns to step 1 — a standing value-discovery loop.
Preconditions: the level 3 talent ladder is in place; company-wide information assets are connected; department walls are down; the CEO is willing to re-examine the existing business; the CTO tracks the moving frontier of AI capability; agents are allowed to question existing tasks; agents are allowed to drive experiments and resource reallocation; performance shifts from "how many tasks completed" to "how much value created."
3.3 Six dimensions that decide your level
Strip away all the examples, and what actually determines where a company (or a product) sits are six things:
| Determinant | The question it asks |
|---|---|
| 1 · How far context reaches | Individual, department, whole company, or inside-plus-outside |
| 2 · How far results close the loop | Delivering assets, assisting processes, owning a role, or creating business increments |
| 3 · Who holds the working attention | Whether humans still remember tasks, enter systems, and push processes |
| 4 · Whether the system keeps learning | Whether feedback, failure, and business outcomes change future behavior |
| 5 · How much permission to act it holds | Whether it can read and act across departments and systems |
| 6 · Whether process and value get redefined | Speeding up old processes, or re-deciding what work is worth doing |
The progression of the six dimensions across the four levels, as one matrix:
| Dimension | Level 1 | Level 2 | Level 3 | Level 4 |
|---|---|---|---|---|
| Context reach | What the user pastes in | One department's knowledge, process, history | Company-wide information assets | Inside + continuous absorption of the outside |
| Result closure | Delivers single assets | Delivers process efficiency | Delivers role responsibility | Delivers business increments |
| Attention ownership | Humans hold all context, tasks, responsibility | Humans remember what to ask AI | AI remembers, and informs humans | AI finds more worthwhile goals itself |
| Continuous learning | No accumulation; experience leaves with people | In-department data feedback and asset iteration | State + preference + skill learning | Plus outcome-strategy learning |
| Permissions & org | Personal tools, no enterprise access | In-department data access | Cross-department permissions; governance changes with it | Company-wide assets connected; agents may question tasks |
| Process & value definition | Processes unchanged | Optimize what exists; value undisturbed | Redefine processes and the human–agent split | Redefine what work is worth doing |
Dimension 1: context reach
Level 1: AI has only what the user pastes in. Level 2: AI reads one department's knowledge, processes, and history. Level 3: AI connects the information assets of sales, operations, product, delivery, R&D, support, finance, and HR. Level 4: AI understands not just the inside of the company but continuously absorbs the market environment, competitors, user behavior, technology shifts, culture, supply chains, and industry change.
Dimension 2: result closure and pricing class
What decides a product's level is not how smart the model is, but:
What result it actually delivers, and how far that result closes the loop.
| Level | Deliverable | Example |
|---|---|---|
| Level 1 | A single asset | Generate one post, one image, one summary |
| Level 2 | Process efficiency | Help the PM finish topics, copy, design, reports faster |
| Level 3 | Role responsibility | Keep the review queue stocked for the next seven days, continuously |
| Level 4 | Business increments | Turn an ordinary set of brand accounts into a content asset that compounds brand awareness and sales |
Commercial value can be understood as:
Dimension 3: who holds the working attention
This is the decisive boundary between levels 2 and 3 — the abstract version of the second and third mornings that opened this essay.
Humans still must:
- Remember what today requires
- Enter the system unprompted
- Create tasks
- Order the tasks
- Track what is unfinished
- Decide when to continue
AI saves execution labor, but the open loops still live in a human head.
The digital employee owns:
- The work calendar
- The to-do queue
- Content inventory
- Open items
- Next actions
- Wake-up timing
- Task priority
Humans stop hunting for work — the system informs them.
Level 2 is humans entering a system to find tasks. Level 3 is the system pushing the judgment calls to the human. The true mark of attention transfer is not how much content AI generates — it is that no human needs to remember tasks, maintain to-dos, track open items, or decide when to continue.
Which means AI has taken over not just the labor, but the attention of management.
Dimension 4: the capacity to keep learning
Continuous learning is not just online training of model weights. An enterprise agent's learning has at least five layers (the retainer agent's four kinds in 1.5 are the concrete form of the first four):
Remember where the work stands
Remember how clients and managers edit the output
Turn successful experience into rules, skills, and workflows
Let business outcomes change future resource allocation and behavior
Update model weights directly without forgetting old capabilities
An enterprise digital employee needs at least the first three; the fourth underpins the value-discovery engine. The fifth remains frontier research on base models — not a precondition for the product to work.
Dimension 5: permissions and org structure
What an agent can create is bounded by the information and permissions it holds. An HR agent that can only read résumés can only optimize screening. Give it access to performance data, project margins, client reviews, delivery quality, collaboration graphs, and future business plans — and it may redefine what the company actually lacks.
That is why moving from level 2 to level 3 is not merely a technical upgrade but a joint shift in data governance, permission governance, department boundaries, incentive design, and decision rights.
Dimension 6: observe, distill, define
Genuine AI-native product definition is not collecting feature requests from business units. It is three things:
Sit beside PMs, copywriters, designers, HR, sales — watch the real work:
- Where information comes from
- How people make judgments
- Where context keeps switching
- Which work is mere habit
- Which problems actually move profit
Distill messy labor into:
- Goals
- State
- Constraints
- Judgment criteria
- Exception types
- Responsibility boundaries
- Verifiable results
Redefine:
- Which tasks deserve to stay
- Which processes should disappear
- Which responsibilities go to agents
- Where humans must step in
- What the product finally delivers
Final summary
The four levels of the AI-native enterprise, the four levels of software products, and the retainer agent's levels compress into one table:
| Level | AI-native enterprise | What the product sells | The Retainer agent |
|---|---|---|---|
| Level 1 | Employees use AI tools | Features | Everyone clicks away by hand |
| Level 2 | AI helps departments do existing work faster | Process and efficiency | The human decides what to produce, then works the system |
| Level 3 | AI takes over role responsibility, working attention, and open items — redesigning how work happens | Closed-loop responsibility, role capacity, attention transfer | The agent decides and produces daily, then calls the human to review |
| Level 4 | AI re-judges from first principles what work is worth doing, and discovers and creates greater value | Growth, profit, and value discovery | The agent runs the accounts steadily and keeps realizing brand and sales value far beyond social presence |
So the most accurate definition of a digital employee:
A digital employee is not an agent that completes tasks, but a software entity that permanently occupies a role, holds the working attention, answers for long-term responsibility, and keeps learning from experience.
And a value-discovery engine:
It doesn't just keep doing the job — it keeps discovering how much more the job was worth all along.
Appendix A · How the market and frontier labs define the digital employee
A.1 There is no unified definition
The previous generation of automation vendors defined digital workers as software robots that execute enterprise processes end to end. Automation Anywhere emphasizes a "virtual workforce" of AI, machine learning, RPA, and analytics completing sequences of business tasks; UiPath now describes a single AI agent as an autonomous digital worker that completes a specific task from start to finish. (Automation Anywhere)
Salesforce uses the concept of "digital labor," stressing agents with defined roles, enterprise knowledge, the ability to act, and around-the-clock availability; Microsoft stresses managing agents like employees: define the role, limit the permissions, supervise continuously. (Salesforce)
So today's de facto market floor is roughly:
Anything that autonomously completes a set of tasks a role used to own can be packaged as a digital employee.
This essay uses a stricter definition:
A digital employee is a software entity accountable for a defined, ongoing role responsibility. It maintains working state and open items; proactively inspects its responsibility domain on schedules and business events; discovers and creates its own tasks; orchestrates workflows dynamically; pushes only necessary judgments and final results to humans; and updates its memory, skills, and policies from feedback and business outcomes.
The heart of it is not "a model running 24 hours a day," but:
The responsibility persists, and the system can resume its own work whenever needed.
A.2 What the frontier labs are building toward
Frontier agent research is no longer chasing single-shot reasoning alone; it is filling in the capabilities a long-lived working entity requires:
Can an agent push one goal across hours, sessions, even days? OpenAI has productized long-horizon tasks, scheduled resumption, and multi-day continuation; Anthropic treats cross-context-window state handoff, long-task recovery, and harness design as the key problems of long-horizon agents. (OpenAI Developers)
An agent needs to know more than "what happened": where the work stands, what the current plan is, which paths already failed, why it paused, what it is waiting for, what remains open. Microsoft's STATE-Bench now evaluates whether agents improve later enterprise tasks with past experience; work like PlugMem turns raw interaction into reusable long-term knowledge. (Microsoft)
Waking on schedules; waking on business events; resuming paused work when conditions are met; carrying responsibility without a direct prompt. OpenAI's Workspace Agents and Codex Automations already productize scheduled execution, repeated runs, retained context, and future automatic resumption. (OpenAI)
Google's ReasoningBank distills successes and failures into reusable reasoning strategies; Titans, MIRAS, and Nested Learning explore longer-horizon memory updates and continual learning. (Google Research)
The frontier labs increasingly recognize that capability comes not only from the base model but from everything around it: context organization, task state, tool environments, memory systems, execution logs, eval sets, failure recovery, human approval. Anthropic emphasizes that harness design significantly affects an agent's ability to advance work across sessions; OpenAI treats durable state and long-task harnesses as core agent engineering. (Anthropic)
Which is to say:
A digital employee is not a bigger prompt, nor a model that never stops running. It is a production system composed of model capability, durable state, wake-up systems, dynamic planning, tool execution, evals, and continuous learning.
Appendix B · From discovery to operations: how the product gets built
B.1 Finding the need
Need-finding requires the CEO, the CMO, or someone with deep market experience. The goal is not collecting feature suggestions but understanding: who hurts most; whether the problem is frequent; how it is solved today; why existing solutions fail; who owns the budget; who bears the risk of failure; how much the problem costs in profit; and why AI can solve it now.
The final deliverable is not a feature list but: a set of testable value hypotheses.
B.2 Defining the product
Product definition takes the CTO, the AI product lead, AI-native engineers, and business experts together. The core method: observe → distill → define (the same method as dimension 6, here as the working version for product definition).
- Read the company's information assets
- Sit beside different roles and watch
- Run targeted interviews
- Run focus groups
- Debrief wins and failures
- Log context switches and rework
Distill the real work into:
- Responsibility
- Inputs
- State
- Constraints
- Judgments
- Exceptions
- Outputs
- Acceptance criteria
- What the product is responsible for
- What humans are responsible for
- Which processes get rebuilt
- Which processes get cancelled
- When the agent may act autonomously
- When a human must decide
- Whether this is SaaS, a digital employee, or a value-discovery engine
This stage requires every CXO to converge on which level the product is.
B.3 Design and engineering
A senior product-engineering team covering four blocks:
Data ingestion, structuring documents and chats, permissions, knowledge bases, project state, long-term memory.
Triggers, planner, task queues, tool calling, dynamic workflows, approvals, failure recovery, evaluator, stop conditions.
Frontend, backend, enterprise integrations, auditing, observability, cost control, multi-tenancy, data isolation.
Quality rubrics, golden datasets, error ledgers, regression tests, online sampling, comparisons across models and workflows.
B.4 Running it in production
An AI-native product needs continuous operations after launch — and operations here is not sales plus support. It is running:
The core work: measure the unedited-into-review rate, the unedited-direct-publish rate, the rejection/rework rate; analyze failure traces; keep an error ledger; fold repeat errors into the regression set; update brand memory, rules, and skills; study why humans take over; prove value with real project financials.
Early go-to-market should lean on hard numbers from lighthouse projects, not agent demos: how many labor hours dropped, how many management minutes per brand dropped, how many more brands one team can serve, how much gross margin improved, whether delivery quality held.