Four Mornings of a Retainer Team: The Four Levels of Enterprise AI-Native Maturity

AI-native is not a scale of how much of the old org you have automated — it is a redesign of how people, agents, information, and decisions relate. From four mornings to why a digital employee can charge $20K–200K a year.

Cover Image for Four Mornings of a Retainer Team: The Four Levels of Enterprise AI-Native Maturity

First, a word of vocabulary. A retainer is the advertising industry's long-term service model: a brand pays a fixed monthly or yearly fee and hands ongoing work — running its social media accounts, producing content — to an agency, as opposed to one-off, per-project campaigns. A typical retainer engagement has the agency permanently responsible for the brand's Instagram, TikTok, and X accounts: content planning, copy and design, publishing, community interaction, and performance reviews.

Now picture such a retainer team. All of them "use AI" — yet their mornings can look like four entirely different things.

Level 1 morning: everyone clicks away by hand

The company has no system. AI is a chat window each person opens on their own. Copywriters draft with AI, designers generate images with AI, the PM summarizes meetings with AI — but every task is still remembered, created, and pushed forward by a human.

Human-drivenAgent / system-driven

The PM recalls from memory what is due today

Human

Digs through Slack, email, and meeting notes for client requests

Human

Opens a general-purpose LLM and pastes the brand background in — again

Human

The model produces a first draft

AI

Rewrites heavily, sources images, manages versions alone

Human

Publishes by hand; the prompts and know-how stay in the PM's head

Human

AI appears only at step 4, and even the brand background has to be spoon-fed every time — prompts and experience walk out the door with each employee. AI capability never becomes an organizational asset.

Level 2 morning: the PM decides what to produce, then works the system

The company now runs a workflow system: brand knowledge base, content calendar, generation tools — all in place. And yet:

The PM remembers, unprompted, that work needs doing today

Human

The PM opens the system and checks the schedule

Human

The PM decides how many posts and image sets are needed

Human

The PM creates or picks the tasks

Human

The system generates candidates

AI

The PM selects, splices, and edits

Human

The PM pushes everything to the next stage

Human

Six of the seven steps are driven by the PM; AI appears only at step 5.

The human is the engine of the project; AI is a production tool and a workbench. Execution labor has moved — management attention has not.

Level 3 morning: the agent has already decided and produced — the PM reviews

What the PM sees in the morning is not a piece of software waiting to be operated, but a report on work already advanced:

THE DIGITAL EMPLOYEE'S MORNING BRIEF

Today needs two posts and three image sets. I have produced all of them. One involves celebrity likeness rights and needs your judgment; the rest passed internal QA — please review them in one batch.

The human no longer needs to know in advance how much to produce today. Behind this stands a task loop the agent runs on its own (all 11 steps in 1.2 below).

Level 4 morning: the agent doesn't just run the accounts — it lands a hit

The level 4 morning goes one step further: the PM picks up the phone and finds that a piece the agent made has gone viral.

THE VALUE-DISCOVERY ENGINE'S MORNING BRIEF

Last night I noticed one strand of user sentiment in the brand's comment sections warming up for the third day straight. I judged the angle had breakout potential, produced and published a piece within my authorized scope. It is now the brand's most-engaged post in six months and is lifting brand search and discussion. I have distilled the success into a reusable content pattern and recommend developing it into a recurring series — plan and data attached, your call.

The viral hit itself is not the point. The point is that this piece was on no agreed schedule: the system found the opportunity, validated it, and then distilled the success into a reusable pattern — turning virality from a lucky accident into a manageable capability (what it takes, in 1.6 below).

The dividing line of this whole essay

The difference between these four mornings is the dividing line this essay is about:

THE ONE SENTENCE THAT MATTERS MOST

At level 2, the human remembers every day what to ask AI to do. At level 3, the AI remembers what needs doing, and tells the human how far it has gotten. At level 4, the AI doesn't just complete expected work — it keeps delivering results that surprise you.

One thing needs saying up front: an AI-native enterprise is not a rating scale for how much of the old organization you have automated — it is a redesign of the relationship between people, agents, information, and decisions. You cannot measure a company's AI-native level by how many models it has integrated, how many agents it has shipped, or how many employees use AI tools. What actually determines the level are six things — context reach, result closure, attention ownership, continuous learning, permissions to act, and whether processes and value get redefined (each unpacked in Part III · the six dimensions).

On that basis, this essay divides enterprise AI-native maturity into four levels, with software products mapping level for level:

LEVEL 1
Personal tool use

Employees use AI tools on their own; capability never becomes an organizational asset

Product sells: Point features
LEVEL 2
Department workflow gains

AI speeds up existing processes; each department runs its own agent (where leading companies are today)

Product sells: Software and process efficiency
LEVEL 3
Process rebuild & digital employees

AI redefines how work should happen: digital employees proactively, continuously own responsibility and deliver

Product sells: Role capacity, closed-loop responsibility, attention transfer
LEVEL 4
Value discovery & creation

AI redefines what work is worth doing, and proactively discovers new value

Product sells: Revenue, profit, and new value opportunities

Figure 1 · The four levels of enterprise AI-native maturity, and what the corresponding product sells at each level.

LevelEnterprise AI-native maturityCorresponding productTypical pricing
Level 1Employees use AI tools on their ownAny AI image / text / video generation toolPer tool or per call
Level 2AI speeds up existing processes; each department has its own agentSaaS$20–200 a month
Level 3AI redefines how work should happen: digital employees proactively, continuously own responsibility and deliverDigital employeeCan sell for $20K–200K a year
Level 4AI redefines what work is worth doing and proactively discovers new valueValue-discovery-and-realization engineA share of the sales lift it creates

The essay starts with the concrete retainer case, generalizes to the value levels of software products, and ends with the full enterprise-side framework.


Part I · A close look at the Retainer agent

1.1 What makes retainer work hard

A typical retainer engagement staffs one PM, half a copywriter, and one to two designers, plus shared hours from strategists and managers.

The pain is not just slow asset production. It includes:

  • The PM holds a large set of open items in memory every day
  • Client requests scatter across Slack, email, and meetings
  • Copy and design go through endless revision rounds
  • Constant context switching
  • Client feedback resists being structured
  • Project knowledge lives in individual memory
  • Strong people get consumed by low-value work
  • One team can only carry so many brands

In other words: the cost of a retainer is not just writing copy and making images — it is the cognitive burden of a PM continuously remembering, scheduling, coordinating, chasing, and closing open items.

1.2 The four product levels of a Retainer agent

Level 1: content generation tools

PMs, designers, and copywriters use scattered AI tools: writing headlines, generating copy, generating images, summarizing trends, drafting report outlines. Humans remain responsible for: remembering tasks, supplying context, organizing the process, editing outputs, managing versions, owning the outcome. What gets delivered: individual content assets.

Level 2: Retainer copilot / workflow SaaS

The PM enters the brand, the month, and the product focus, picks topics; the agent generates a few directions of copy and design; the PM selects, splices, edits, and reviews. The system covers trend-watching, topic selection, copy, artwork, publishing, and retrospectives — product capabilities include a brand knowledge base, monthly roadmaps, topic and content calendars, copy generation, visual briefs, review and version management, data rollups, report generation — but the PM is still the engine of the whole project.

The typical way of working is exactly the seven steps of the level 2 morning above.

Level 3: the Retainer digital employee

The real change at level 3 is not adding more generation modules. It is:

The digital employee takes over the working attention and the open loops of running the accounts.

The agent proactively and continuously owns the operating responsibility — for example, a standing duty like:

Keep the next seven days stocked with enough content that meets the internal first-review bar, at all times.

Every day, or whenever an event fires, it checks content inventory, publishing schedule, client feedback, approval status, trends, and performance data; decides how many posts and image sets today needs; and completes production, QA, and process advancement on its own. One full round of work is a closed loop (the dashed line means: when done, return to step 1 and wait for the next wake-up):

Human-drivenAgent / system-driven

Wakes daily or on events

Agent

Checks content inventory, schedule, and approvals

Agent

Checks new client requests, trends, and competitor moves

Agent

Determines the current responsibility gap

Agent

Creates and prioritizes tasks on its own

Agent

Calls topic, copy, design, and QA tools

Agent

Completes production and a first review pass

Agent

Pushes only the items that need judgment to the PM for review

Agent → human

Revises according to the approvals

Agent

Updates task state and experience

Agent

Stops once safety stock is reached

Agent

The whole loop is driven by the agent; the human appears only at step 8 — the PM does not need to know in advance what today requires, and simply receives: "This week is short two pieces. I have finished them; one needs your judgment on celebrity likeness rights, the rest can be reviewed in one batch."

This is the complete transfer of attention:

Level 2Level 3
The human remembers that work needs doingThe agent remembers that work needs doing
The human enters the system to find tasksThe agent pushes what matters to the human
The human maintains the to-do queueThe agent maintains the to-do queue
The human decides production volumeThe agent decides from the responsibility gap
The human pushes the process step by stepThe agent advances autonomously to the review point
AI offers candidate materialThe agent delivers acceptance-ready results
The human tracks open itemsThe agent closes open items

What level 3 delivers: a portfolio of social accounts that runs steadily with minimal human involvement. (What level 4 looks like: see 1.6.)

1.3 Is level 3 just level 2 plus a scheduler

Structurally, yes — level 3 can be built on top of level 2. It reduces to:

Level 3=Level 2 business capability+Durable state+Timer/event wake-up+Responsibility-gap detection+Autonomous task creation+Dynamic workflows+Result acceptance+Continuous learning
Level 2 already provides the hands

Trend tools, topic tools, copy tools, image tools, publishing tools, data tools.

Level 3 adds the job-holder's brain

What is my responsibility right now; what is the current state; how far to the target; should I act; what tasks should I create; which capabilities should I call; when should I ask a human; when should I stop.

If the system merely "generates three posts at nine o'clock every day", it is still scheduled automation. A digital employee is:

At nine o'clock, inspect the responsibility domain; based on current inventory, schedule, trends, approvals, and targets, decide whether to produce, what to produce, how much, and where to stop.

So the timer only wakes the agent up. The real capability of a digital employee is:

Waking up knowing why it works, what to do, and when to stand down.

1.4 Does a Retainer agent need to run nonstop

No. The right concept is not "continuous execution" but:

Continuous responsibility, intermittent execution.

A retainer agent can wake on:

ON A SCHEDULE
  • Daily content-inventory checks
  • Weekly prep for next week
  • Monthly roadmap generation
  • Monthly retrospectives
  • Quarterly strategy reviews
ON EVENTS
  • New client feedback arrives
  • New assets are uploaded
  • Content gets rejected
  • A publish fails
  • A new trend emerges
  • A competitor makes a major move
  • Data goes anomalous
  • A sensitive comment appears
ON STATE
  • Inventory falls below the safety line
  • A task nears its deadline
  • An approval waits too long
  • A content type keeps failing
  • An experiment has enough data
  • Monthly metrics drift off target

On waking, the agent can reach one of three conclusions: it must act now; it must wait for a condition; or there is nothing worth doing.

A mature digital employee must be able to judge that "now is not the time to keep working." It should never manufacture content and tasks just to look busy.

1.5 How a Retainer agent keeps learning

Start with concrete things worth learning, then sort them into layers. The system should remember, for instance: this client dislikes loud emoji; "premium" means more restraint, not more complexity; a certain product-exposure ratio always gets bounced; the client wants two complete directions per review — that is client-preference learning. Layered, a level 3 retainer digital employee needs at least four kinds:

State learning

Remember where the work stands, what is waiting, what needs resuming, what comes next

Client-preference learning

The examples above: emoji, what "premium" means, exposure ratios, two directions per review

Skill learning

Turn experience into skills: how to handle vague feedback, how to rewrite for each platform, how to spot authenticity risks, how to judge whether a trend is worth chasing, how to keep moving when assets are missing

Outcome-strategy learning

Let results change: topic weights, content mix, publishing cadence, creative direction, platform roles, experiment priorities

Level 3 requires at least the first three; level 4 demands the fourth.

1.6 Level 4: the Retainer value-discovery engine

What the brand originally asked for may be nothing more than: keep up the social presence, publish on time, support campaigns, avoid PR incidents, send a monthly report.

LEVEL 3 OWNS

Getting the agreed account operations done, reliably.

LEVEL 4 OWNS

Discovering and realizing value in the account portfolio the brand never knew was there.

The agent doesn't just maintain the brand's social presence — it keeps finding and realizing more value across the account portfolio. It proactively spots content directions, user sentiments, and brand narratives with breakout potential; grows one high-performing piece into a recurring series, a content IP, and a stable brand persona; turns virality from a lucky outcome into a capability that can be learned, reused, and managed; and ultimately produces verifiable lift in brand awareness, search, user preference, and sales.

The system keeps discovering
  • Which content has genuine breakout potential
  • Which user sentiment is worth owning long-term
  • What distinctive social persona the brand can form
  • Which comments can grow into recurring segments
  • Which success can be replicated into a series
  • What role each platform should play
  • Which content drives search and purchase, not just engagement
  • Which routine content has no value and should stop
  • Which content can become a long-term brand asset
The eventual results include
  • A rising share of high-performing content
  • A markedly higher hit rate
  • Users following and discussing unprompted
  • A stable brand social persona
  • Long-running segments and content IP
  • The accounts reinforcing one another as a matrix
  • Growth in brand search, preference, and discussion
  • Verifiable incremental impact on sales

"Make every post go viral" works as an internal ambition, not as a per-post contractual promise. The more accurate definition of level 4 is:

The system continuously discovers, validates, and reuses high-value content patterns, gradually turning virality and brand lift from chance events into a manageable capability.

1.7 What justifies $20K–200K a year

Level 2 enterprise SaaS can also reach $20K a year, especially bundled with customization, deployment, data integration, and services. But to charge on the logic of a "digital employee" rather than "expensive SaaS," you must prove three things:

1 · Attention actually transferred
  • The PM no longer enters the system to find work
  • The agent creates the necessary tasks
  • The agent pushes items for review
  • Open items no longer depend on human memory
2 · Headcount or capacity actually changed
  • Hours per brand per month drop visibly
  • One PM supervises more brands
  • Copy and design lose their repetitive load
  • Project knowledge no longer lives in one person
  • The same team serves more brands
3 · Results stable and acceptable
  • On-time delivery rate
  • Unedited-into-review rate
  • Rejection / rework rate
  • Human-takeover rate
  • High-risk error rate
  • Client satisfaction
  • Project gross margin

A suggested core dashboard for a digital employee:

MetricWhat it measures
Share of tasks created by humansWhether work still relies on someone remembering it
Share of agent-created tasks that prove usefulWhether initiative creates real value
Human management minutes per brand per weekWhether the attention burden is falling
Times humans enter the system unpromptedWhether the product is pull or push
Auto-closure rate of open itemsWhether the agent pushes work through to the end
Average autonomous steps before reviewHow long a responsibility chain the agent carries
Unedited-into-review rateHow close output is to deliverable
Rejection / rework rateQuality stability
Repeat-error rateWhether it truly keeps learning
Brands supervised per PMWhether organizational capacity changed
Gross margin per brandWhether financial value materialized

The pricing logic in one sentence:

No client pays $20K–200K a year because you have seven agent modules. They pay because they no longer need to schedule, remember, and push this work every day.

1.8 Team and resources at each stage

Core team size grows with product level
People — each bar spans the min–max core team size per stage
16128402–3Stage 1Generation tools4–6Stage 2Copilot7–10Stage 3Digital employee10–15Stage 4Value-discovery engine
STAGE 1 · GENERATION TOOLS (2–3 PEOPLE)

Roles: one AI-native full-stack engineer; one advertising-business or content product lead; design support as needed.

Goals: validate copy and image quality; find high-frequency point tasks; get a usable prototype fast.

STAGE 2 · RETAINER COPILOT (4–6 CORE MEMBERS)

Roles: one product lead who deeply knows retainer work; one senior full-stack or tech lead; one to two AI application engineers; one knowledge-and-evals engineer; one product designer or agent-operations person; PMs, copywriters, and designers as a standing co-creation group.

Goals: build the brand knowledge base; connect the content workflow; cut generation and collaboration costs; measure true labor-hour changes precisely; find the responsibility loop best suited for a digital employee to take over.

STAGE 3 · RETAINER DIGITAL EMPLOYEE (7–10 CORE MEMBERS)

Roles: one AI product architect; one agent systems architect; two to three AI-native full-stack engineers; one state/memory/platform engineer; one evals-and-learning engineer; one to two agent-operations people; one senior retainer business lead (role duties and cost assumptions in Part III · the talent ladder).

Key capabilities: event-driven systems, durable state machines, durable execution, timer and event triggers, task queues, dynamic workflows, forward-looking tasks, idempotency/retries/rollback, human-in-the-loop approval, agent observability, memory and skill updates, production-grade evals.

The minimum responsibility loop: do not open by promising to "fully take over the brand's accounts." A more realistic first duty is — keep the next seven days stocked with enough content that meets the internal first-review bar, at all times. It naturally contains: daily inspection, task discovery, content production, QA, approval, restocking, and a stop condition.

STAGE 4 · RETAINER VALUE-DISCOVERY ENGINE (10–15 CORE MEMBERS)

Added on top of stage 3: one brand strategy lead; one senior creative lead; one to two marketing data scientists; one data engineer; one experimentation/growth engineer; one multimodal content-understanding engineer; more content-quality and agent-operations support.

  • Brand strategy lead: brand equity, user culture, platform context, content motifs, social persona, long-term brand value.
  • Senior creative lead: turns the signals the system finds into segments, series, content IP, creative motifs, and cross-platform matrices.
  • Marketing data scientists: content-to-sales attribution, incrementality experiments, causal analysis, content and audience clustering, portfolio optimization, separating vanity metrics from real value.

The scarcest person at level 4 is not the agent engineer, but:

someone who can read a data anomaly as a brand opportunity, and turn that opportunity into a long-term content asset.


Part II · The value levels of software products

The four retainer levels are not specific to that industry. Generalize "generation tool → copilot → digital employee → value-discovery engine" and you get the value levels of software products at large:

Level 1 sells features. Level 2 sells process efficiency. Level 3 sells role capacity and attention transfer. Level 4 sells growth and value discovery.

Product levelFormDivision of laborTypical pricingWhat it sells
Level 1 · Point AI toolsAny AI image / text / video tool, summarizers, simple chatbots, one-shot automation pluginsThe product delivers one result; the user holds all context, tasks, and responsibilityPriced like software tools, model calls, conveniencePoint features
Level 2 · Ordinary SaaS / copilotCentralized context, standardized process, AI generation, approval and analyticsHumans still remember tasks, create tasks, push the process, own the outcome$20–200 a month; complex enterprise editions higherSoftware capability and process efficiency
Level 3 · Digital employeeA software entity responsible for a defined, ongoing roleThe agent creates and advances tasks on its own; humans review the key judgments$2K–10K a month, benchmarked against role cost, capacity, and management attentionRole capacity, closed-loop responsibility, attention transfer
Level 4 · Value-discovery engineProactively finds better goals and new opportunitiesThe agent proposes and validates value hypotheses; the company reallocates resourcesBenchmarked against incremental revenue and margin; can take a share of the sales liftThe sustained ability to find and realize growth

The four levels also differ in their basic unit:

The unit of a tool is the feature. The unit of a workflow is the process. The unit of an agent is the task. The unit of a digital employee is the responsibility of a role.

Level 1: point AI tools

Typical products: copy generators, image generators, video tools, summarizers, simple chatbots, one-shot automation plugins. The product delivers a result, but the user still holds all context, tasks, and responsibility. Priced against software tools, model calls, and convenience.

Level 2: ordinary SaaS / copilot

The internal workflow products of leading companies sit roughly here — in retainer terms, the copilot form of 1.2.

The product can
  • Centralize context
  • Standardize the process
  • Generate with AI
  • Provide approval and analytics
  • Reduce tool switching
  • Raise the productivity of existing staff
But humans still
  • Remember the tasks
  • Create the tasks
  • Enter the system
  • Push the process
  • Track open items
  • Decide the next step
  • Own the final outcome

Typical pricing: $20–200 a month, more for complex enterprise editions. What this level sells: software capability and process efficiency.

Level 3: the digital employee

The basic definition of a digital employee:

A software entity that holds a defined, ongoing role responsibility.

It maintains

Long-lived working state, a work calendar, open items, forward-looking tasks, priorities, durable context, approval rules, stop conditions.

It can

Be woken by timers, business events, or state changes; judge whether a responsibility gap exists; create tasks autonomously; orchestrate workflows dynamically; call tools to execute; handle routine exceptions; push only what matters to humans; keep learning from feedback and outcomes; stand down once the responsibility is met.

A digital employee is not required to be busy all the time. It should be: continuously responsible, intermittently executing, waking on demand, sleeping when done.

A digital employee is not an agent that completes tasks. It is a software entity that permanently occupies a role, answers for long-term goals, and keeps finding its own next piece of work. Its core capability is not automated execution but taking over working attention. What is truly expensive is not the labor of generation — it is having someone keep the job on their mind.

Typical pricing: $2K–10K a month, benchmarked against role cost, capacity, and management attention (how to prove you deserve it: the three proofs in 1.7). What this level sells: sustained role capacity, closed-loop responsibility, attention transfer, changed staffing, delivery reliability.

Level 4: the value-discovery engine

A value-discovery engine carries the original responsibility and, beyond it, proactively finds: which goals are more worth pursuing, which work has no value, which opportunities the company has not seen, which experiments could create new revenue and profit.

Its pricing no longer benchmarks only software and labor. It can benchmark: incremental revenue, margin improvement, cost reduction, new business opportunities, avoided bad investments. Business models can include: base service fees, deployment fees, ongoing platform fees, revenue share on increments, profit-improvement share.

What this level sells: the sustained ability to discover and realize growth opportunities.


Part III · Enterprise AI-native maturity

Finally, the enterprise side: where does a whole company sit? First watch two departments across the four levels, then the full definitions of the four levels and the six dimensions that decide the answer.

3.1 Four levels through two departments

HR, level by level

Level 1

HR uses scattered AI tools — writing JDs, summarizing interviews, generating interview questions — but the capability depends on individuals and never accrues to the organization.

Level 2

An HR agent matches people to roles, screens résumés, builds the talent pool, assists interviews and funnel analytics — but the definition of talent still comes from business units' gut feel. The business says "we need a senior PM," and the agent helps HR find a senior PM faster.

Level 3

AI connects HR with sales, operations, delivery, and finance across the company, and redefines how talent demand is generated. It no longer waits for departments to file hiring requests; it continuously analyzes project margins, client reviews, delivery quality, employee performance, and workload to surface the real capability bottlenecks. It might conclude that growth is limited not by headcount but by senior PMs' scarce capacity being consumed by low-value coordination — and proactively propose a new structure of "senior PM + digital employees + junior staff," then keep driving hiring, training, role redesign, and agent substitution.

Level 4

AI redefines HR's value from first principles: HR is not hiring and managing people, but keeping the organization's capability supply matched to the demands of value creation. It judges which capabilities to hire, which to train, which to outsource, and which to hand to agents, continuously reallocating organizational resources so scarce talent always works on the most valuable things.

Marketing ops, level by level

Level 1

Ops uses scattered AI tools — writing copy, generating images, cutting video, summarizing data — but everyone still manages their own tasks and context.

Level 2

An ops agent maintains the prompt pool, mass-produces video and ad creative, feeds data back, analyzes and iterates automatically — making the existing production and delivery pipeline more efficient.

Level 3

AI no longer waits for operators to type in topics and requirements; it continuously owns the operating responsibility for an account, a channel, or a growth target. It watches inventory, schedules, trends, competitors, and performance; creates tasks, produces content, initiates reviews, tracks publishing and retrospectives on its own; and pushes only strategic questions and high-risk items to the operators. Humans stop reminding themselves what to produce — the agent tells them what has been advanced and what still needs their judgment.

Level 4

AI redefines the value of operations: not producing content continuously, but continuously discovering and occupying value gaps. It analyzes user sentiment, competitors' blind spots, platform shifts, and business results; finds new content motifs, audience segments, distribution scenarios, and growth opportunities; and validates them through experiments. Content production is merely the last step after a value judgment lands.

3.2 The four maturity levels defined

Level 1: the personal-tools company

Employees use general-purpose LLMs, AI coding tools, generation tools, and automation scripts on their own. AI capability is not yet an organizational asset: results depend on individual skill; prompts and experience leave with employees; data never flows back; workflows do not change; the financial contribution cannot be measured. At bottom it is still a traditional company in which some employees happen to use AI.

Level 2: the department-workflow company

Leading companies sit roughly here. The company can speed up each department's existing workflows with agents, producing cost, capacity, and financial improvements visible at group level — the level 2 of the HR and ops examples above. But the premise of this level is:

Do not redefine the department's value; only optimize what the department already does.

What level 2 takes, organizationally
  • One or two senior developers, or AI-native young developers
  • One PM able to model the department's workflows
  • A department head willing to share process and data
  • A degree of CEO will
  • Basic data access and permissions
How level 2 is measured

Task time, manual steps removed, adoption rate of generated content, department throughput, per-department labor cost, unit output cost, financial impact on revenue and expenses.

What this level sells: doing existing things faster and cheaper.

Level 3: the process-rebuilt AI-native company

Level 3 stops treating existing departments and processes as immovable premises. It redefines from scratch: how information flows, how work gets created, how departments collaborate, how humans and agents divide labor, which processes stay, which processes go — the level where, in the examples above, HR discovers "the growth ceiling is senior-PM scarcity" and ops proactively owns account responsibility.

Break the department walls; build enterprise-level context. The reason those level 3 examples are impossible at level 2 is that a traditional company's information is scattered across nine domains: sales, operations, product, R&D, production and delivery, support, finance, HR, management. Each department sees only local facts, so it can only optimize locally. The first thing level 3 builds is not more agents but a unified enterprise context, so the system can continuously understand:

  • Why customers buy or churn
  • Which projects earn revenue but lose money
  • Which people and capabilities actually drive delivery
  • Which problems keep recurring across departments
  • Which internal goals conflict
  • Which link constrains the whole company's growth
  • Which business assumptions have expired

This is fundamentally a question of CEO resolve and CTO field of view. The CTO and the AI product lead must: ① observe every department's real workflows; ② map the information and efficiency gaps; ③ redefine the data structures; ④ work with the CEO to break departmental isolation; ⑤ turn information assets into durable context agents can use.

Continuous responsibility and the autonomous task loop. A level 3 system no longer depends on people entering software to create tasks. One full round of its work (isomorphic to the retainer employee's 11-step loop):

Continuously maintains goals and working state

Agent

Wakes on schedule or on business events

Agent

Judges the gap between current state and the responsibility target

Agent

Creates and prioritizes tasks autonomously

Agent

Orchestrates execution dynamically

Agent

Calls tools to complete the work

Agent

Asks humans to handle high-risk exceptions

Agent → human

Accepts and verifies results

Agent

Updates state and experience

Agent

Stops when the responsibility is met, until the next wake-up

Agent

The human appears only at step 7, as the handler of high-risk exceptions; the dashed line means each round returns to step 1.

So the core change of level 3 is:

Working attention transfers from humans to the system.

Continuous, autonomous learning. A level 3 agent must keep learning from: human edits, content acceptance and rejection, project wins and failures, client satisfaction, business results, execution anomalies, and its own errors. What it learns must be written into: long-term memory, business rules, prompts, eval sets, skills, model routing, dynamic workflows, decision policies. Early on the model need not modify its own weights — but one thing must hold:

The same mistake is never repeated forever; the same success can be reproduced reliably.

The talent ladder level 3 requires:

AI product architect / process-modeling lead

Observes workflows across the company, defines responsibility boundaries, builds the business state model, sets the human–agent division of labor, defines what "closed loop" and "deliverable" mean, aligns the CXOs. This role cannot be a collector of feature requests.

Agent engineer≈ $140K–200K a year

Scope: agent orchestration, memory systems, tool calling, RAG, long-task state, dynamic planning, human-in-the-loop, failure recovery, quality control, the learning loop. Needs to build three layers: ① grounded reasoning; ② quality checks; ③ experience learning.

The goal: a production-grade, 95-point closed loop that finishes the job itself — not a thin wrapper over a model API. One of the hardest roles to hire, and one that most determines the product's moat.

Prompt / algorithm / evals engineer≈ $140K–200K a year

Scope: prompt engineering, model selection and routing, business scoring rubrics, offline evals, online quality monitoring, error ledgers, regression eval sets, failure attribution, training and fine-tuning.

The metrics that matter most: loop-closure rate, unedited-into-review rate, unedited-direct-publish rate (only where contract and risk allow), rejection/rework rate, human-takeover rate, cross-version regression rate.

Evals remain a scarce, dedicated craft in today's market — and a digital employee's price rests precisely on proving these numbers.

Senior AI-native developer≈ $100K a year

Scope: data-ingestion pipelines, structuring the company's information assets, backend, frontend product engineering, permissions, auditing, infrastructure, enterprise integrations, production stability. The value is shipping the complete product from data to systems to user experience — and data ingestion and structuring is usually the single largest block of work.

Young AI-native developers / interns≈ $20K–60K a year

Scope: prototype exploration, tool integration, testing, data labeling, routine pipeline runs, eval-sample preparation, new-model experiments, edge automation. They can carry high-frequency experiments and the grunt work, but should not own evals, and should not own production architecture alone.

Agent operations

Scope: catching false and missed wake-ups, catching junk tasks, analyzing failure traces, cleaning contaminated memory, maintaining client-specific rules, updating skills and eval sets, and watching whether the agent truly reduces human attention.

At the current state of the art, "continuous learning" does not complete itself. Agent operations is the training, management, and performance department for digital employees.

Organizational preconditions: strong CEO backing; a CTO with a cross-department field of view; accessible information assets; walls that can actually be broken; solid permissions and auditing; a real business as a continuous training ground; a management team willing to redesign roles and processes.

Level 4: the value-discovering AI-native company

LEVEL 3 ANSWERS

How does the company keep running in a better way?

LEVEL 4 ANSWERS

What value should the company create — and what value has no one discovered yet?

Level 4 does not merely hand tasks to digital employees — it redefines, from first principles, why each department exists: the level where, in the examples above, HR becomes a "capability allocation system" and ops becomes a "market-cognition, value-validation, opportunity-occupation system".

The level 4 enterprise loop:

Continuously ingest signals from inside and outside the company

Discover value gaps

Form business hypotheses

Design low-cost experiments

Organize execution

Judge the true increment

Scale the successful patterns

Write the experience back into the company's cognition system

Seek the next, higher-value target

The dashed line means each round returns to step 1 — a standing value-discovery loop.

Preconditions: the level 3 talent ladder is in place; company-wide information assets are connected; department walls are down; the CEO is willing to re-examine the existing business; the CTO tracks the moving frontier of AI capability; agents are allowed to question existing tasks; agents are allowed to drive experiments and resource reallocation; performance shifts from "how many tasks completed" to "how much value created."

3.3 Six dimensions that decide your level

Strip away all the examples, and what actually determines where a company (or a product) sits are six things:

DeterminantThe question it asks
1 · How far context reachesIndividual, department, whole company, or inside-plus-outside
2 · How far results close the loopDelivering assets, assisting processes, owning a role, or creating business increments
3 · Who holds the working attentionWhether humans still remember tasks, enter systems, and push processes
4 · Whether the system keeps learningWhether feedback, failure, and business outcomes change future behavior
5 · How much permission to act it holdsWhether it can read and act across departments and systems
6 · Whether process and value get redefinedSpeeding up old processes, or re-deciding what work is worth doing

The progression of the six dimensions across the four levels, as one matrix:

DimensionLevel 1Level 2Level 3Level 4
Context reachWhat the user pastes inOne department's knowledge, process, historyCompany-wide information assetsInside + continuous absorption of the outside
Result closureDelivers single assetsDelivers process efficiencyDelivers role responsibilityDelivers business increments
Attention ownershipHumans hold all context, tasks, responsibilityHumans remember what to ask AIAI remembers, and informs humansAI finds more worthwhile goals itself
Continuous learningNo accumulation; experience leaves with peopleIn-department data feedback and asset iterationState + preference + skill learningPlus outcome-strategy learning
Permissions & orgPersonal tools, no enterprise accessIn-department data accessCross-department permissions; governance changes with itCompany-wide assets connected; agents may question tasks
Process & value definitionProcesses unchangedOptimize what exists; value undisturbedRedefine processes and the human–agent splitRedefine what work is worth doing

Dimension 1: context reach

Level 1: AI has only what the user pastes in. Level 2: AI reads one department's knowledge, processes, and history. Level 3: AI connects the information assets of sales, operations, product, delivery, R&D, support, finance, and HR. Level 4: AI understands not just the inside of the company but continuously absorbs the market environment, competitors, user behavior, technology shifts, culture, supply chains, and industry change.

Dimension 2: result closure and pricing class

What decides a product's level is not how smart the model is, but:

What result it actually delivers, and how far that result closes the loop.

LevelDeliverableExample
Level 1A single assetGenerate one post, one image, one summary
Level 2Process efficiencyHelp the PM finish topics, copy, design, reports faster
Level 3Role responsibilityKeep the review queue stocked for the next seven days, continuously
Level 4Business incrementsTurn an ordinary set of brand accounts into a content asset that compounds brand awareness and sales

Commercial value can be understood as:

Chargeable valueScope of responsibility×Loop-closure rate×Quality stability×Degree of attention transfer×Business impact

Dimension 3: who holds the working attention

This is the decisive boundary between levels 2 and 3 — the abstract version of the second and third mornings that opened this essay.

Level 2: attention lives in the human

Humans still must:

  • Remember what today requires
  • Enter the system unprompted
  • Create tasks
  • Order the tasks
  • Track what is unfinished
  • Decide when to continue

AI saves execution labor, but the open loops still live in a human head.

Level 3: attention lives in the system

The digital employee owns:

  • The work calendar
  • The to-do queue
  • Content inventory
  • Open items
  • Next actions
  • Wake-up timing
  • Task priority

Humans stop hunting for work — the system informs them.

Level 2 is humans entering a system to find tasks. Level 3 is the system pushing the judgment calls to the human. The true mark of attention transfer is not how much content AI generates — it is that no human needs to remember tasks, maintain to-dos, track open items, or decide when to continue.

Which means AI has taken over not just the labor, but the attention of management.

Dimension 4: the capacity to keep learning

Continuous learning is not just online training of model weights. An enterprise agent's learning has at least five layers (the retainer agent's four kinds in 1.5 are the concrete form of the first four):

State continuity

Remember where the work stands

Preference learning

Remember how clients and managers edit the output

Skill learning

Turn successful experience into rules, skills, and workflows

Strategy learning

Let business outcomes change future resource allocation and behavior

Model-level continual learning

Update model weights directly without forgetting old capabilities

An enterprise digital employee needs at least the first three; the fourth underpins the value-discovery engine. The fifth remains frontier research on base models — not a precondition for the product to work.

Dimension 5: permissions and org structure

What an agent can create is bounded by the information and permissions it holds. An HR agent that can only read résumés can only optimize screening. Give it access to performance data, project margins, client reviews, delivery quality, collaboration graphs, and future business plans — and it may redefine what the company actually lacks.

That is why moving from level 2 to level 3 is not merely a technical upgrade but a joint shift in data governance, permission governance, department boundaries, incentive design, and decision rights.

Dimension 6: observe, distill, define

Genuine AI-native product definition is not collecting feature requests from business units. It is three things:

Observe

Sit beside PMs, copywriters, designers, HR, sales — watch the real work:

  • Where information comes from
  • How people make judgments
  • Where context keeps switching
  • Which work is mere habit
  • Which problems actually move profit
Distill

Distill messy labor into:

  • Goals
  • State
  • Constraints
  • Judgment criteria
  • Exception types
  • Responsibility boundaries
  • Verifiable results
Define

Redefine:

  • Which tasks deserve to stay
  • Which processes should disappear
  • Which responsibilities go to agents
  • Where humans must step in
  • What the product finally delivers

Final summary

The four levels of the AI-native enterprise, the four levels of software products, and the retainer agent's levels compress into one table:

LevelAI-native enterpriseWhat the product sellsThe Retainer agent
Level 1Employees use AI toolsFeaturesEveryone clicks away by hand
Level 2AI helps departments do existing work fasterProcess and efficiencyThe human decides what to produce, then works the system
Level 3AI takes over role responsibility, working attention, and open items — redesigning how work happensClosed-loop responsibility, role capacity, attention transferThe agent decides and produces daily, then calls the human to review
Level 4AI re-judges from first principles what work is worth doing, and discovers and creates greater valueGrowth, profit, and value discoveryThe agent runs the accounts steadily and keeps realizing brand and sales value far beyond social presence

So the most accurate definition of a digital employee:

A digital employee is not an agent that completes tasks, but a software entity that permanently occupies a role, holds the working attention, answers for long-term responsibility, and keeps learning from experience.

And a value-discovery engine:

It doesn't just keep doing the job — it keeps discovering how much more the job was worth all along.


Appendix A · How the market and frontier labs define the digital employee

A.1 There is no unified definition

The previous generation of automation vendors defined digital workers as software robots that execute enterprise processes end to end. Automation Anywhere emphasizes a "virtual workforce" of AI, machine learning, RPA, and analytics completing sequences of business tasks; UiPath now describes a single AI agent as an autonomous digital worker that completes a specific task from start to finish. (Automation Anywhere)

Salesforce uses the concept of "digital labor," stressing agents with defined roles, enterprise knowledge, the ability to act, and around-the-clock availability; Microsoft stresses managing agents like employees: define the role, limit the permissions, supervise continuously. (Salesforce)

So today's de facto market floor is roughly:

Anything that autonomously completes a set of tasks a role used to own can be packaged as a digital employee.

This essay uses a stricter definition:

A digital employee is a software entity accountable for a defined, ongoing role responsibility. It maintains working state and open items; proactively inspects its responsibility domain on schedules and business events; discovers and creates its own tasks; orchestrates workflows dynamically; pushes only necessary judgments and final results to humans; and updates its memory, skills, and policies from feedback and business outcomes.

The heart of it is not "a model running 24 hours a day," but:

The responsibility persists, and the system can resume its own work whenever needed.

A.2 What the frontier labs are building toward

Frontier agent research is no longer chasing single-shot reasoning alone; it is filling in the capabilities a long-lived working entity requires:

1 · Long-horizon autonomy

Can an agent push one goal across hours, sessions, even days? OpenAI has productized long-horizon tasks, scheduled resumption, and multi-day continuation; Anthropic treats cross-context-window state handoff, long-task recovery, and harness design as the key problems of long-horizon agents. (OpenAI Developers)

2 · Durable state and execution memory

An agent needs to know more than "what happened": where the work stands, what the current plan is, which paths already failed, why it paused, what it is waiting for, what remains open. Microsoft's STATE-Bench now evaluates whether agents improve later enterprise tasks with past experience; work like PlugMem turns raw interaction into reusable long-term knowledge. (Microsoft)

3 · Proactive wake-up and forward-looking work

Waking on schedules; waking on business events; resuming paused work when conditions are met; carrying responsibility without a direct prompt. OpenAI's Workspace Agents and Codex Automations already productize scheduled execution, repeated runs, retained context, and future automatic resumption. (OpenAI)

4 · Learning from experience

Google's ReasoningBank distills successes and failures into reusable reasoning strategies; Titans, MIRAS, and Nested Learning explore longer-horizon memory updates and continual learning. (Google Research)

5 · Harness, evals, and reliability

The frontier labs increasingly recognize that capability comes not only from the base model but from everything around it: context organization, task state, tool environments, memory systems, execution logs, eval sets, failure recovery, human approval. Anthropic emphasizes that harness design significantly affects an agent's ability to advance work across sessions; OpenAI treats durable state and long-task harnesses as core agent engineering. (Anthropic)

Which is to say:

A digital employee is not a bigger prompt, nor a model that never stops running. It is a production system composed of model capability, durable state, wake-up systems, dynamic planning, tool execution, evals, and continuous learning.


Appendix B · From discovery to operations: how the product gets built

B.1 Finding the need

Need-finding requires the CEO, the CMO, or someone with deep market experience. The goal is not collecting feature suggestions but understanding: who hurts most; whether the problem is frequent; how it is solved today; why existing solutions fail; who owns the budget; who bears the risk of failure; how much the problem costs in profit; and why AI can solve it now.

The final deliverable is not a feature list but: a set of testable value hypotheses.

B.2 Defining the product

Product definition takes the CTO, the AI product lead, AI-native engineers, and business experts together. The core method: observe → distill → define (the same method as dimension 6, here as the working version for product definition).

Observe
  • Read the company's information assets
  • Sit beside different roles and watch
  • Run targeted interviews
  • Run focus groups
  • Debrief wins and failures
  • Log context switches and rework
Distill

Distill the real work into:

  • Responsibility
  • Inputs
  • State
  • Constraints
  • Judgments
  • Exceptions
  • Outputs
  • Acceptance criteria
Define
  • What the product is responsible for
  • What humans are responsible for
  • Which processes get rebuilt
  • Which processes get cancelled
  • When the agent may act autonomously
  • When a human must decide
  • Whether this is SaaS, a digital employee, or a value-discovery engine

This stage requires every CXO to converge on which level the product is.

B.3 Design and engineering

A senior product-engineering team covering four blocks:

Data & context

Data ingestion, structuring documents and chats, permissions, knowledge bases, project state, long-term memory.

The agent system

Triggers, planner, task queues, tool calling, dynamic workflows, approvals, failure recovery, evaluator, stop conditions.

Product & infrastructure

Frontend, backend, enterprise integrations, auditing, observability, cost control, multi-tenancy, data isolation.

Evals

Quality rubrics, golden datasets, error ledgers, regression tests, online sampling, comparisons across models and workflows.

B.4 Running it in production

An AI-native product needs continuous operations after launch — and operations here is not sales plus support. It is running:

Loop closure×Result score×Business impact

The core work: measure the unedited-into-review rate, the unedited-direct-publish rate, the rejection/rework rate; analyze failure traces; keep an error ledger; fold repeat errors into the regression set; update brand memory, rules, and skills; study why humans take over; prove value with real project financials.

Early go-to-market should lean on hard numbers from lighthouse projects, not agent demos: how many labor hours dropped, how many management minutes per brand dropped, how many more brands one team can serve, how much gross margin improved, whether delivery quality held.