Skip to main content
Digital FTE (AI Employee) Development

Hire an AI employee that clocks in and never clocks out 

A Digital FTE is not a chatbot. It is a scoped role that reads your inboxes, your ERP, and your ticket queue, does the repeatable part of a job end to end, and escalates the rest with its reasoning attached. Built spec-first, so you approve what it will and won't do before it touches anything.

Shift coverage, no rota
24/7

Shift coverage, no rota

Years building production systems
3+

Years building production systems

Projects shipped end to end
12+

Projects shipped end to end

A Digital FTE at a console, running Gmail, LinkedIn, Instagram, Facebook, Odoo, and X in parallel around the clock
One role · scoped, permissioned, logged
Full Digital FTE dashboard showing vault status, post pipeline, and system health for every connected service

Built on

  • OpenAI Agents SDK
  • LangGraph
  • General Agents
  • n8n
  • Model Context Protocol
  • TypeScript
  • Python
  • Postgres + pgvector
An operations team standing over a connected map of the systems their work passes between

Where the payroll hours actually go

Your team is not slow. The handoffs are. 

Every operations team I have looked at loses the same hours in the same four places, and none of them are the part of the job you hired that person for.

01

The work stalls between systems, not inside them

CRM, ERP, inbox, and spreadsheet each do their job fine. A person is the glue: copying a value, checking a policy, chasing a status. That cost never shows up on any tool's invoice.

02

Rule-based automation dies on the first exception

A recorded script breaks the moment a field moves or a supplier sends a different format, so someone ends up reviewing the automation as well as doing the work.

03

A chatbot answers; it does not finish the job

Explaining the refund policy is not issuing the refund. Deflection numbers look good while the queue length stays exactly where it was.

04

Nobody signs off on work they cannot audit

The blocker on most agent projects is not accuracy. It is accountability. Without scoped permissions, a log, and a human checkpoint on the risky steps, the rollout stops at the security review.

What a Digital FTE actually is

A role you scope and staff, not a tool you have to drive 

The unit of delivery here is a job description, not a feature list. We write down the role, the systems it may touch, the decisions it may make alone, and the ones that stop for a human. Then it runs that role on your infrastructure.

A chatbot

Answers a question and hands the work back to a person.

Never writes to your systems of record.

An RPA script

Replays a fixed sequence of clicks against a fixed screen.

Cannot reason about an exception it has not seen before.

A Digital FTE

Owns a scoped role: reads the situation, decides, acts through permissioned tools, and reports what it did.

Still escalates anything outside its written boundary.

Inside one shift

Observe, decide, act, report, then round again 

An agent is not one call to a model. It is a loop: it reads the whole situation, writes down the next move and why, executes against your real systems, then records what actually happened so the next pass is better informed.

An operations control room monitoring autonomous agent activity
One loop · four stages · every pass logged

Stage 01

Observe

Read the whole situation

Pulls the record from every system that matters: the invoice, the PO, the contract terms, the prior tickets. It does not reason from a single message.

Stage 02

Decide

Write down the next move

Chooses the next action and states the reason in plain language. That reasoning is stored, so a disputed decision can be read back rather than guessed at.

Stage 03

Act

Call the real tools

Executes through a permissioned tool layer wired into your ERP, CRM, ticketing, or inbox. Anything above the threshold you set stops at an approval gate with a named owner.

Stage 04

Report

Close the loop out loud

Logs the outcome, updates the record, and files the exception queue, so the morning briefing is a summary of work done rather than a request for instructions.

Agent trace · shape of one pass
  • obsinvoice INV-88301 · PO-4471 · goods receipt GR-9920
  • thinkline 3 over-billed by 4.2% → outside 2% tolerance
  • toolerp.hold_payment(inv_88301) → ok
  • gate>$25k requires approval → queued to finance lead
  • logexception filed with full reasoning attached

The loop exits on one of two conditions: the goal is met, or the agent hits a boundary you defined and escalates to a named human with its full reasoning attached. It does not guess its way past a wall.

How it is actually engineered

Loop, harness, and graph engineering: the three layers that decide reliability 

Picking a good model is the easy part, and it is not where agents fail. What separates a demo from something an operations lead will sign off on is the engineering around the model: how often it is allowed to think again, what it is handed and what it may touch, and whether its control flow is written down or improvised. These are the three layers I build every Digital FTE on.

A branching agent graph showing steps, conditions, and state moving between them
Nodes · conditions · checkpointed state
Diagram of a closed agent loop cycling through learn, observe, optimize, and perform
Layer 01

Loop Engineering

When it thinks again, and when it stops

The loop is the agent's heartbeat: observe, decide, act, then re-read the result and decide whether it is done. The engineering is in the exits, not the iterations: a step budget, a token and cost ceiling, a convergence check so it does not re-run the same call with the same input, idempotency keys so a retry cannot double-post an invoice, and a hard stop that escalates instead of guessing.

  • Step, token, and cost ceilings enforced per run
  • Idempotency keys on every write so retries are safe
  • Repeat-call detection instead of a silent infinite loop
  • Termination on goal met, boundary hit, or budget spent, never on luck
Diagram contrasting prompt engineering, context engineering, and harness engineering
Layer 02

Harness Engineering

What surrounds the model call

The harness is everything that is not the model: which records get assembled into context, what the tool schemas look like, which credentials the call runs under, how the output is validated before it is trusted, and what gets written to the trace. Most production failures I have debugged were harness failures (a stale record, a loose tool contract, an unvalidated field) rather than reasoning failures.

  • Context assembled from systems of record, not pasted history
  • Typed tool schemas with validation on the way in and the way out
  • Scoped credentials per agent, checked at the tool layer
  • Prompt-injection filtering on anything arriving from outside the org
  • Every prompt, call, and outcome traced and replayable
Agent graph with plan, code, and retrieve nodes feeding a deterministic validation step
Layer 03

Graph Engineering

Control flow you can read on a page

Past a certain complexity, a single loop stops being honest about what the agent does. The work becomes a graph: nodes are steps, edges are conditions, and the state is checkpointed at each hop, so a run can pause at an approval gate on Friday and resume on Monday exactly where it stopped. It also makes review possible: the diagram is the behaviour, and a failed node has a deterministic fallback rather than a retry loop.

  • LangGraph state machines instead of one open-ended prompt
  • Checkpointed state, so runs survive a restart or an approval wait
  • Human approval gates modelled as nodes, not bolted on afterwards
  • Deterministic fallback branch on every step that can fail
  • Supervisor plus specialists when one role needs several skills

The three are one system. A tight loop with no harness burns budget quickly and confidently. A strong harness with no graph works right up to the first exception that needs two steps and a human. The reason a Digital FTE can be handed to your team is that all three are written down before the build starts.

Roles you can staff

Seven jobs a Digital FTE can hold down 

Each of these is scoped as a role with a written boundary, not a feature bolted onto an existing tool. Most engagements start with one and add the next once it is running unattended.

An executive reviewing an automatically generated operations briefing
Role 01

Operations analyst

The Monday-morning role. It reconciles the week across your systems, audits what moved against what should have moved, and lands a briefing in your inbox before anyone opens a laptop, with the exceptions ranked and the reasoning attached.

  • Cross-system reconciliation on a schedule
  • Ranked exception queue instead of a raw report
  • CEO briefing generated from live records, not a template
A finance team reviewing invoice and purchase order documents
Role 02

Finance & invoice clerk

Three-way matching between invoice, purchase order, and goods receipt, with the clean cases going straight through and only genuine discrepancies reaching a person. Approvals above your threshold still stop for a named owner.

  • Invoice, PO, and receipt matched automatically
  • Discrepancies routed with their evidence
  • Threshold-based human approval, never a silent auto-approve
An AI assistant handling multi-channel business conversations
Role 03

Inbox & comms handler

Watches Gmail, WhatsApp, and the shared inboxes; classifies what arrives, drafts the reply in your voice against your own policies, files the attachment where it belongs, and flags anything that needs a human before it goes out.

  • Multi-channel monitoring (Gmail, WhatsApp, shared inboxes)
  • Drafts grounded in your own policy documents
  • Send-gate on anything customer-facing you mark sensitive
A support specialist working through a queue of customer requests
Role 04

Customer support agent

Grounded in your help centre and past tickets, so it resolves the routine request end to end by issuing the change, not just explaining it, and escalates the rest with the full context already attached.

  • Retrieval over your real help centre and ticket history
  • Executes the change through scoped API permissions
  • Escalations arrive with context, not a transcript
An operator supervising automated data flows between business systems
Role 05

Back-office data operator

The copy-paste role nobody wants: pulling values between Oracle, Odoo, spreadsheets, and your CRM, validating them against the source, and keeping both sides in step without a nightly export ritual.

  • ERP and CRM write-backs with idempotent retries
  • Validation against the system of record before write
  • Sync failures surfaced as alerts, never swallowed
A developer building an automated task routing pipeline
Role 06

Task & project coordinator

Ingests incoming work, validates it, and routes it to the right owner with a deadline. The intake-to-assignment pipeline runs as an event stream instead of a person triaging a queue.

  • Event-driven intake with automatic validation
  • Routing rules you can read and change
  • Auto-generated task lists and follow-up chasing
Enterprise infrastructure running a team of coordinated AI agents
Role 07

Multi-agent operations team

When one role is not enough: a supervisor that decomposes the job, specialists that own each step, and shared memory so context survives the handoff. It is the same structure you would use if you were hiring three people instead of one.

  • Supervisor plus specialist topology
  • Persistent memory across steps and sessions
  • Deterministic fallback when a specialist fails

Not sure which role should go autonomous first?

Bring the process you want off your team's desk. You leave the call knowing what it would cost to run, what it would take to build, and whether an agent is even the right answer.

A human and a robotic hand meeting on a working agreement

Built for what actually gets approved

The parts nobody demos, that decide whether it ships 

Model quality is rarely what stops an agent reaching production. These are the layers that get it through a security review.

  • Tool and API calling

    Typed, permissioned calls into your CRM, ERP, database, and third-party tools. Never screen scraping.

  • Human-in-the-loop gates

    Manual review checkpoints built into the high-stakes steps, with a named owner rather than a shared inbox.

  • Role-based access control

    Each agent reaches only the data and actions it was explicitly authorised for, even if it is asked otherwise.

  • Audit trail and logging

    A replayable record of every decision, tool call, prompt, and outcome, retained on your side.

  • Persistent memory

    State and context that survive across runs, so an agent picks up a long-running task where it left off.

  • Fallback behaviour

    When it hits a case it was not designed for, it stops and routes to a person with the trace attached. It does not improvise.

  • Local-first deployment

    Cloud, on-premise, or hybrid, including open-weight models when the data genuinely cannot leave your network.

  • Evaluation before rollout

    Every change is measured against a fixed test set built from your real cases before it reaches production.

Headcount maths

The honest comparison, including where a person still wins 

A Digital FTE is not a replacement for judgement, relationships, or accountability. It is a replacement for the repeatable half of a role, and it is worth being precise about which half.

 Human hireDigital FTE
Hours9–5, weekends off24 / 7
Ramp-upWeeks of onboardingDeployed with your SOPs
CostSalary + benefits + overheadFixed subscription, no overhead
ScalingHire againClone the agent
ConsistencyVaries by daySame output, every run
SupervisionNeeds active managementSelf-operating, exception-only

Keep this with a person

  • Judgement calls with no precedent to reason from
  • Relationships, negotiation, and anything reputational
  • Owning the outcome when something goes badly wrong

Hand this to the agent

  • The same decision made a hundred times, identically
  • Work that arrives at 2am, on a Sunday, in peak season
  • Cross-system checking nobody has ever enjoyed doing

Industries

Where the first Digital FTE usually lands 

The pattern repeats across sectors: high-volume repeatable work, a system of record nobody wants to touch twice, and a person acting as the bridge between two tools.

  • E-commerce & retail ops
  • Logistics & supply chain
  • Fintech & finance ops
  • Insurance back office
  • Healthcare administration
  • Real estate
  • Legal ops
  • Professional services
  • Manufacturing
  • SaaS customer operations
  • Education administration
  • Agencies & studios
A warehouse operator checking stock against the system of record
Server infrastructure running the agent runtime behind the role

Tech stack

The runtime behind the role 

Chosen for boring reasons: everything here is something I have run in production and can hand over without a translation layer.

Agent runtime

OpenAI Agents SDK · LangGraph · General Agents · MCP

Reasoning models

GPT family · Claude · Gemini · Open-weight for private deploys

Memory & retrieval

PostgreSQL · pgvector · Qdrant · Redis

Integrations

Oracle · Odoo · Gmail / WhatsApp · REST & webhooks · n8n

Application

Next.js · TypeScript · Python · FastAPI

Operations

Docker · Vercel · GitHub Actions · Structured tracing

How the build runs

Seven steps from role audit to handover 

Every stage has an exit you can take if the numbers stop making sense. Nothing is built before you have approved what it is for.

Role audit

We map the job as it is really performed, including the shortcuts nobody documented, and score each step by volume, cost, and risk to find where an agent earns its keep.

Written spec

A job description for the agent: scope, systems, decisions it owns, decisions it escalates, and what it will refuse to do. You approve it before any code exists.

Evaluation set

A scored set of your real cases, agreed up front, so quality is a number that moves rather than an opinion argued about later.

Prototype on real data

A working agent runs against your sample data early, so you see its actual behaviour, and its actual failure modes, while changing course is still cheap.

Integration & guardrails

Scoped credentials, tool layer, approval gates, fallback paths, and tracing go in before the pilot, not after someone asks for them.

Supervised pilot

The agent runs beside your existing process on live traffic. You watch its decisions, its cost, and its edge cases with nothing at stake operationally.

Rollout & handover

Staged cutover with monitoring, cost ceilings, and a documented rollback. Then the codebase, prompts, and runbooks are handed to your team. You own all of it.

Control

Autonomy your security review will actually approve 

Autonomy with a hand on the brake

Every Digital FTE ships with the controls that make an operations lead comfortable signing off on it.

  • Explicit permission boundaries defined before deployment, not discovered after
  • High-stakes actions pause for human approval with the reasoning attached
  • Every prompt, tool call, and outcome logged and replayable
  • Scoped credentials per agent, so it gets only the access its role requires
  • Prompt-injection filtering on anything arriving from outside your org
  • Local-first and on-premise options when data cannot leave your network
  • Cost ceilings and rate limits so a runaway loop is capped, not invoiced
  • Your data stays in your environment and is never used to train models
Governance and access controls applied to a production AI system

Spec before code, every time

You approve a written role definition first, so the argument about what was promised never happens on delivery day.

Built for production, not for a demo

Integrated with your real systems from the first sprint. A prototype that only works on curated sample data is not a milestone.

You own the whole thing

Codebase, prompts, evaluation set, and infrastructure access hand over on delivery. No vendor lock, no black box.

I will tell you when it is the wrong answer

If a scheduled job, a better form, or a smaller script solves your problem, you hear that on the first call rather than after the invoice.

Engagements

Three ways to put one on the payroll 

Starter Agent

$2,500/project

  • Single workflow
  • Up to 2 integrations
  • Basic monitoring
  • 2-week delivery
Most common

Production Agent

$6,000/project

  • Multi-step workflows
  • Up to 5 integrations
  • Full monitoring & alerts
  • Human-in-the-loop approvals
  • 4-week delivery

Agent Retainer

$1,500/month

  • Ongoing operation
  • Continuous tuning
  • Priority support
  • Monthly reporting

Digital FTEs, answered

The questions that come up on the first call 

An employee that finishes the work. Not another pilot.

Bring one role. You get a written scope, an honest cost, and a straight answer on whether it should be an agent at all, before anyone commits to a build.