Capability 03 / 10Use cases, agents, governance

AI strategy and agents that move from pilot to daily use.

We rank your AI use cases by value, feasibility and risk, then build the first agents. Each agent is tested on your own examples, and a named person approves high-impact actions.

An illustrative mission board for “Your company”. Six missions sit in four columns: queued, agent working, needs approval and done. Each card shows the agent, the person who owns the work, its autonomy level, spend against a budget and progress through its steps. A billing dispute waits for the Service lead to approve a $240 credit before the agent writes it.

Autonomy levels on the board

  1. L1Suggestrecommends only
  2. L2Drafta person sends it
  3. L3Act with approvala person approves
  4. L4Act within limitsescalates outside
Typical length
Strategy in 3 weeks · first agent live by week 8
Scope
Use cases · agents · copilots
Standard
Tested before rollout

01Why now

Why AI pilots stall, and what we settle before building.

Pilots rarely fail on the model. They stall when nobody owns the outcome, the data is not ready, the risk is unclear or the running cost is unknown.

Pilot funnel · per 100 ideas

Illustrative
  1. Ideas raisedEvery team has a list 100
  2. 01 Stalls on: No owner, no measure −58
  3. Prototypes builtA demo on sample data 42
  4. 02 Stalls on: Data not ready −24
  5. Pilots with usersLive data, a handful of people 18
  6. 03 Stalls on: Risk unclear −10
  7. Approved for rolloutOwner, risk and cost signed off 8
  8. 04 Stalls on: Unit cost unknown −3
  9. In daily useMeasured every week 5
A presenter walks colleagues through a chart on a large screen in a glass-walled meeting room
The demoThe demo is easy. Daily use is the hard part.

Agreed before anything is built

  • A named owner for the work
  • A baseline and a target measure
  • Data readiness, scored
  • A risk tier and an autonomy level
  • A cost-per-task ceiling
  1. 01

    No owner, no measure

    Nobody owns the process the demo changes or the number it should move.

    How it is closedEach use case gets a named owner, a baseline and a target before build.

    The portfolio
  2. 02

    Data not ready

    The prototype ran on a clean sample. Production data is scattered, stale or has no API.

    How it is closedData readiness is scored per use case, and every gap goes on the roadmap with an owner.

    The workshop
  3. 03

    Risk unclear

    Nobody can say what the agent may touch or who approves it, so legal says no.

    How it is closedA risk tier, an autonomy level, approvals and an audit log, mapped to NIST AI RMF and the EU AI Act.

    Governance
  4. 04

    Unit cost unknown

    Nobody knows what each task costs at full volume compared with doing it by hand.

    How it is closedWe measure cost per task from day ten and compare it with the manual baseline.

    Agent testing

02Use-case portfolio

Use cases ranked, and the top three costed.

We score each idea on four criteria. You get a ranked roadmap by quarter and business cases for the top three.

Handwritten sticky notes on a printed canvas listing customer problems and pains
Value before buildEvery use case starts from a named problem and a costed value case, before anyone builds an agent.Photo: Daria Nepriakhina / Unsplash
  • Value

    Hours saved or revenue affected, based on your teams’ own volumes.

  • Feasibility

    Data readiness and integration effort: where the data lives and how clean it is.

  • Risk

    EU AI Act tier, personal data involved and the harm if the agent is wrong.

  • Time to value

    Weeks to a measured pilot, data and integration work included. Each week costs 0.05 points.

A chart of twelve use cases by value and feasibility. The three in the high-value, high-feasibility corner are marked to build first: reconcile vendor invoices, triage inbound RFPs and draft renewal quotes. The ranked list that follows holds the same information.

Bubble size = hours of work a month · axes 3–10Illustrative portfolio

score = 0.45·value + 0.35·feasibility + risk − 0.05·weeksrisk: +1 minimal · +0.4 limited · −1 high

#Use caseScoreRisk

What the portfolio hands over

  • Ranked roadmapAll twelve, sequenced by quarter with dependencies
  • Business case × 3Baseline, target, build and run cost, payback
  • Owners and measuresOne accountable owner and one number per use case

Selected · rank 01 · score 7.5

Reconcile vendor invoices

Finance · owner Finance lead · 4–6 wks to a measured pilot

Value
360 hours a month of three-way matching across about 2,900 invoices.
Feasibility
Invoices and purchase orders already sit in the ERP. About 18% arrive as PDFs and need extraction.
Risk Minimal risk
Minimal tier. Supplier data only, and payment release stays with a person.

DecisionBuild first

Pilot starts atL2 · Draft

Autonomy ceilingL4 · only after 90 days at L3

03Agent demo · Illustrative

Agent autonomy, set by you and logged at every step.

Pick a task, set the level and limits, and run the agent. The level changes only the last step. Set a limit below the agent’s proposal and it escalates to a person.

agent-runner · Your company · sandbox Live demo · touch any control to take over

Interactive agent run: choose a task, an autonomy level and limits, then press Run. The trace lists each step; the run ends in a suggestion, a draft, an approval request, an action within limits, an escalation or a halt, and adds an audit row. Outcomes are pre-written; no live model is called.

1 · Task
2 · Autonomy level
3 · Limits
Agent will propose 8%

run_7f3a21 · renewals-agent · prompt renewals v14.1

Renew a customer contract

L3 · Act with approval
Latency
3.9 s
Tokens
4.6k → 800
Records
3
Spend
$0.020

Run complete · approved by Sales ops lead · quote sent · audit entry written

  1. Planplannerrouter-small

    inGoal: renew account 20-418 before 14 Oct

    out4 sub-tasks · tools kb, crm, calc

    640 ms1.2k → 180$0.0003

  2. Retrieve policykb.searchreadembed

    in“renewal discount policy”

    outPricing policy v4 §3.2 · discounts up to the approval limit

    410 ms40 → 0<$0.0001

  3. Read systemcrm.get_accountread

    inaccount_id = 20-418

    outARR $48,000 · 3 seats added · usage +18%

    280 ms——

  4. Calculatecalc.renewal

    inARR, seats, 24-month term

    outPrice $51,840 · discount 8%

    35 ms——

  5. Draftllm.draftreasoner-large

    inQuote and cover email in the account’s tone

    outQuote Q-7731 · email of 146 words

    2100 ms3.4k → 620$0.020

  6. Policy checkpolicy.check

    indiscount, records, spend vs limits

    outDiscount 8% ≤ 10% · records 2+1 ≤ 25 · spend $0.020 ≤ $0.25 · cites §3.2

    90 ms——

  7. Write actioncrm.update_quotewrite

    inQ-7731 → status “sent”

    outApproved by Sales ops lead · Q-7731 sent

    320 ms——

Ready

Press Run agent to start. Change a limit first to see a guard step in.

Suggestion · no write tools at L1

Renew account 20-418 for 24 months at $51,840 with an 8% discount. Sources: pricing policy v4 §3.2 and CRM usage.

Sent to Sales ops lead. A person does the work.

Draft saved · crm.save_draft

Editable. Nothing leaves until Sales ops lead sends it.

Approval needed · Sales ops lead

Send quote Q-7731 to the customer?

Evidence: pricing policy v4 §3.2 · CRM usage +18% · 3 records affected

Executed after approval

Quote Q-7731 sent. Approved by Sales ops lead.

Write
crm.update_quote · Q-7731 → status “sent”
Approver
Sales ops lead
Waited for a person
14 s
Audit row
run_7f3a21 · append-only

Run stopped by a guard

The run passed its limit and stopped before acting.

Handed to Sales ops lead with the trace so far.

Stopped at
—
Written
Nothing
Spend so far
—
Audit row
—

Audit log · append-only · newest first

  1. 10:42:07run_7f3a21L3Renew a customer contractExecuted after approval

    Approver
    Sales ops lead
    Models
    router-small · embed · reasoner-large
    Prompt
    renewals v14.1
    Tools
    kb.search:r · crm.get_account:r · crm.update_quote:w
    Records
    3
    Spend
    $0.020
  • Tool allow-list

    Only the tools each level grants.

  • Read and write scopes

    Broad reads. Narrow, short-lived writes.

  • Spend and record limits

    Checked before every step and every write.

  • Policy check, cited

    Each write checked against the policy it quotes.

  • Human approval

    Consequential actions wait for a named person.

  • Full audit log

    Model, prompt, tools, spend and approver.

OWASP Top 10 for LLM Applications — Security risks in generative AI applications

LLM06 · Excessive Agency. We apply OWASP’s mitigations: minimal tools and permissions, approval for high-impact actions and every call checked outside the model. Prompt injection (LLM01) is tested on the same run.

04How it works

What a production agent is made of.

A demo needs a prompt and a model. A production agent needs eight parts, each replaceable, tested and monitored. We choose the technology for each part per client.

  • Models

    plan

    Chosen per step: a small model to route, a larger one where reasoning pays. Hosted APIs or open-weight models in your region.

    • Anthropic
    • Google Gemini
    • Meta Llama
    • Mistral AI
    • OpenAI
    • vLLM
    • Hugging Face
  • Workflow

    plan

    Each run has checkpoints and retries, and resumes after a failure.

    • LangGraph
    • LangChain
    • Temporal
    • Python
  • Memory & knowledge

    plan

    Session state in Redis; knowledge in a vector index that keeps a link to every source.

    • PostgreSQL
    • pgvector
    • Qdrant
    • Redis
    • LlamaIndex
  • Identity & permissions

    act

    Each agent has its own identity, with short-lived, least-privilege credentials.

    • Okta
    • Auth0
    • Vault
    • OpenID

A diagram of an agent loop: plan, act, observe, repeated until the task is done. Around it sit eight parts: models, orchestration, memory and knowledge, and identity and permissions on one side; tools and APIs, guardrails, evals and observability on the other. Each part lists the technologies used for it.

  • Tools & APIs

    act

    MCP servers wrap your systems once, so any agent can use them with scoped permissions.

    • Model Context Protocol
    • FastAPI
    • OpenAPI Initiative
    • HubSpot
    • Salesforce
  • Automatic checks

    act

    Policy rules on every tool call, personal data redacted, and input and output checks run outside the model.

    • Open Policy Agent
    • Microsoft Presidio
    • NeMo Guardrails
    • Llama Guard
  • Testing

    observe

    Your evaluation set, tool-order checks and calibrated judge-model scoring, run on every change before merge.

    • promptfoo
    • Ragas
    • MLflow
    • GitHub Actions
  • Monitoring

    observe

    Every step traced, with latency, tokens, cost and tool inputs, in the tools your engineers already use.

    • OpenTelemetry
    • Langfuse
    • Grafana
    • Datadog
  • Models you can swap

    Swapping a model changes one route and reruns the full test suite before release.

  • Your data stays yours

    Enterprise API terms or self-hosted open-weight models. Nothing trains on your data.

  • Region where required

    Models run and data is stored in the region your policy, GDPR or the DPDP Act requires.

05Agent testing

Agents are tested like software, on every change.

Each agent has an evaluation set built from your own cases. Changes merge only when they pass it, and a person still signs each release.

eval-harness · renewals-agent · golden set v10 · 218 cases CI gate · merge allowed

An illustrative evaluation dashboard for one agent: five gate metrics, success per scenario against its gate, a comparison of three models, the cost and energy saved by routing a small and a large model per step, and a regression timeline in which a prompt change lets the agent write before its policy check, the gate blocks the merge, the rule moves into the tool router and the re-run passes.

  • Task success

    94.2%

    gate ≥ 92%

  • Tool-call accuracy

    98.1%

    gate ≥ 97%

  • Policy violations

    0

    must be 0

  • Cost per task

    $0.021

    manual $4.80

  • p95 time per task

    6.8 s

    gate ≤ 10 s

Success by scenarioIllustrative

  1. Standard renewal64 cases 98%gate ≥ 95%
  2. Seats changed mid-term38 cases 96%gate ≥ 92%
  3. Discount over the limitmust escalate25 cases 100%gate = 100%
  4. Usage data missing25 cases 88%gate ≥ 85%
  5. Prompt injection in the account notesmust refuse · LLM0130 cases 100%gate = 100%
  6. Ambiguous contract termmust ask36 cases 84%gate ≥ 80%
Judge agreement
κ 0.84 vs human labels
Double-labelled
40 of 218 cases
Last calibrated
release 0418

Model comparison · every step on one modelSame 218 cases

Metricrouter-smallhosted · fastreasoner-largehosted · strongestopen-weight 70Bself-hosted · in-region
Task success86.1%94.6%91.7%
Tool-call accuracy97.4%98.3%97.9%
Cost per task$0.006$0.058$0.031
p95 time per task3.9 s7.4 s8.1 s
Policy violations000
Used forRoutes, classifies, extractsDrafts and reasonsFallback where data must stay in-region

Right-sized · cost and carbon per taskSCI · ISO/IEC 21031:2024 — Software Carbon Intensity

  • Every step on reasoner-large $0.0582.4 Wh est. · 94.6% success
  • Routed per step · context cached $0.0210.9 Wh est. · 94.2% success

SCI = ((E × I) + M) per R with R = one completed task. The manual process is the ceiling: about 4.8 Wh a task (4.8 minutes at a workstation). An agent that does not beat it on cost and carbon does not ship.

Regression timeline · last four runs · trajectory checks on

  1. run 0412 renewals v13 214 / 214 gates passed · every trajectory in the allowed tool order merge allowed
  2. run 0418 renewals v14 Trajectory check failed: crm.update_quote called before policy.check in 6 of 25 “discount over the limit” cases, so the agent wrote instead of escalating merge blocked expected…draftpolicy.checkcrm.update_quote got…draftcrm.update_quotepolicy.check
  3. fix renewals v14.1 Write tools now unlock only after policy.check passes, enforced in the tool router rather than the prompt · 4 trajectory cases added (218) awaiting re-run
  4. run 0419 renewals v14.1 218 / 218 gates passed · 0 out-of-order tool calls · released after sign-off by the Sales ops lead merge allowed
  • Test cases from your experts

    Cases from your own tickets, labelled by the people who do the work. New edge cases from production join every sprint.

  • Judge model, calibrated

    A judge model scores every case; its agreement with human labels (Cohen’s κ, target ≥ 0.8) is checked each release.

  • Adversarial cases

    Prompt injection, over-limit actions and missing data. Must-refuse and must-escalate cases must all pass.

  • Test gate in CI

    Every prompt, tool, model or retrieval change runs the suite, including tool-order checks. A failure blocks the merge.

06AI governance map

AI governance that lets legal say yes.

We write down what each agent may do, who approves it and which rules apply, mapped to the frameworks your auditors use, before the first agent goes live.

Rows of numbered keys hanging on a board, each on its own hook
Only the keys it needsEvery permission is numbered, owned and handed out one at a time: an agent gets only the keys its job needs.Photo: tokyo rattus / Unsplash

18 artefacts across the four functions of NIST AI RMF 1.0 · frameworks we build to, with no certification claimed

01GOVERN

Govern

Policies, roles and accountability, so every decision about AI has an owner.

  • AI use policy and principles

    NIST AI RMF GOVERN 1ISO/IEC 42001 A.2EU AI Act Art. 4

  • Roles, approval rights and a RACI

    NIST AI RMF GOVERN 2ISO/IEC 42001 A.3

  • AI system inventory

    NIST AI RMF GOVERN 1.6ISO/IEC 42001 A.4

  • Training record for owners and approvers

    EU AI Act Art. 4 AI literacyNIST AI RMF GOVERN 2

No artefact in this function for the chosen framework.

02MAP

Map

Each use case in context: who it affects, what could go wrong, which rules apply.

  • Context sheet and EU AI Act risk tier

    NIST AI RMF MAP 1ISO/IEC 23894 §6.3 contextEU AI Act Art. 6 · Annex III

  • Risk assessment with treatment plan

    ISO/IEC 42001 6.1.2 · 6.1.3ISO/IEC 23894 §6.4 · §6.5NIST AI RMF MAP 5

  • Impact assessment where high-risk

    ISO/IEC 42001 A.5EU AI Act Art. 27GDPR Art. 35 DPIADPDP Act 2023 §10

  • Data map and lawful basis

    GDPR Art. 6DPDP Act 2023 §4 · §5ISO/IEC 42001 A.7

  • Threat model for the agent

    OWASP LLM Top 10 LLM01 · LLM06NIST AI RMF MAP 5

No artefact in this function for the chosen framework.

03MEASURE

Measure

Evidence, not opinion: every risk has a test, a number and a trend.

  • Eval suite, golden sets and red-team results

    NIST AI RMF MEASURE 2EU AI Act Art. 15OWASP LLM Top 10 LLM01–LLM10

  • Model cards per model and version

    EU AI Act Art. 11 · Art. 13ISO/IEC 42001 A.8

  • Bias and fairness tests where people are affected

    EU AI Act Art. 10GDPR Art. 22NIST AI RMF MEASURE 2

  • Monitoring: success, cost, escalations, drift

    NIST AI RMF MEASURE 3EU AI Act Art. 26

No artefact in this function for the chosen framework.

04MANAGE

Manage

Controls in place, kept working after launch, and a plan for when they fail.

  • Human-oversight design: the four autonomy levels

    EU AI Act Art. 14NIST AI RMF MANAGE 1ISO/IEC 42001 A.9

  • Guardrail and approval policy

    OWASP LLM Top 10 LLM06NIST AI RMF MANAGE 2

  • Disclosure: people are told they are talking to AI

    EU AI Act Art. 50DPDP Act 2023 §5 notice

  • Audit log and retention rules

    EU AI Act Art. 12 · Art. 26GDPR Art. 5DPDP Act 2023 §8

  • Change, rollback and incident response

    ISO/IEC 42001 A.6NIST AI RMF MANAGE 4DPDP Act 2023 §8(6)GDPR Art. 33

No artefact in this function for the chosen framework.

Frameworks we build to

  • NIST AI RMF 1.0AI Risk Management Framework
  • ISO/IEC 42001:2023Artificial intelligence management systems
  • ISO/IEC 23894:2023AI risk management guidance
  • EU AI ActArtificial Intelligence Act (EU) 2024/1689
  • OWASP Top 10 for LLM ApplicationsSecurity risks in generative AI applications
  • DPDP Act 2023Digital Personal Data Protection Act, 2023
  • GDPRGeneral Data Protection Regulation (EU) 2016/679

EU AI Act · obligations phase in

  1. 2 Feb 2025 in forceProhibited practices and AI literacy dutiesArt. 4 · Art. 5
  2. 2 Aug 2025 in forceGeneral-purpose model obligations and governanceChapter V
  3. 2 Aug 2026 in forceTransparency duties and most remaining provisionsArt. 50
  4. 2 Dec 2026 aheadContent marking for systems already on the market, and a new prohibited practiceArt. 50(2) · Art. 5
  5. 2 Dec 2027 aheadHigh-risk systems listed in Annex III, deferred by the 2026 amendmentArt. 6(2) · Annex III
  6. 2 Aug 2028 aheadHigh-risk systems embedded in regulated productsArt. 6(1) · Annex I

Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force since 27 July 2026. It moved the high-risk dates later, set 2 December 2026 for content marking by systems already on the market and added a prohibited practice from the same day. We plan against the text in force.

What this covers

The artefacts are the evidence a certification body or regulator asks for; whether to certify is your decision. GDPR and India’s DPDP Act 2023 are in the same map, so one register covers privacy and AI risk.

Security and compliance in depth

07Typical timelines

From idea to a working agent in weeks.

We start from a reference platform and let test results decide. Your team’s speed at labelling cases and approving steps sets the calendar. Ranges shown are typical, agreed in the workshop.

A stopwatch drawn to scale over eight weeks, with five milestones: week 1 two-day workshop, week 2 prototype on your data, week 3 eval baseline and the go-to-prove gate, week 5 pilot with ten users, week 8 production at act-with-approval after the production gate. The milestone buttons that follow describe each one.

Stepping through
Typical range

What exists on week 1

  • The long list scored and the first three chosen, each with a named owner
  • Baseline measured on the manual process
  • Sandbox with masked copies of your data

workshop · 12 scored · 3 chosen · baseline 4.8 min per task · sandbox ready

Who decidesOwner and sponsor agree the measure

What exists on week 2

  • One system wired through an MCP server, read scope only
  • 40 golden cases taken from real tickets
  • First runs at L1 Suggest, traced end to end

prototype · 1 tool · 40 cases · L1 Suggest

Who decidesThe team who does the work reviews the first outputs

What exists on week 3

  • 120 cases labelled by your experts; judge calibrated (κ 0.82)
  • Cost per task and p95 time measured against the manual baseline
  • Gate G1, go to prove: decided on the numbers

evals · 120 cases · κ 0.82 · $0.03 per task · p95 8.6 s · G1 passed

Who decidesSponsor, owner and legal sign G1 on evidence, not on the demo

What exists on week 5

  • Shadow mode first, then L2 Draft for ten named users
  • Guardrails, spend limits and the approval flow live
  • Weekly review of escalations feeds the golden set (214 cases by week 7)

pilot · 10 users · L2 Draft · 214 cases · 0 violations

Who decidesLegal and security sign the guardrail policy

What exists on week 8

  • Act with approval for the whole team, after gate G2
  • Audit log, monitoring and runbooks handed over
  • Decision record: the criteria for moving to L4 later

production · L3 · G2 passed · audit log · runbooks · review in 90 days

Who decidesSponsor signs the release; the owner runs it

Model swap · evaluating

reasoner-large v2 behind the same route

  1. 1New model behind the same routeNo code change: the route points at the new model in the sandbox.0 h
  2. 2Eval suite re-run: 218 casesTwo regressions on “ambiguous contract term”, fixed with one prompt line and re-run.day 1
  3. 3Rolled out at L3 with sign-offAudit log records the model version; the old route stays warm for a rollback.day 2

Model swap · rolled out

Re-evaluated and live in two days

Cases re-run
218 / 218
Regressions
2 found · 2 fixed
Policy violations
0
Cost per task
$0.021 → $0.017

Model version recorded in the audit log · old route kept warm for rollback

  • Sandbox on day one

    Masked copies of your data in an isolated environment, so nothing waits on a production connection.

  • A reference platform

    MCP servers, test suite, automatic checks, audit log and tracing, ready to configure for your systems.

  • Test results decide

    Every gate is a number from the tests, so go or no-go is a dashboard check with the owner.

08The workshop

AI strategy with the people who run the work.

Two days in the first week with operations, legal, security, finance and the team that does the work. Your AI ideas are scored and become a one-page strategy the room has agreed.

Two colleagues place handwritten sticky notes on a glass wall divided into quarters
Before scoringThe long list on the wall, sorted by quarter.

The week before · a sourced pre-readAgents summarise your process documents, a sample of tickets and the system inventory into a short pre-read, every line linked to its source, so the workshop is spent deciding.

Hands sort handwritten yellow notes across a blue table
Day 1Sorting the long list.
  1. Day 1 · Morning

    How the work runs today

    Three to five processes walked end to end by the people who do them.

    Process maps with baseline times

  2. Day 1 · Afternoon

    The long list, scored

    Every idea on the wall, scored on value, feasibility, risk and time.

    A scored long list, usually a dozen

Two colleagues arrange sticky notes into columns on a whiteboard in a meeting room
Day 2The candidates sorted into columns.
  1. Day 2 · Morning

    Risk and data, per candidate

    Legal sets the tier, security the scopes, system owners the data readiness.

    Risk tier, autonomy ceiling, data gaps

  2. Day 2 · Afternoon

    Pick the first three

    Three to build first, each with an owner, a baseline and a target.

    The one-page strategy, signed

AI strategy · one page · Your company · v1Illustrative

What leaves the room

  1. 1Ambition and measures

    Return about 700 hours a month to sales ops and finance by the end of Q2; cost per task under a quarter of manual.

  2. 2Portfolio · build first

    Reconcile vendor invoices · Draft renewal quotes · Triage inbound RFPs. Nine more sequenced by quarter.

  3. 3Operating model

    Embedded: agents owned by the teams that run the process, with a small central platform and review group.

  4. 4Data readiness gaps

    Billing history split across two systems · 18% of invoices arrive as PDFs · CRM firmographics patchy.

  5. 5Budget and measures

    Build, run and review budget per use case; one owner and one number each; quarterly review.

  6. 6Risk tier and governance

    Two limited-risk use cases carry disclosure duties; two high-risk candidates parked with a compliance plan.

Agreed in the room

  • COO
  • CFO
  • CISO
  • Head of sales ops

09How we work with you

Assess, prove and scale, with a gate after each phase.

Work moves to the next phase only when the numbers meet criteria agreed in week one and the named people sign.

  1. 013 weeks

    Assess

    Interviews and the two-day workshop in week one, then a prototype on your data and an evaluation baseline for the first use case.

    Leaves with

    • Scored use-case portfolio
    • One-page AI strategy and roadmap
    • Business cases for the first three
    • Prototype and evaluation baseline

    Autonomy reachedL1 · SuggestA prototype suggests, in a sandbox on masked data

    G1 Go to prove Passed

    • A named owner, a measured baseline and an evaluation baseline for the first use case
    • Data access agreed at read scope, sandbox ready
    • Risk tier set and the autonomy ceiling written down
    • Budget approved, with a cost-per-task ceiling

    Signed bySponsor · owner · legal

  2. 025 weeks

    Prove

    One agent, tested on your experts’ cases, piloted with ten users at L2 from week 5 and released at L3 in week 8, with approval on every write.

    Leaves with

    • Working agent and its MCP servers
    • Test suite, evaluation set, red-team results
    • Check and approval policy

    Autonomy reachedL3 · Act with approvalActs on your systems once a named person approves

    G2 Go to production Passed

    • Test gate passed on the evaluation set, adversarial cases included
    • Zero policy violations across every run
    • Cost per task below the ceiling, 95th-percentile time on target
    • Check policy signed by legal and security
    • Runbooks and on-call handed over

    Signed bySponsor signs · owner runs

  3. 03Ongoing

    Scale

    The next agents from the portfolio, a risk register kept current, and autonomy raised one task at a time on evidence.

    Leaves with

    • Next agents from the roadmap
    • AI inventory and risk register kept current
    • Quarterly review against the measures

    Autonomy reachedL4 · Act within limitsFor named tasks, inside limits, after G3

    G3 Raise autonomy Every 90 days

    • 90 days at L3 with approvers changing under 2% of actions
    • Escalations answered within the agreed time
    • No incident traced to the agent
    • One task moves up one level, with new limits

    Signed byOwner · risk · security

10What you get

Strategy, agent, governance and runbooks, in your accounts.

What we make for you is yours once it is paid for. Tools we already had stay ours, and you get a free, permanent licence to use them.

Tab 1 of 4Your company · AI programme · handover

Strategy

Where AI fits, what comes first and what each use case is worth.

  • ai-strategy-one-page.pdf One-page AI strategy Document
  • use-case-portfolio.xlsx AI opportunity map and use-case portfolio BoardSheet
  • roadmap-business-cases.pptx AI roadmap with business cases DeckSheet
  • operating-model.pdf Operating model, roles and approval rights Document

Handover checklist

  • Signed by the sponsor and the three owners
  • Source files kept in your document store
  • Portfolio scores editable, formula included

Tab 2 of 4Your company · AI programme · handover

Agent

The agent, in your repository and cloud account, with the tests it must pass.

  • agents/renewals-agent/ Agent or copilot in production Sourceyour cloud
  • mcp-servers/ MCP servers for the CRM, billing and documents Source
  • evals/ Evaluation suite, test cases, red-team results Testsreport
  • infra/ Deployment as code for your cloud account Terraform

Handover checklist

  • Repository transferred to your organisation
  • CI, secrets and cloud account in your name
  • Prompts and model routes versioned in the repo

Tab 3 of 4Your company · AI programme · handover

Governance

The written answers legal, security and auditors ask for, kept in one register.

  • guardrails/policy.yaml Check and approval policy Policyconfig
  • ai-governance-framework.pdf AI governance framework Document
  • ai-inventory-risk-register.xlsx AI inventory and risk register Register
  • model-cards/ Model cards per model and version Document

Handover checklist

  • Register owner named, review dates booked
  • Legal and security sign-off recorded
  • Every row linked to its evidence

Tab 4 of 4Your company · AI programme · handover

Operate

Everything your team needs to run, review and change the agent without us in the room.

  • dashboards/agent-actions Agent action and audit log Dashboard
  • runbooks/ Runbooks: escalations, rollback, model swap Document
  • training/ Training for owners, approvers and reviewers Sessionsrecordings
  • reviews/q-review-template.pdf Quarterly review pack Report

Handover checklist

  • On-call rota and escalation contacts agreed
  • Approvers trained and named in the policy
  • First quarterly review in the calendar

Each tab shows a mock of the first artefact in that part of the handover: a roadmap by quarter for strategy, the repository layout for the agent, the AI inventory and risk register for governance, and an escalation runbook for operations.

1190-day review

Six measures decide whether an agent stays.

Every agent is reviewed on six measures, each with a pass line and a hold line. An agent that misses them goes back a level. The lines shown are review targets.

Agent review · day 90

renewals-agentDraft renewal quotesowner Sales ops lead

Measure Hold and pass lines At review Status

Hours returned per week

Handling time saved × volume, against the baseline taken before build

Hold < 30 hPass ≥ 60 h

74 h

Pass

Task success on evals

The golden set, re-run on every change and at each review

Hold < 85%Pass ≥ 92%

94.2%

Pass

Cost per task vs manual

Models, tools, hosting and review time per completed task, as a share of manual cost

Pass ≤ 25%Hold > 50%

13% · $0.62 vs $4.80

Pass

Escalation rate

Share of runs handed to a person, by reason

Pass ≤ 10%Hold > 20%

6.8%

Pass

Policy violations

Writes outside policy, from the audit log and the red-team suite

Pass = 0Hold ≥ 1

0

Pass

Carbon per task vs manual

SCI estimate per completed task, as a share of the manual process on the same grid

Pass ≤ 50%Hold ≥ 100%

19% · 0.9 vs 4.8 Wh est.

Pass

Decision

Stays at L3L3 · Act with approval

Review in 90 days whether quotes with discounts of 5% or less can move to L4, inside limits.

Signed · ownerReviewed · risk and securityLogged · audit trail

Illustrative reviewLines are agreed per agent before build. Cost includes the time people spend approving and reviewing.

Services & packages

Decide where AI pays off, then build the agents that do the work.

We rank your AI use cases, then design and build the agents and copilots, with the governance to run them. Each is tested on your own examples and overseen by your people.

Categories
04
Services
15
Packages
04
Not sure what you need? Describe the problem

How to buy

  1. 01Pick services. Enquire about one, or add several to a brief.
  2. 02Choose a package. A sprint, a fixed project or an ongoing team.
  3. 03Send the brief. We reply within one working day.

Browse by category

Timelines are typical. Every quote follows a written scope.

01AI strategy

4 services
Typical timeline: 3–4 weeks

AI readiness & opportunity mapping

We review your processes, data and systems with the teams who run them and show where AI would cut time, cost or risk.

What’s included

  • Leadership and team interviews
  • Process walk-throughs and data review
  • Opportunity map with value and effort
  • Data and systems readiness review
  • Discovery
  • Opportunity map
  • Readiness

Best forLeadership teams under pressure to act on AI and unsure where to start.

Typical timeline: 2–4 weeks

Use-case portfolio & business case

Every AI idea scored on value, feasibility, data readiness and risk, then sequenced into a costed plan with owners and success measures.

What’s included

  • Scoring model for value, feasibility and risk
  • Cost and benefit estimate per use case
  • Sequenced 12-month roadmap
  • Success measures and baselines
  • Scored
  • Sequenced
  • Business case

Best forTeams with a long list of AI ideas and a limited budget.

Typical timeline: 2–4 weeks

Model & platform selection

We test models and AI platforms on your own examples for quality, speed, cost and where data is stored, then pick one per task.

What’s included

  • Shortlist across OpenAI, Anthropic, Google, Mistral and open-weight models
  • Side-by-side tests on your tasks
  • Cost-per-task and latency estimates
  • Data residency and contract terms reviewed
  • Model-agnostic
  • Tested on your data
  • OpenAI

Best forTeams unsure which model or vendor to commit to.

Typical timeline: 1–2 weeks

Leadership AI briefing

A working session for your board or leadership team on what AI can and cannot do today, based on your sector and your own data.

What’s included

  • Pre-read tailored to your business
  • Live demonstrations on relevant tasks
  • Risk and governance primer
  • Agreed next steps
  • Workshop
  • Leadership

Best forBoards and leadership teams starting their AI plan.

02Agents & copilots

4 services
Typical timeline: 8–12 weeks

Custom AI agent

Software that completes multi-step tasks in your systems, such as checking an order or updating a record, and hands over to a person when needed.

What’s included

  • Task design with scoped tool permissions
  • Integration with your APIs and systems
  • Human approval for consequential actions
  • Full action log and audit trail
  • Evaluation suite before go-live
  • Tool use
  • Human approval
  • Audit log
  • OpenAI

Best forTeams with repetitive, rules-heavy work spread across several systems.

Typical timeline: 6–10 weeks

Copilots for teams

An assistant in Slack, Teams, your CRM or other tools your people already use, answering from your own documents and data.

What’s included

  • Connection to approved knowledge sources
  • Permission-aware answers with citations
  • Deployment in Slack, Teams or a web app
  • Usage and feedback analytics
  • Your data
  • Cited
  • In your tools
  • Microsoft Teams
  • Slack
  • LlamaIndex
  • OpenAI

Best forSales, support, HR and operations teams looking for answers all day.

Typical timeline: 10–16 weeks

Multi-agent workflows

Several specialised agents working together on a larger process, such as research, drafting, checking and filing, with clear handoffs and checkpoints.

What’s included

  • Workflow and agent role design
  • Coordination with state and retries
  • Checkpoints for human review
  • Cost and step limits
  • Coordination
  • Checkpoints
  • Limits

Best forComplex processes that already follow documented steps.

Typical timeline: 4–8 weeks

Agent tools & MCP integrations

We connect agents to your internal systems through defined tools and Model Context Protocol (MCP) servers, with only the access each task needs.

What’s included

  • Tool and MCP server design
  • Authentication and scoped permissions
  • Rate and spend limits
  • Logging of every tool call
  • MCP
  • Least privilege
  • Tool calls

Best forCompanies that want agents to act in internal systems without over-broad access.

03Pilots & evaluation

3 services
Typical timeline: 6–12 weeks

AI pilot with evaluations

A time-boxed pilot of one use case, measured against a baseline, so the go or no-go decision rests on evidence.

What’s included

  • Baseline measure before the pilot
  • Task-level evaluation set built with your experts
  • Shadow-mode run, then supervised live use
  • Decision record and rollout plan
  • Baseline
  • Shadow mode
  • Go / no-go

Best forOrganisations ready to test one use case before scaling.

Typical timeline: 3–6 weeks

Evaluation suite design

Automated tests that score an AI system on accuracy, task completion, safety and cost every time a prompt or model changes.

What’s included

  • Evaluation sets drawn from your own examples
  • Automated scoring calibrated against human review
  • Red-team cases for prompt injection and tool misuse
  • Evaluation gate in your release pipeline
  • Evaluation in CI
  • Regression tests

Best forTeams with live AI that cannot tell whether a change made it better or worse.

Typical timeline: 6–10 weeks

Pilot to production

Take a promising prototype and make it dependable: monitoring, fallbacks, access control, cost limits and support.

What’s included

  • Production architecture review
  • Safety checks, fallbacks and rate limits
  • Monitoring of quality, latency and cost
  • Runbook and handover
  • Hardening
  • Monitoring
  • Cost limits
  • AWS
  • Microsoft Azure

Best forTeams with a pilot that works in demos but is not ready for customers.

04Governance & operating model

4 services
Typical timeline: 4–8 weeks

AI governance framework

Rules on which uses of AI are allowed, who approves them and how they are monitored, aligned with the NIST AI RMF and ISO/IEC 42001:2023.

What’s included

  • AI policy and acceptable-use rules
  • Risk tiers and an approval workflow
  • Inventory of AI systems in use
  • Impact assessment template
  • NIST AI RMF
  • ISO/IEC 42001
  • Policy

Best forOrganisations adopting AI across several teams at once.

Typical timeline: 3–6 weeks

AI operating model & team design

Who owns AI in your organisation (a central team, specialists in business units or both), with the roles, skills and budget process to match.

What’s included

  • Operating model options and a recommendation
  • Roles and responsibilities (RACI)
  • Skills gap and hiring plan
  • Funding and prioritisation process
  • Centre of excellence
  • RACI
  • Skills

Best forCompanies moving from scattered experiments to a managed AI programme.

Typical timeline: 3–6 weeks

EU AI Act & regulatory mapping

Each use case mapped to its EU AI Act risk tier and obligations, alongside GDPR and India’s DPDP Act 2023, with your legal counsel.

What’s included

  • Risk classification per use case
  • Obligations map per AI system
  • Gap list with named owners
  • Readiness plan
  • EU AI Act
  • DPDP Act 2023
  • GDPR

Best forBusinesses selling into Europe or processing personal data with AI.

Typical timeline: 2–6 weeks

AI adoption & training

Training so your people use approved AI tools well and safely: how to ask, how to check the output and when not to use AI.

What’s included

  • Role-based training sessions
  • Prompting and review playbooks
  • Safe-use guidelines
  • Adoption measurement
  • Training
  • Playbooks
  • Safe use

Best forTeams given AI tools without guidance on using them well.

Your brief

Tick “Add to brief” on any service, choose a package, then continue. Or enquire about one service directly.

How we work with you

Ways to engage, from a question to an RFQ.

Ask a quick question, send a project brief or issue a formal RFQ. The lead for the work reads each one in full, and any services already in your brief go with it.

Or book a thirty-minute call

What are you sending?

  1. 01

    About 2 minutes4 required answers

    For a first conversation, a press request or anything that does not need a scope yet.

    You get A reply from a named lead

  2. 02Recommended

    About 8 minutes5 short steps

    Goals, audiences, a budget band and timing, so our first reply can outline the work.

    You get Options and a first scope after one call

  3. 03

    About 15 minutesYour documents attached

    Your documents, deadlines and the procurement and security rules the work must meet.

    You get Receipt confirmed and a named bid lead

How it is priced

Each package shows its pricing model. Work starts once a written scope and quote are agreed.

  • Typical length
    1–3 weeks
    Pricing
    Fixed fee
  • Typical length
    4–12 weeks
    Pricing
    Fixed price
  • Typical length
    3–9 months
    Pricing
    Fixed price per milestone
  • Typical length
    6–18 months
    Pricing
    Programme fee · by statement of work
Compare what each package includes
What each package includes and who it suits
PackageEvery engagement includesBest for
SprintA short, fixed-scope engagement that answers one defined question.
  • Scope and outcome agreed before day one
  • A senior lead plus the specialists needed
  • A working review every week
  • A decision-ready answer or prototype
Discovery, a diagnostic, a prototype or a decision you need to make soon
ProjectA defined scope, delivered for a fixed price.
  • Statement of work with deliverables and acceptance criteria
  • A named project lead and a fixed team
  • A shared plan with dated checkpoints
  • Source files, yours once paid for
Work you can describe up front: an identity, a platform or a set of tools
MilestoneA larger build in phases you approve and pay for one at a time.
  • Phases with their own scope, output and sign-off
  • A go or no-go review at every gate
  • Re-planning between phases as you learn
  • Payment tied to accepted milestones
Programmes too large for one contract, where you want control at each step
EnterpriseA multi-workstream programme with a dedicated team, governance and agreed service levels.
  • An engagement director and a steering group
  • A dedicated team across several workstreams
  • Service levels, reporting and a risk register
  • Security, legal and procurement reviews in the plan
Large organisations running change across markets, portfolios or business units

12Questions

AI Strategy & Agents, asked directly.

The questions leadership, legal and security ask in the first meeting, answered plainly.

Bring to the first call

  • The processes you want to change, with rough weekly volumes
  • Who owns each one, and who approves changes to it
  • The systems and data each process touches
  • Your risk appetite, and any regulator in scope
Book a call

Start with one high-volume process with a clear owner and reachable data. We score your long list, then pilot the top candidate at L1 or L2, where the agent suggests or drafts. Back-office tasks with clear rules, such as invoice matching or quote drafting, usually make better first candidates than open-ended customer chat.

Models from OpenAI, Anthropic, Google, Mistral and open-weight families such as Llama, chosen per task on quality, speed, cost and data residency. We call them through one gateway, so you can switch when the evidence changes. Within one agent, models are chosen per step: a small, fast model routes and extracts, and a larger one is used only where the reasoning pays for itself.

Only at L4, for named tasks and inside limits you set. Limits cover spend per run, records touched and a business limit such as a maximum discount; anything outside them goes to a person. A task reaches L4 only after 90 days at L3 with approvers rarely changing its actions, and the decision is recorded in the register.

Mainly through permissions and approvals, since no filter catches every attack. All model inputs, including documents and emails, are treated as untrusted. Tools are allow-listed per level, each agent has its own scoped credentials, every write is policy-checked outside the model and high-impact actions wait for a person. Each release is red-teamed against OWASP LLM01 (prompt injection) and LLM06 (excessive agency).

Three things: model usage, the platform around it and people’s review time. The platform covers hosting, vector store and tracing. With a small model routing and context cached, model usage for a back-office task is often a few cents, depending on document length and the models chosen. We measure cost per task from the first evaluation baseline against the manual cost.

Both can apply, and we map them in one register. The EU AI Act covers AI placed on the EU market or used in the EU, or whose output is used there. Most business agents are minimal or limited risk, with transparency duties from August 2026. Duties for high-risk Annex III uses, such as recruitment or creditworthiness checks, apply from December 2027 under the 2026 amendment. The DPDP Act 2023 covers personal data of people in India; under the DPDP Rules 2025, most duties apply from May 2027. This is engineering guidance, not legal advice.

Tell us what you need built.

You will speak to a lead who would run the work, and get a straight answer on fit.

Book a call

Three ways to start

Every engagement starts with a written scope and a quote agreed before work begins.

Choose one of the three ways above