Capability 03 / 10Use cases, agents, governance
AI strategy and agents that move from pilot to daily use.
We rank your AI use cases by value, feasibility and risk, then build the first agents. Each agent is tested on your own examples, and a named person approves high-impact actions.
An illustrative mission board for “Your company”. Six missions sit in four columns: queued, agent working, needs approval and done. Each card shows the agent, the person who owns the work, its autonomy level, spend against a budget and progress through its steps. A billing dispute waits for the Service lead to approve a $240 credit before the agent writes it.
Autonomy levels on the board
- L1Suggestrecommends only
- L2Drafta person sends it
- L3Act with approvala person approves
- L4Act within limitsescalates outside
01Why now
Why AI pilots stall, and what we settle before building.
Pilots rarely fail on the model. They stall when nobody owns the outcome, the data is not ready, the risk is unclear or the running cost is unknown.
Pilot funnel · per 100 ideas
Illustrative- Ideas raisedEvery team has a list 100
- 01 Stalls on: No owner, no measure −58
- Prototypes builtA demo on sample data 42
- 02 Stalls on: Data not ready −24
- Pilots with usersLive data, a handful of people 18
- 03 Stalls on: Risk unclear −10
- Approved for rolloutOwner, risk and cost signed off 8
- 04 Stalls on: Unit cost unknown −3
- In daily useMeasured every week 5

Agreed before anything is built
- A named owner for the work
- A baseline and a target measure
- Data readiness, scored
- A risk tier and an autonomy level
- A cost-per-task ceiling
-
01
No owner, no measure
Nobody owns the process the demo changes or the number it should move.
How it is closedEach use case gets a named owner, a baseline and a target before build.
The portfolio -
02
Data not ready
The prototype ran on a clean sample. Production data is scattered, stale or has no API.
How it is closedData readiness is scored per use case, and every gap goes on the roadmap with an owner.
The workshop -
03
Risk unclear
Nobody can say what the agent may touch or who approves it, so legal says no.
How it is closedA risk tier, an autonomy level, approvals and an audit log, mapped to NIST AI RMF and the EU AI Act.
Governance -
04
Unit cost unknown
Nobody knows what each task costs at full volume compared with doing it by hand.
How it is closedWe measure cost per task from day ten and compare it with the manual baseline.
Agent testing
02Use-case portfolio
Use cases ranked, and the top three costed.
We score each idea on four criteria. You get a ranked roadmap by quarter and business cases for the top three.

Value
Hours saved or revenue affected, based on your teams’ own volumes.
Feasibility
Data readiness and integration effort: where the data lives and how clean it is.
Risk
EU AI Act tier, personal data involved and the harm if the agent is wrong.
Time to value
Weeks to a measured pilot, data and integration work included. Each week costs 0.05 points.
A chart of twelve use cases by value and feasibility. The three in the high-value, high-feasibility corner are marked to build first: reconcile vendor invoices, triage inbound RFPs and draft renewal quotes. The ranked list that follows holds the same information.
Minimal Limited · transparency duties High risk · Annex III
Bubble size = hours of work a month · axes 3–10Illustrative portfolio
score = 0.45·value + 0.35·feasibility + risk − 0.05·weeksrisk: +1 minimal · +0.4 limited · −1 high
What the portfolio hands over
- Ranked roadmapAll twelve, sequenced by quarter with dependencies
- Business case × 3Baseline, target, build and run cost, payback
- Owners and measuresOne accountable owner and one number per use case
Selected · rank 01 · score 7.5
Reconcile vendor invoices
Finance · owner Finance lead · 4–6 wks to a measured pilot
- Value
- 360 hours a month of three-way matching across about 2,900 invoices.
- Feasibility
- Invoices and purchase orders already sit in the ERP. About 18% arrive as PDFs and need extraction.
- Risk Minimal risk
- Minimal tier. Supplier data only, and payment release stays with a person.
DecisionBuild first
Pilot starts atL2 · Draft
Autonomy ceilingL4 · only after 90 days at L3
03Agent demo · Illustrative
Agent autonomy, set by you and logged at every step.
Pick a task, set the level and limits, and run the agent. The level changes only the last step. Set a limit below the agent’s proposal and it escalates to a person.
Interactive agent run: choose a task, an autonomy level and limits, then press Run. The trace lists each step; the run ends in a suggestion, a draft, an approval request, an action within limits, an escalation or a halt, and adds an audit row. Outcomes are pre-written; no live model is called.
run_7f3a21 · renewals-agent · prompt renewals v14.1
Renew a customer contract
- Latency
- 3.9 s
- Tokens
- 4.6k → 800
- Records
- 3
- Spend
- $0.020
Run complete · approved by Sales ops lead · quote sent · audit entry written
-
Plan
plannerrouter-smallinGoal: renew account 20-418 before 14 Oct
out4 sub-tasks · tools kb, crm, calc
640 ms1.2k → 180$0.0003
-
Retrieve policy
kb.searchreadembedin“renewal discount policy”
outPricing policy v4 §3.2 · discounts up to the approval limit
410 ms40 → 0<$0.0001
-
Read system
crm.get_accountreadinaccount_id = 20-418
outARR $48,000 · 3 seats added · usage +18%
280 ms——
-
Calculate
calc.renewalinARR, seats, 24-month term
outPrice $51,840 · discount 8%
35 ms——
-
Draft
llm.draftreasoner-largeinQuote and cover email in the account’s tone
outQuote Q-7731 · email of 146 words
2100 ms3.4k → 620$0.020
-
Policy check
policy.checkindiscount, records, spend vs limits
outDiscount 8% ≤ 10% · records 2+1 ≤ 25 · spend $0.020 ≤ $0.25 · cites §3.2
90 ms——
-
Write action
crm.update_quotewriteinQ-7731 → status “sent”
outApproved by Sales ops lead · Q-7731 sent
320 ms——
Ready
Press Run agent to start. Change a limit first to see a guard step in.
Suggestion · no write tools at L1
Renew account 20-418 for 24 months at $51,840 with an 8% discount. Sources: pricing policy v4 §3.2 and CRM usage.
Sent to Sales ops lead. A person does the work.
Draft saved · crm.save_draft
Editable. Nothing leaves until Sales ops lead sends it.
Approval needed · Sales ops lead
Send quote Q-7731 to the customer?
Evidence: pricing policy v4 §3.2 · CRM usage +18% · 3 records affected
Executed after approval
Quote Q-7731 sent. Approved by Sales ops lead.
- Write
- crm.update_quote · Q-7731 → status “sent”
- Approver
- Sales ops lead
- Waited for a person
- 14 s
- Audit row
- run_7f3a21 · append-only
Run stopped by a guard
The run passed its limit and stopped before acting.
Handed to Sales ops lead with the trace so far.
- Stopped at
- —
- Written
- Nothing
- Spend so far
- —
- Audit row
- —
Audit log · append-only · newest first
-
10:42:07run_7f3a21L3Renew a customer contractExecuted after approval
- Approver
- Sales ops lead
- Models
- router-small · embed · reasoner-large
- Prompt
- renewals v14.1
- Tools
- kb.search:r · crm.get_account:r · crm.update_quote:w
- Records
- 3
- Spend
- $0.020
Tool allow-list
Only the tools each level grants.
Read and write scopes
Broad reads. Narrow, short-lived writes.
Spend and record limits
Checked before every step and every write.
Policy check, cited
Each write checked against the policy it quotes.
Human approval
Consequential actions wait for a named person.
Full audit log
Model, prompt, tools, spend and approver.
LLM06 · Excessive Agency. We apply OWASP’s mitigations: minimal tools and permissions, approval for high-impact actions and every call checked outside the model. Prompt injection (LLM01) is tested on the same run.
04How it works
What a production agent is made of.
A demo needs a prompt and a model. A production agent needs eight parts, each replaceable, tested and monitored. We choose the technology for each part per client.
-
Models
planChosen per step: a small model to route, a larger one where reasoning pays. Hosted APIs or open-weight models in your region.
- Anthropic
- Google Gemini
- Meta Llama
- Mistral AI
- OpenAI
- vLLM
- Hugging Face
-
Workflow
planEach run has checkpoints and retries, and resumes after a failure.
- LangGraph
- LangChain
- Temporal
- Python
-
Memory & knowledge
planSession state in Redis; knowledge in a vector index that keeps a link to every source.
- PostgreSQL
- pgvector
- Qdrant
- Redis
- LlamaIndex
-
Identity & permissions
actEach agent has its own identity, with short-lived, least-privilege credentials.
- Okta
- Auth0
- Vault
- OpenID
A diagram of an agent loop: plan, act, observe, repeated until the task is done. Around it sit eight parts: models, orchestration, memory and knowledge, and identity and permissions on one side; tools and APIs, guardrails, evals and observability on the other. Each part lists the technologies used for it.
-
Tools & APIs
actMCP servers wrap your systems once, so any agent can use them with scoped permissions.
- Model Context Protocol
- FastAPI
- OpenAPI Initiative
- HubSpot
- Salesforce
-
Automatic checks
actPolicy rules on every tool call, personal data redacted, and input and output checks run outside the model.
- Open Policy Agent
- Microsoft Presidio
- NeMo Guardrails
- Llama Guard
-
Testing
observeYour evaluation set, tool-order checks and calibrated judge-model scoring, run on every change before merge.
- promptfoo
- Ragas
- MLflow
- GitHub Actions
-
Monitoring
observeEvery step traced, with latency, tokens, cost and tool inputs, in the tools your engineers already use.
- OpenTelemetry
- Langfuse
- Grafana
- Datadog
Models you can swap
Swapping a model changes one route and reruns the full test suite before release.
Your data stays yours
Enterprise API terms or self-hosted open-weight models. Nothing trains on your data.
Region where required
Models run and data is stored in the region your policy, GDPR or the DPDP Act requires.
05Agent testing
Agents are tested like software, on every change.
Each agent has an evaluation set built from your own cases. Changes merge only when they pass it, and a person still signs each release.
An illustrative evaluation dashboard for one agent: five gate metrics, success per scenario against its gate, a comparison of three models, the cost and energy saved by routing a small and a large model per step, and a regression timeline in which a prompt change lets the agent write before its policy check, the gate blocks the merge, the rule moves into the tool router and the re-run passes.
-
Task success
94.2%
gate ≥ 92%
-
Tool-call accuracy
98.1%
gate ≥ 97%
-
Policy violations
0
must be 0
-
Cost per task
$0.021
manual $4.80
-
p95 time per task
6.8 s
gate ≤ 10 s
Success by scenarioIllustrative
- Standard renewal64 cases 98%gate ≥ 95%
- Seats changed mid-term38 cases 96%gate ≥ 92%
- Discount over the limitmust escalate25 cases 100%gate = 100%
- Usage data missing25 cases 88%gate ≥ 85%
- Prompt injection in the account notesmust refuse · LLM0130 cases 100%gate = 100%
- Ambiguous contract termmust ask36 cases 84%gate ≥ 80%
- Judge agreement
- κ 0.84 vs human labels
- Double-labelled
- 40 of 218 cases
- Last calibrated
- release 0418
Model comparison · every step on one modelSame 218 cases
| Metric | router-smallhosted · fast | reasoner-largehosted · strongest | open-weight 70Bself-hosted · in-region |
|---|---|---|---|
| Task success | 86.1% | 94.6% | 91.7% |
| Tool-call accuracy | 97.4% | 98.3% | 97.9% |
| Cost per task | $0.006 | $0.058 | $0.031 |
| p95 time per task | 3.9 s | 7.4 s | 8.1 s |
| Policy violations | 0 | 0 | 0 |
| Used for | Routes, classifies, extracts | Drafts and reasons | Fallback where data must stay in-region |
Right-sized · cost and carbon per taskSCI · ISO/IEC 21031:2024 — Software Carbon Intensity
- Every step on reasoner-large $0.0582.4 Wh est. · 94.6% success
- Routed per step · context cached $0.0210.9 Wh est. · 94.2% success
SCI = ((E × I) + M) per R with R = one completed task. The manual process is the ceiling: about 4.8 Wh a task (4.8 minutes at a workstation). An agent that does not beat it on cost and carbon does not ship.
Regression timeline · last four runs · trajectory checks on
- run 0412 renewals v13 214 / 214 gates passed · every trajectory in the allowed tool order merge allowed
-
run 0418
renewals v14
Trajectory check failed: crm.update_quote called before policy.check in 6 of 25 “discount over the limit” cases, so the agent wrote instead of escalating
merge blocked
expected…
draft→policy.check→crm.update_quotegot…draft→crm.update_quote→policy.check - fix renewals v14.1 Write tools now unlock only after policy.check passes, enforced in the tool router rather than the prompt · 4 trajectory cases added (218) awaiting re-run
- run 0419 renewals v14.1 218 / 218 gates passed · 0 out-of-order tool calls · released after sign-off by the Sales ops lead merge allowed
Test cases from your experts
Cases from your own tickets, labelled by the people who do the work. New edge cases from production join every sprint.
Judge model, calibrated
A judge model scores every case; its agreement with human labels (Cohen’s κ, target ≥ 0.8) is checked each release.
Adversarial cases
Prompt injection, over-limit actions and missing data. Must-refuse and must-escalate cases must all pass.
Test gate in CI
Every prompt, tool, model or retrieval change runs the suite, including tool-order checks. A failure blocks the merge.
06AI governance map
AI governance that lets legal say yes.
We write down what each agent may do, who approves it and which rules apply, mapped to the frameworks your auditors use, before the first agent goes live.

18 artefacts across the four functions of NIST AI RMF 1.0 · frameworks we build to, with no certification claimed
01GOVERN
Govern
Policies, roles and accountability, so every decision about AI has an owner.
-
AI use policy and principles
-
Roles, approval rights and a RACI
-
AI system inventory
-
Training record for owners and approvers
No artefact in this function for the chosen framework.
02MAP
Map
Each use case in context: who it affects, what could go wrong, which rules apply.
-
Context sheet and EU AI Act risk tier
-
Risk assessment with treatment plan
-
Impact assessment where high-risk
-
Data map and lawful basis
-
Threat model for the agent
No artefact in this function for the chosen framework.
03MEASURE
Measure
Evidence, not opinion: every risk has a test, a number and a trend.
-
Eval suite, golden sets and red-team results
-
Model cards per model and version
-
Bias and fairness tests where people are affected
-
Monitoring: success, cost, escalations, drift
No artefact in this function for the chosen framework.
04MANAGE
Manage
Controls in place, kept working after launch, and a plan for when they fail.
-
Human-oversight design: the four autonomy levels
-
Guardrail and approval policy
-
Disclosure: people are told they are talking to AI
-
Audit log and retention rules
-
Change, rollback and incident response
No artefact in this function for the chosen framework.
Frameworks we build to
- NIST AI RMF 1.0AI Risk Management Framework
- ISO/IEC 42001:2023Artificial intelligence management systems
- ISO/IEC 23894:2023AI risk management guidance
- EU AI ActArtificial Intelligence Act (EU) 2024/1689
- OWASP Top 10 for LLM ApplicationsSecurity risks in generative AI applications
- DPDP Act 2023Digital Personal Data Protection Act, 2023
- GDPRGeneral Data Protection Regulation (EU) 2016/679
EU AI Act · obligations phase in
- 2 Feb 2025 in forceProhibited practices and AI literacy dutiesArt. 4 · Art. 5
- 2 Aug 2025 in forceGeneral-purpose model obligations and governanceChapter V
- 2 Aug 2026 in forceTransparency duties and most remaining provisionsArt. 50
- 2 Dec 2026 aheadContent marking for systems already on the market, and a new prohibited practiceArt. 50(2) · Art. 5
- 2 Dec 2027 aheadHigh-risk systems listed in Annex III, deferred by the 2026 amendmentArt. 6(2) · Annex III
- 2 Aug 2028 aheadHigh-risk systems embedded in regulated productsArt. 6(1) · Annex I
Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force since 27 July 2026. It moved the high-risk dates later, set 2 December 2026 for content marking by systems already on the market and added a prohibited practice from the same day. We plan against the text in force.
What this covers
The artefacts are the evidence a certification body or regulator asks for; whether to certify is your decision. GDPR and India’s DPDP Act 2023 are in the same map, so one register covers privacy and AI risk.
Security and compliance in depth07Typical timelines
From idea to a working agent in weeks.
We start from a reference platform and let test results decide. Your team’s speed at labelling cases and approving steps sets the calendar. Ranges shown are typical, agreed in the workshop.
A stopwatch drawn to scale over eight weeks, with five milestones: week 1 two-day workshop, week 2 prototype on your data, week 3 eval baseline and the go-to-prove gate, week 5 pilot with ten users, week 8 production at act-with-approval after the production gate. The milestone buttons that follow describe each one.
What exists on week 1
- The long list scored and the first three chosen, each with a named owner
- Baseline measured on the manual process
- Sandbox with masked copies of your data
workshop · 12 scored · 3 chosen · baseline 4.8 min per task · sandbox ready
Who decidesOwner and sponsor agree the measure
What exists on week 2
- One system wired through an MCP server, read scope only
- 40 golden cases taken from real tickets
- First runs at L1 Suggest, traced end to end
prototype · 1 tool · 40 cases · L1 Suggest
Who decidesThe team who does the work reviews the first outputs
What exists on week 3
- 120 cases labelled by your experts; judge calibrated (κ 0.82)
- Cost per task and p95 time measured against the manual baseline
- Gate G1, go to prove: decided on the numbers
evals · 120 cases · κ 0.82 · $0.03 per task · p95 8.6 s · G1 passed
Who decidesSponsor, owner and legal sign G1 on evidence, not on the demo
What exists on week 5
- Shadow mode first, then L2 Draft for ten named users
- Guardrails, spend limits and the approval flow live
- Weekly review of escalations feeds the golden set (214 cases by week 7)
pilot · 10 users · L2 Draft · 214 cases · 0 violations
Who decidesLegal and security sign the guardrail policy
What exists on week 8
- Act with approval for the whole team, after gate G2
- Audit log, monitoring and runbooks handed over
- Decision record: the criteria for moving to L4 later
production · L3 · G2 passed · audit log · runbooks · review in 90 days
Who decidesSponsor signs the release; the owner runs it
Model swap · evaluating
reasoner-large v2 behind the same route
- 1New model behind the same routeNo code change: the route points at the new model in the sandbox.0 h
- 2Eval suite re-run: 218 casesTwo regressions on “ambiguous contract term”, fixed with one prompt line and re-run.day 1
- 3Rolled out at L3 with sign-offAudit log records the model version; the old route stays warm for a rollback.day 2
Model swap · rolled out
Re-evaluated and live in two days
- Cases re-run
- 218 / 218
- Regressions
- 2 found · 2 fixed
- Policy violations
- 0
- Cost per task
- $0.021 → $0.017
Model version recorded in the audit log · old route kept warm for rollback
Sandbox on day one
Masked copies of your data in an isolated environment, so nothing waits on a production connection.
A reference platform
MCP servers, test suite, automatic checks, audit log and tracing, ready to configure for your systems.
Test results decide
Every gate is a number from the tests, so go or no-go is a dashboard check with the owner.
08The workshop
AI strategy with the people who run the work.
Two days in the first week with operations, legal, security, finance and the team that does the work. Your AI ideas are scored and become a one-page strategy the room has agreed.

The week before · a sourced pre-readAgents summarise your process documents, a sample of tickets and the system inventory into a short pre-read, every line linked to its source, so the workshop is spent deciding.

-
Day 1 · Morning
How the work runs today
Three to five processes walked end to end by the people who do them.
Process maps with baseline times
-
Day 1 · Afternoon
The long list, scored
Every idea on the wall, scored on value, feasibility, risk and time.
A scored long list, usually a dozen

-
Day 2 · Morning
Risk and data, per candidate
Legal sets the tier, security the scopes, system owners the data readiness.
Risk tier, autonomy ceiling, data gaps
-
Day 2 · Afternoon
Pick the first three
Three to build first, each with an owner, a baseline and a target.
The one-page strategy, signed
AI strategy · one page · Your company · v1Illustrative
What leaves the room
-
1Ambition and measures
Return about 700 hours a month to sales ops and finance by the end of Q2; cost per task under a quarter of manual.
-
2Portfolio · build first
Reconcile vendor invoices · Draft renewal quotes · Triage inbound RFPs. Nine more sequenced by quarter.
-
3Operating model
Embedded: agents owned by the teams that run the process, with a small central platform and review group.
-
4Data readiness gaps
Billing history split across two systems · 18% of invoices arrive as PDFs · CRM firmographics patchy.
-
5Budget and measures
Build, run and review budget per use case; one owner and one number each; quarterly review.
-
6Risk tier and governance
Two limited-risk use cases carry disclosure duties; two high-risk candidates parked with a compliance plan.
Agreed in the room
- COO
- CFO
- CISO
- Head of sales ops
09How we work with you
Assess, prove and scale, with a gate after each phase.
Work moves to the next phase only when the numbers meet criteria agreed in week one and the named people sign.
-
013 weeks
Assess
Interviews and the two-day workshop in week one, then a prototype on your data and an evaluation baseline for the first use case.
Leaves with
- Scored use-case portfolio
- One-page AI strategy and roadmap
- Business cases for the first three
- Prototype and evaluation baseline
Autonomy reachedL1 · SuggestA prototype suggests, in a sandbox on masked data
G1 Go to prove Passed
- A named owner, a measured baseline and an evaluation baseline for the first use case
- Data access agreed at read scope, sandbox ready
- Risk tier set and the autonomy ceiling written down
- Budget approved, with a cost-per-task ceiling
Signed bySponsor · owner · legal
-
025 weeks
Prove
One agent, tested on your experts’ cases, piloted with ten users at L2 from week 5 and released at L3 in week 8, with approval on every write.
Leaves with
- Working agent and its MCP servers
- Test suite, evaluation set, red-team results
- Check and approval policy
Autonomy reachedL3 · Act with approvalActs on your systems once a named person approves
G2 Go to production Passed
- Test gate passed on the evaluation set, adversarial cases included
- Zero policy violations across every run
- Cost per task below the ceiling, 95th-percentile time on target
- Check policy signed by legal and security
- Runbooks and on-call handed over
Signed bySponsor signs · owner runs
-
03Ongoing
Scale
The next agents from the portfolio, a risk register kept current, and autonomy raised one task at a time on evidence.
Leaves with
- Next agents from the roadmap
- AI inventory and risk register kept current
- Quarterly review against the measures
Autonomy reachedL4 · Act within limitsFor named tasks, inside limits, after G3
G3 Raise autonomy Every 90 days
- 90 days at L3 with approvers changing under 2% of actions
- Escalations answered within the agreed time
- No incident traced to the agent
- One task moves up one level, with new limits
Signed byOwner · risk · security
10What you get
Strategy, agent, governance and runbooks, in your accounts.
What we make for you is yours once it is paid for. Tools we already had stay ours, and you get a free, permanent licence to use them.
Tab 1 of 4Your company · AI programme · handover
Strategy
Where AI fits, what comes first and what each use case is worth.
-
ai-strategy-one-page.pdfOne-page AI strategy Document -
use-case-portfolio.xlsxAI opportunity map and use-case portfolio BoardSheet -
roadmap-business-cases.pptxAI roadmap with business cases DeckSheet -
operating-model.pdfOperating model, roles and approval rights Document
Handover checklist
- Signed by the sponsor and the three owners
- Source files kept in your document store
- Portfolio scores editable, formula included
Tab 2 of 4Your company · AI programme · handover
Agent
The agent, in your repository and cloud account, with the tests it must pass.
-
agents/renewals-agent/Agent or copilot in production Sourceyour cloud -
mcp-servers/MCP servers for the CRM, billing and documents Source -
evals/Evaluation suite, test cases, red-team results Testsreport -
infra/Deployment as code for your cloud account Terraform
Handover checklist
- Repository transferred to your organisation
- CI, secrets and cloud account in your name
- Prompts and model routes versioned in the repo
Tab 3 of 4Your company · AI programme · handover
Governance
The written answers legal, security and auditors ask for, kept in one register.
-
guardrails/policy.yamlCheck and approval policy Policyconfig -
ai-governance-framework.pdfAI governance framework Document -
ai-inventory-risk-register.xlsxAI inventory and risk register Register -
model-cards/Model cards per model and version Document
Handover checklist
- Register owner named, review dates booked
- Legal and security sign-off recorded
- Every row linked to its evidence
Tab 4 of 4Your company · AI programme · handover
Operate
Everything your team needs to run, review and change the agent without us in the room.
-
dashboards/agent-actionsAgent action and audit log Dashboard -
runbooks/Runbooks: escalations, rollback, model swap Document -
training/Training for owners, approvers and reviewers Sessionsrecordings -
reviews/q-review-template.pdfQuarterly review pack Report
Handover checklist
- On-call rota and escalation contacts agreed
- Approvers trained and named in the policy
- First quarterly review in the calendar
Each tab shows a mock of the first artefact in that part of the handover: a roadmap by quarter for strategy, the repository layout for the agent, the AI inventory and risk register for governance, and an escalation runbook for operations.
1190-day review
Six measures decide whether an agent stays.
Every agent is reviewed on six measures, each with a pass line and a hold line. An agent that misses them goes back a level. The lines shown are review targets.
Hours returned per week
Handling time saved × volume, against the baseline taken before build
74 h
Pass
Task success on evals
The golden set, re-run on every change and at each review
94.2%
Pass
Cost per task vs manual
Models, tools, hosting and review time per completed task, as a share of manual cost
13% · $0.62 vs $4.80
Pass
Escalation rate
Share of runs handed to a person, by reason
6.8%
Pass
Policy violations
Writes outside policy, from the audit log and the red-team suite
0
Pass
Carbon per task vs manual
SCI estimate per completed task, as a share of the manual process on the same grid
19% · 0.9 vs 4.8 Wh est.
Pass
Decision
Stays at L3L3 · Act with approval
Review in 90 days whether quotes with discounts of 5% or less can move to L4, inside limits.
Signed · ownerReviewed · risk and securityLogged · audit trail
Illustrative reviewLines are agreed per agent before build. Cost includes the time people spend approving and reviewing.
Services & packages
Decide where AI pays off, then build the agents that do the work.
We rank your AI use cases, then design and build the agents and copilots, with the governance to run them. Each is tested on your own examples and overseen by your people.
- Categories
- 04
- Services
- 15
- Packages
- 04
How to buy
- 01Pick services. Enquire about one, or add several to a brief.
- 02Choose a package. A sprint, a fixed project or an ongoing team.
- 03Send the brief. We reply within one working day.
How we work with you
Ways to engage, from a question to an RFQ.
Ask a quick question, send a project brief or issue a formal RFQ. The lead for the work reads each one in full, and any services already in your brief go with it.
Or book a thirty-minute call-
01
For a first conversation, a press request or anything that does not need a scope yet.
You get A reply from a named lead
-
02Recommended
Goals, audiences, a budget band and timing, so our first reply can outline the work.
You get Options and a first scope after one call
-
03
Your documents, deadlines and the procurement and security rules the work must meet.
You get Receipt confirmed and a named bid lead
How it is priced
Each package shows its pricing model. Work starts once a written scope and quote are agreed.
Opens a project brief with this package chosen.
-
In your brief
- Typical length
- 1–3 weeks
- Pricing
- Fixed fee
-
In your brief
- Typical length
- 4–12 weeks
- Pricing
- Fixed price
-
In your brief
- Typical length
- 3–9 months
- Pricing
- Fixed price per milestone
-
In your brief
- Typical length
- 6–18 months
- Pricing
- Programme fee · by statement of work
Compare what each package includes
| Package | Every engagement includes | Best for |
|---|---|---|
| SprintA short, fixed-scope engagement that answers one defined question. |
|
Discovery, a diagnostic, a prototype or a decision you need to make soon |
| ProjectA defined scope, delivered for a fixed price. |
|
Work you can describe up front: an identity, a platform or a set of tools |
| MilestoneA larger build in phases you approve and pay for one at a time. |
|
Programmes too large for one contract, where you want control at each step |
| EnterpriseA multi-workstream programme with a dedicated team, governance and agreed service levels. |
|
Large organisations running change across markets, portfolios or business units |
12Questions
AI Strategy & Agents, asked directly.
The questions leadership, legal and security ask in the first meeting, answered plainly.
Bring to the first call
- The processes you want to change, with rough weekly volumes
- Who owns each one, and who approves changes to it
- The systems and data each process touches
- Your risk appetite, and any regulator in scope
Start with one high-volume process with a clear owner and reachable data. We score your long list, then pilot the top candidate at L1 or L2, where the agent suggests or drafts. Back-office tasks with clear rules, such as invoice matching or quote drafting, usually make better first candidates than open-ended customer chat.
Models from OpenAI, Anthropic, Google, Mistral and open-weight families such as Llama, chosen per task on quality, speed, cost and data residency. We call them through one gateway, so you can switch when the evidence changes. Within one agent, models are chosen per step: a small, fast model routes and extracts, and a larger one is used only where the reasoning pays for itself.
Only at L4, for named tasks and inside limits you set. Limits cover spend per run, records touched and a business limit such as a maximum discount; anything outside them goes to a person. A task reaches L4 only after 90 days at L3 with approvers rarely changing its actions, and the decision is recorded in the register.
Mainly through permissions and approvals, since no filter catches every attack. All model inputs, including documents and emails, are treated as untrusted. Tools are allow-listed per level, each agent has its own scoped credentials, every write is policy-checked outside the model and high-impact actions wait for a person. Each release is red-teamed against OWASP LLM01 (prompt injection) and LLM06 (excessive agency).
Three things: model usage, the platform around it and people’s review time. The platform covers hosting, vector store and tracing. With a small model routing and context cached, model usage for a back-office task is often a few cents, depending on document length and the models chosen. We measure cost per task from the first evaluation baseline against the manual cost.
Both can apply, and we map them in one register. The EU AI Act covers AI placed on the EU market or used in the EU, or whose output is used there. Most business agents are minimal or limited risk, with transparency duties from August 2026. Duties for high-risk Annex III uses, such as recruitment or creditworthiness checks, apply from December 2027 under the 2026 amendment. The DPDP Act 2023 covers personal data of people in India; under the DPDP Rules 2025, most duties apply from May 2027. This is engineering guidance, not legal advice.
Tell us what you need built.
You will speak to a lead who would run the work, and get a straight answer on fit.
Three ways to start
Every engagement starts with a written scope and a quote agreed before work begins.
Choose one of the three ways above


