AI opportunity mapping
The journey read for the moments where AI removes real effort, and the ones where it only adds another box to type in.
AI Product Strategy & Development
AI Product Strategy & Development decides where AI belongs in your product and builds it: the interaction patterns, the disclosure, the controls and the evaluation loop that decide whether an AI feature is still used in week four.
An illustrative AI feature specification for drafting support replies: three model options measured against five evaluation thresholds, with the fallbacks that handle each gap.
What it covers · 6 parts
Every feature is designed around its failure modes first: what the person sees when confidence is low, how they correct it, and what the product learns from the correction.

The journey read for the moments where AI removes real effort, and the ones where it only adds another box to type in.
Streaming, citations, confidence, suggestion against action, undo and escalation to a person: chosen per task and added to your design system.
Prototypes wired to real models and your own content, so the concept is judged on what the model actually returns rather than on a scripted demo.
Task-level eval sets agreed with your experts, plus preference and trust testing with the people who will use the feature.
Clear disclosure that output is AI-generated, sources shown, results editable, an obvious way back, and a log of what the system did.
The feature built with your engineers: retrieval or tool plumbing, guardrails, human approval steps, feedback capture and quality monitoring.
How it runs
The prototype meets a real model in week two. What it gets wrong there shapes the interface, not the launch note.
The task, the people doing it today, the cost of a wrong answer and the data available. Success measures and a first evaluation set drafted with your experts.
Gate · Is AI the right tool for this task?
Interaction patterns drawn and wired to live models on real content, then tested for usefulness, trust and what people do when the answer is wrong.
Gate · Does the prototype clear the eval set?
Built with engineering: retrieval or tools, guardrails, approval steps and feedback capture, with evaluation thresholds gating the release.
Gate · Are guardrails and approvals in place?
Staged rollout, adoption and correction rates watched, the eval set grown from real traffic, and the patterns fed back into the design system.
Gate · Is it still used, and still correct?
Timings are typical and shorten when the evidence already exists.
Eval scorecard · what you will see
The scorecard each release is held to: the threshold agreed with you, how each model option scores on your own test set and what carries the gap.
| Check | Threshold | Option A · large model | Option B · small model + retrieval | Fallback |
|---|---|---|---|---|
| Answer grounded in the help centre | ≥ 95% | 97.1% (Passes) | 95.4% (Passes) | Cite or decline |
| Declines out-of-scope requests | ≥ 98% | 98.6% (Passes) | 96.2% (Fails) | Route to a person |
| Prompt-injection suite (OWASP LLM01) | 0 passes through | 0 of 60 (Passes) | 1 of 60 (Fails) | Block and log |
| Personal data in output | 0 cases | 0 (Passes) | 0 (Passes) | Redact |
| Latency, p95 | ≤ 2.5 s | 3.1 s (Watch) | 1.4 s (Passes) | Stream the draft |
What you keep
Everything needed to ship the feature and keep it honest: the spec, the eval set it must pass and the fallback when it cannot.

/handover/04-ai-product/
/handover/README
What changes
AI features earn their place when people keep using them, low-confidence answers are caught and the controls exist on day one. Baseline, target and instrument are agreed in week one.
Adoption measured after the launch spike, alongside correction, abandonment and escalation rates.
Low confidence, missing sources and refusals designed as states, so a mistake is recoverable rather than alarming.
Risk tier, disclosure, human oversight and logging decided during design, which is where transparency duties are cheapest to meet.
The stack
Model providers, orchestration and product tools we use for AI features. Your accounts, your keys. Technologies we work with, not partnerships.
Frameworks we build to
Frameworks every AI feature is classified, tested and documented against. Frameworks we build to, not certifications we hold.
A risk-based regime: prohibited practices, obligations for high-risk systems, transparency duties and rules for general-purpose AI models.
How we apply itEach use case classified by risk tier early, with transparency notices, logging and human oversight designed in.
Four functions for trustworthy AI: Govern, Map, Measure and Manage, with a companion profile for generative AI (NIST AI 600-1).
How we apply itRisks mapped per use case, measured with evals, and managed with named owners and release thresholds.
Requirements for establishing, running and improving an AI management system: AI policy, impact assessment, data and lifecycle controls.
How we apply itAn AI inventory, impact assessments and lifecycle controls are part of how every model and agent is shipped.
Risks specific to LLM systems, including prompt injection (LLM01), sensitive information disclosure (LLM02), excessive agency (LLM06) and vector and embedding weaknesses (LLM08).
How we apply itRed-team suites for prompt injection, data leakage and tool misuse run in CI before any model or agent release.
Perceivable, operable, understandable and robust content, including 2.2 criteria such as focus not obscured, target size and accessible authentication.
How we apply itContrast, focus order and target size checked in the design file, then keyboard and screen-reader passes on every key journey before sign-off.
Lawful basis, data-subject rights, data protection by design and by default, breach notification and DPIAs for high-risk processing.
How we apply itResearch participants give recorded, specific consent; recordings and notes have a retention date and are deleted on it.
Notice and consent, duties of data fiduciaries, rights of data principals, breach intimation and added duties for significant data fiduciaries, with the DPDP Rules.
How we apply itConsent notices for India-based participants and users, with data-principal requests answerable from the research log.
Services & packages
Commission a short opportunity read, a prototype on live models, or the design and build of an AI feature inside your product. Every engagement is judged on real model output and tested with the people who will use it.
How to buy
Start a project
Three ways in, from a two-minute question to a formal RFQ. Each is read in full by the lead for the work, and anything already in your brief goes with it.
Or book a thirty-minute call01
For a first conversation, a press request, or anything that does not need a scope yet.
You get A reply from a lead, not a sales queue
02Most useful
Goals, audiences, a budget band and timing. Enough for us to come back with a shape, not only questions.
You get Options and a first scope after one call
03
Your pack, your deadlines, and the procurement and security rules the work must meet.
You get Receipt confirmed and a named bid lead
How it is priced
Each package shows how it is priced. Every engagement starts with a written scope and a quote agreed before work begins.
Opens a project brief with this package chosen.
In your brief
In your brief
In your brief
In your brief
| Package | Every engagement includes | Best for |
|---|---|---|
| SprintOne fixed question, answered in one to three weeks. |
|
Discovery, a diagnostic, a prototype or a decision you need to make soon |
| ProjectA defined scope, delivered for a fixed price. |
|
Work you can describe up front: an identity, a system, a set of tools |
| MilestoneA larger build, split into gated phases you approve and pay for one at a time. |
|
Programmes too big to fix in one contract, and teams that want control at each step |
| SquadA dedicated team working inside your stack, tools and sprint cadence. |
|
Teams with a clear roadmap that need more senior hands, fast |
Questions
Anything else goes straight to the people who would do the work.
Ask a questionThe same systems, approached from opposite ends. Technology & Intelligence engineers the models, retrieval, agents and infrastructure. This capability designs the product around them: where AI belongs in the journey, what the person sees, and how the feature is evaluated with real users. Most AI products need both, and the two teams work as one.
Saying plainly that it is AI, showing where the answer came from, making the result editable, keeping an obvious undo, and never taking a consequential action without a person approving it. Trust is a set of interface decisions, not a tone of voice.
Yes, where a brand voice exists. Tone, vocabulary and prohibited phrasing become part of the system prompt and part of the evaluation set, so the output is checked against them rather than hoped for.
With eval sets rather than fixed expected strings: graded criteria per task, scored automatically and calibrated against human review, run over many inputs so the result is a distribution rather than a single pass or fail. Preference and trust tests with users sit on top of that.
It can, including for organisations outside the EU when the system is placed on the EU market or its output is used there. Most product features sit in the transparency tier, which asks mainly that people are told they are interacting with AI and that generated content is marked. We classify the use case early and design the notices in rather than bolting them on.
Keep going
Experience Design makes the interaction clear; Strategy decides where AI is worth it at all.
03 · Pairs well
Designed, tested, built
Experience Design & Development
Fast, rigorous iterations that shape and validate an experience before big bets get made.
Explore Experience
02 · Pairs well
What to build next
Product Strategy & Vision
Spotting the product experiences worth building next to meet real business ambitions.
Explore Strategy
The discipline
From signal to shipped
Product & Experience Design
All five capabilities, the loop they share and how an engagement moves through it.
Back to the overview
Tell us what you’re building. We’ll answer straight.
Three ways to start
Every engagement starts with a written scope and a quote agreed before work begins.
Choose one of the three ways above