Capability 04 of 05 · AI people keep using

AI Product Strategy & Development

The model is not the product. The experience around it is.

AI Product Strategy & Development decides where AI belongs in your product and builds it: the interaction patterns, the disclosure, the controls and the evaluation loop that decide whether an AI feature is still used in week four.

Typical length
4-week concept · 12-week build
Approach
Design-led AI delivery
Standard
Evaluated with users
AI feature spec · Draft reply assistantIllustrative

An illustrative AI feature specification for drafting support replies: three model options measured against five evaluation thresholds, with the fallbacks that handle each gap.

Small + retrievalSmall model with retrieval: all five thresholds met. Ships behind a flag to 10% of agents, with the fallbacks below.

What it covers · 6 parts

Designed for the times the model is wrong

Every feature is designed around its failure modes first: what the person sees when confidence is low, how they correct it, and what the product learns from the correction.

Two people reviewing a product demo on a laptop
01 / 06

AI opportunity mapping

The journey read for the moments where AI removes real effort, and the ones where it only adds another box to type in.

Journey-led
02 / 06

AI interaction patterns

Streaming, citations, confidence, suggestion against action, undo and escalation to a person: chosen per task and added to your design system.

Patterns · design system
03 / 06

Prototyping on live models

Prototypes wired to real models and your own content, so the concept is judged on what the model actually returns rather than on a scripted demo.

Real outputs
04 / 06

Evaluation with users

Task-level eval sets agreed with your experts, plus preference and trust testing with the people who will use the feature.

Evals · preference tests
05 / 06

Trust, transparency & control

Clear disclosure that output is AI-generated, sources shown, results editable, an obvious way back, and a log of what the system did.

Disclosure · audit log
06 / 06

AI feature delivery

The feature built with your engineers: retrieval or tool plumbing, guardrails, human approval steps, feedback capture and quality monitoring.

Shipped · monitored

How it runs

A concept in four weeks, a feature people keep using

The prototype meets a real model in week two. What it gets wrong there shapes the interface, not the launch note.

  1. Frame Wk 01–02

    The task, the people doing it today, the cost of a wrong answer and the data available. Success measures and a first evaluation set drafted with your experts.

    • Use-case brief
    • Eval set v1
    • Risk classification

    Gate · Is AI the right tool for this task?

  2. Prototype Wk 02–04

    Interaction patterns drawn and wired to live models on real content, then tested for usefulness, trust and what people do when the answer is wrong.

    • Live prototype
    • Test findings
    • Pattern decisions

    Gate · Does the prototype clear the eval set?

  3. Build Wk 04–12

    Built with engineering: retrieval or tools, guardrails, approval steps and feedback capture, with evaluation thresholds gating the release.

    • Production feature
    • Guardrails
    • Eval gate

    Gate · Are guardrails and approvals in place?

  4. Improve Wk 12+

    Staged rollout, adoption and correction rates watched, the eval set grown from real traffic, and the patterns fed back into the design system.

    • Rollout
    • Quality dashboard
    • Pattern updates

    Gate · Is it still used, and still correct?

Timings are typical and shorten when the evidence already exists.

Eval scorecard · what you will see

A feature ships when the evals pass. Not when the demo looks good.

The scorecard each release is held to: the threshold agreed with you, how each model option scores on your own test set and what carries the gap.

Eval scorecard · Draft reply assistant · 420 test casesIllustrative
Eval scorecard · Draft reply assistant · 420 test cases
CheckThresholdOption A · large modelOption B · small model + retrievalFallback
Answer grounded in the help centre≥ 95%97.1% (Passes)95.4% (Passes)Cite or decline
Declines out-of-scope requests≥ 98%98.6% (Passes)96.2% (Fails)Route to a person
Prompt-injection suite (OWASP LLM01)0 passes through0 of 60 (Passes)1 of 60 (Fails)Block and log
Personal data in output0 cases0 (Passes)0 (Passes)Redact
Latency, p95≤ 2.5 s3.1 s (Watch)1.4 s (Passes)Stream the draft
Ship A with streaming; retest B after tuningIllustrative example of an eval scorecard. Thresholds are set per feature with you before any model is compared.

What you keep

An AI feature with its numbers attached. Spec, evals and guardrails.

Everything needed to ship the feature and keep it honest: the spec, the eval set it must pass and the fallback when it cannot.

A hand holds a phone showing an app screen in front of an open laptop and a notebook
The featuremeasured by its evals

/handover/04-ai-product/

AI Product Strategy & Development

  • AI opportunity map across the journeyBoard · Figma
  • AI interaction pattern setFigma · Storybook
  • Prototype on live modelsCoded prototype
  • Evaluation set & resultsDatasets · report
  • Trust & disclosure guidelinesDocument
  • Production AI featureSource · your cloud
  • Adoption & quality dashboardDashboard

/handover/README

In every handover

  • Decision record, with evidence linksDoc
  • Assumption tracker, final statesSheet
  • Research consent & retention logSheet
  • Walkthrough session, recordedVideo

What changes

Used, trusted and within its limits. Three measures that prove it.

AI features earn their place when people keep using them, low-confidence answers are caught and the controls exist on day one. Baseline, target and instrument are agreed in week one.

  1. Outcome 01

    Used past the novelty week

    Adoption measured after the launch spike, alongside correction, abandonment and escalation rates.

  2. Outcome 02

    Wrong answers handled well

    Low confidence, missing sources and refusals designed as states, so a mistake is recoverable rather than alarming.

  3. Outcome 03

    Governed from the first sketch

    Risk tier, disclosure, human oversight and logging decided during design, which is where transparency duties are cheapest to meet.

Outcome review · Your companyIllustrative
  • 01 Active users in week 6Product analytics cohort after the launch spike
    46%from 20%target 40%
  • 02 Low-confidence answers routed wellFallback shown or handed to a person, from logs
    95%from 40%target 95%
  • 03 Risk controls in place at launchDisclosure, oversight, logging and eval checks met
    100%from 30%target 100%
Scale 0–100 on every row; row numbers match the outcomes. Figures are illustrative examples of the measures we set, not client results.

The stack

Model-agnostic by design. Chosen per task, swappable later.

Model providers, orchestration and product tools we use for AI features. Your accounts, your keys. Technologies we work with, not partnerships.

  • Figma
  • Anthropic
  • Google Gemini
  • LangGraph
  • Python
  • TypeScript
  • React
  • Next.js
  • Storybook
  • PostHog
  • OpenAI
  • LlamaIndex
  • Playwright

Frameworks we build to

Risk classified before the first prompt. Oversight designed in.

Frameworks every AI feature is classified, tested and documented against. Frameworks we build to, not certifications we hold.

  • EU AI ActArtificial Intelligence Act (EU) 2024/1689

    A risk-based regime: prohibited practices, obligations for high-risk systems, transparency duties and rules for general-purpose AI models.

    How we apply itEach use case classified by risk tier early, with transparency notices, logging and human oversight designed in.

  • NIST AI RMF 1.0AI Risk Management Framework

    Four functions for trustworthy AI: Govern, Map, Measure and Manage, with a companion profile for generative AI (NIST AI 600-1).

    How we apply itRisks mapped per use case, measured with evals, and managed with named owners and release thresholds.

  • ISO/IEC 42001:2023Artificial intelligence management systems

    Requirements for establishing, running and improving an AI management system: AI policy, impact assessment, data and lifecycle controls.

    How we apply itAn AI inventory, impact assessments and lifecycle controls are part of how every model and agent is shipped.

  • OWASP Top 10 for LLM ApplicationsLLM and generative AI security risks

    Risks specific to LLM systems, including prompt injection (LLM01), sensitive information disclosure (LLM02), excessive agency (LLM06) and vector and embedding weaknesses (LLM08).

    How we apply itRed-team suites for prompt injection, data leakage and tool misuse run in CI before any model or agent release.

  • WCAG 2.2 AAWeb Content Accessibility Guidelines

    Perceivable, operable, understandable and robust content, including 2.2 criteria such as focus not obscured, target size and accessible authentication.

    How we apply itContrast, focus order and target size checked in the design file, then keyboard and screen-reader passes on every key journey before sign-off.

  • GDPRGeneral Data Protection Regulation (EU) 2016/679

    Lawful basis, data-subject rights, data protection by design and by default, breach notification and DPIAs for high-risk processing.

    How we apply itResearch participants give recorded, specific consent; recordings and notes have a retention date and are deleted on it.

  • DPDP Act 2023Digital Personal Data Protection Act, 2023

    Notice and consent, duties of data fiduciaries, rights of data principals, breach intimation and added duties for significant data fiduciaries, with the DPDP Rules.

    How we apply itConsent notices for India-based participants and users, with data-principal requests answerable from the research log.

Services & packages

Design the AI feature people keep using.

Commission a short opportunity read, a prototype on live models, or the design and build of an AI feature inside your product. Every engagement is judged on real model output and tested with the people who will use it.

Categories
04
Services
15
Packages
04
Not sure what you need? Describe the problem

How to buy

  1. 01Pick services. Enquire about one, or add several to a brief.
  2. 02Choose how to engage. A sprint, a fixed project or an ongoing team.
  3. 03Send the brief. We reply within one working day.

Browse by category

Timelines are typical. Every quote follows a written scope.

01Find the opportunity

3 services
Typical timeline: 2–4 weeks

AI opportunity mapping

Where AI genuinely removes effort in your product, and where it would only add another box to type in.

What’s included

  • Journey read for the AI moments
  • Value, feasibility and risk score per idea
  • Content and data readiness check
  • Shortlist with a first use case to prototype
  • Journey-led
  • Scored

Best forProduct teams under pressure to add AI without a clear reason yet.

Typical timeline: 2–3 weeks

AI concept sprint

Two weeks from idea to a tested concept: patterns drawn, a rough prototype on a live model, and sessions with the people it is for.

What’s included

  • Framing of the task and its failure modes
  • Concept and interaction patterns
  • Rough prototype on a live model
  • Customer sessions and a recommendation
  • 2 weeks
  • Live model
  • OpenAI

Best forTeams with an AI idea and no evidence either way.

Typical timeline: 2–4 weeks

AI feasibility & data check

Whether the content, data and permissions exist to make the feature work, checked before a team is committed to building it.

What’s included

  • Content and data inventory for the use case
  • Quality, coverage and permission assessment
  • Retrieval or model approach options
  • Cost-per-request estimate with assumptions
  • Data readiness
  • Cost model
  • LlamaIndex
  • pgvector

Best forTeams whose AI plan assumes the content is better than it is.

02Design the experience

4 services
Typical timeline: 4–8 weeks

AI feature design

The full design of an AI feature, starting from what happens when the model is wrong, uncertain or refuses.

What’s included

  • Task and failure-mode analysis
  • Interaction design including low-confidence and refusal states
  • Disclosure, citation and correction flows
  • Handover with acceptance criteria for engineering
  • Failure states first
  • Handover-ready

Best forTeams with model access and no design for what surrounds it.

Typical timeline: 4–8 weeks

AI interaction pattern set

Reusable patterns for AI in your products: streaming, citations, confidence, suggestion against action, undo and escalation to a person.

What’s included

  • Pattern set with usage guidance
  • Disclosure and transparency rules
  • Components added to your design system
  • Review checklist for every new AI feature
  • Design system
  • Reusable

Best forOrganisations where several teams are shipping AI features at once.

Typical timeline: 3–6 weeks

Prototype on live models

A working prototype wired to real models and your own content, so the idea is judged on what the model actually returns.

What’s included

  • Interaction patterns designed for the task
  • Prototype on live models and real content
  • Sessions with the people who would use it
  • Findings, costs and a build or stop recommendation
  • Live models
  • Tested
  • OpenAI

Best forTeams deciding whether an AI feature deserves a quarter of engineering.

Typical timeline: 3–5 weeks

AI voice & tone guidelines

How the product should sound when it is the one speaking: register, hedging, refusals, apologies and the phrasing it must never use.

What’s included

  • Voice rules written for prompts, not posters
  • Refusal, uncertainty and error phrasing
  • Prohibited phrasing and claim rules
  • Checks added to the evaluation set
  • Voice
  • In the eval set

Best forBrands whose assistant sounds like a different company.

03Prove it

4 services
Typical timeline: 3–6 weeks

Evaluation set & scoring

A graded test set built with your experts, so quality is a number that moves rather than an impression after a demo.

What’s included

  • Task list and grading criteria with your experts
  • Scored eval set with automated runs
  • Calibration against human review
  • Release thresholds and a regression run in CI
  • Evals
  • Release gate
  • OpenAI

Best forTeams shipping AI changes with no way to tell if quality dropped.

Typical timeline: 2–4 weeks

Trust & preference testing

Sessions that show whether people believe the output, notice when it is wrong and know what to do next.

What’s included

  • Test design around trust and correction
  • Sessions with the intended users
  • Preference tests between approaches
  • Design changes prioritised from what they did
  • With users
  • Trust

Best forFeatures where a wrong answer accepted quietly would cost something.

Typical timeline: 2–4 weeks

Transparency & disclosure review

A review of what your AI features tell people: that it is AI, where the answer came from, what was stored and how to reach a human.

What’s included

  • Review of disclosure across the journey
  • Citation, sourcing and data-use statements
  • Route to a person where it is missing
  • Prioritised changes with copy written
  • Disclosure
  • Sourcing

Best forProducts that shipped AI quickly and want to check what it tells people.

Typical timeline: 3–5 weeks

AI use-case risk classification

Each AI use case classified by risk tier against the EU AI Act and NIST AI RMF, with the obligations that follow written as design requirements.

What’s included

  • Inventory of AI use cases in the product
  • Risk tier and reasoning per use case
  • Transparency, oversight and logging requirements
  • Design and engineering backlog to meet them
  • EU AI Act
  • NIST AI RMF

Best forProducts with EU users or an internal AI governance requirement.

04Build & run

4 services
Typical timeline: 8–14 weeks

AI feature build

The feature shipped inside your product, with guardrails, human approval for consequential actions and quality watched after launch.

What’s included

  • Retrieval or tool integration with your systems
  • Guardrails and approval steps
  • Evaluation thresholds gating the release
  • Feedback capture and a quality dashboard
  • Shipped
  • Human in the loop
  • OpenAI
  • LlamaIndex

Best forProducts ready to put an AI feature in front of real customers.

Typical timeline: 8–14 weeks

In-product assistant

An assistant inside your product that answers from your own content and can act on your systems, with sources shown and a clear way back to a person.

What’s included

  • Grounded answers with citations
  • Permission-aware access to your content
  • Actions with confirmation and undo
  • Escalation path to a human with context
  • Grounded
  • Cited
  • Escalation
  • LlamaIndex
  • pgvector
  • OpenAI

Best forProducts whose users search documentation more than they use the product.

Typical timeline: Ongoing · monthly

Adoption & quality monitoring

What happens after launch: adoption past the novelty week, correction and abandonment rates, and an eval set that grows from real traffic.

What’s included

  • Adoption, correction and escalation measures
  • Quality dashboard with alerting
  • Eval set grown from production traffic
  • Monthly read and improvement backlog
  • Post-launch
  • Monitored

Best forTeams whose AI feature launched well and has not been measured since.

Typical timeline: Ongoing · monthly

Embedded AI product squad

Product designers, AI engineers and a researcher working inside your team on your AI roadmap.

What’s included

  • Named specialists matched to your roadmap
  • Your repositories, models and cadence
  • Evaluation and guardrail practice built in
  • Knowledge transfer from week one
  • Squad
  • Time & materials
  • OpenAI

Best forTeams with an AI roadmap and too few people who have shipped one.

Your brief

Tick “Add to brief” on any service, choose a package, then continue. Or enquire about one service directly.

Start a project

Ways to engage. Same team, same standard.

Three ways in, from a two-minute question to a formal RFQ. Each is read in full by the lead for the work, and anything already in your brief goes with it.

Or book a thirty-minute call

What are you sending?

  1. 01

    About 2 minutes4 required answers

    For a first conversation, a press request, or anything that does not need a scope yet.

    You get A reply from a lead, not a sales queue

  2. 02Most useful

    About 8 minutes5 short steps

    Goals, audiences, a budget band and timing. Enough for us to come back with a shape, not only questions.

    You get Options and a first scope after one call

  3. 03

    About 15 minutesYour documents attached

    Your pack, your deadlines, and the procurement and security rules the work must meet.

    You get Receipt confirmed and a named bid lead

How it is priced

Each package shows how it is priced. Every engagement starts with a written scope and a quote agreed before work begins.

  • Typical length
    1–3 weeks
    Pricing
    Fixed fee
  • Typical length
    4–12 weeks
    Pricing
    Fixed price
  • Typical length
    3–9 months
    Pricing
    Fixed price per milestone
  • Typical length
    Ongoing · 3-month minimum
    Pricing
    Time & materials
Compare what each package includes
What every engagement package includes, and who it suits
PackageEvery engagement includesBest for
SprintOne fixed question, answered in one to three weeks.
  • Scope and outcome agreed before day one
  • One senior lead and the specialists the question needs
  • A working review every week
  • A decision-ready output, not a status deck
Discovery, a diagnostic, a prototype or a decision you need to make soon
ProjectA defined scope, delivered for a fixed price.
  • Statement of work with deliverables and acceptance criteria
  • A named project lead and a fixed team
  • A shared plan with dated checkpoints
  • Source files and IP transferred on delivery
Work you can describe up front: an identity, a system, a set of tools
MilestoneA larger build, split into gated phases you approve and pay for one at a time.
  • Phases with their own scope, output and sign-off
  • A go or no-go review at every gate
  • Re-planning between phases as you learn
  • Payment tied to accepted milestones
Programmes too big to fix in one contract, and teams that want control at each step
SquadA dedicated team working inside your stack, tools and sprint cadence.
  • Named specialists matched to your roadmap
  • Your tools, your rituals, your backlog
  • Scale the team up or down each month
  • Knowledge transfer built in from week one
Teams with a clear roadmap that need more senior hands, fast

Questions

Asked before we pick a model.

Anything else goes straight to the people who would do the work.

Ask a question
Q01How is this different from the AI work in Technology & Intelligence?

The same systems, approached from opposite ends. Technology & Intelligence engineers the models, retrieval, agents and infrastructure. This capability designs the product around them: where AI belongs in the journey, what the person sees, and how the feature is evaluated with real users. Most AI products need both, and the two teams work as one.

Q02What makes an AI feature feel trustworthy?

Saying plainly that it is AI, showing where the answer came from, making the result editable, keeping an obvious undo, and never taking a consequential action without a person approving it. Trust is a set of interface decisions, not a tone of voice.

Q03Can the output follow our brand voice?

Yes, where a brand voice exists. Tone, vocabulary and prohibited phrasing become part of the system prompt and part of the evaluation set, so the output is checked against them rather than hoped for.

Q04How do you test something that answers differently every time?

With eval sets rather than fixed expected strings: graded criteria per task, scored automatically and calibrated against human review, run over many inputs so the result is a distribution rather than a single pass or fail. Preference and trust tests with users sit on top of that.

Q05Does the EU AI Act affect a feature like this?

It can, including for organisations outside the EU when the system is placed on the EU market or its output is used there. Most product features sit in the transparency tier, which asks mainly that people are told they are interacting with AI and that generated content is marked. We classify the use case early and design the notices in rather than bolting them on.

Let’s build what happens next.

Tell us what you’re building. We’ll answer straight.

Book a discovery call

Three ways to start

Every engagement starts with a written scope and a quote agreed before work begins.

Choose one of the three ways above