01
Designed for the wrong answer
Every screen has a state for uncertain, slow and incorrect. People recover from a bad answer instead of losing trust in the product.
Recovery after a wrong answer
Measured by Moderated sessions on the prompt set
Capability 01 / 04 · Interfaces for systems that guess
AI Application Design shapes what an AI product feels like to use: assistants, agentic flows and multimodal interfaces designed across Gemini, OpenAI, Anthropic and open-weight models. Every screen accounts for the uncertain answer, the slow answer and the wrong one.
Illustration: an assistant conversation. A user asks which enterprise contracts renew before March and whether pricing terms changed. The assistant shows a read-only plan of three tool calls, answers with three numbered citations, says two scanned contracts could not be read, and proposes drafting reminders to three account owners, which waits for the user to approve.
What it covers
Each one is a decision about control: what the system does on its own, what it asks for, and what it shows so a person can tell whether to trust it.
The first message, the questions it asks, the refusals and the recovery. Written as a script and tested with real users before it reaches a model.
Scripted · tested
How a plan is shown, paused, edited and approved. The user sees each step, keeps the stop button and can undo what the agent did.
Plan · approve · undo
Voice, camera, document and screen input designed together, with a clear hand-back to typing when the room is loud or the light is poor.
Voice · vision · document
Citations, sources, confidence, edit history and a plain account of what the system used. Designed so people check rather than assume.
Cited · inspectable
Streaming, progressive disclosure, partial results and honest empty states. What the interface does at four seconds, and what it does when the model is wrong.
Streaming · graceful failure
Task-based testing with real users against a fixed set of prompts and edge cases, scored for completion, correction effort and trust.
Task success · trust
How it runs
A working prototype on a live model, early. Paper cannot tell you how a probabilistic system feels. Timings are typical, set per engagement.
Users, jobs to be done, the data the system may see and the actions it may take. Failure modes named before features.
Flows, conversation scripts and interface design built in the design system, prototyped against a live model rather than a static mock.
Task-based sessions with real users on real prompts, including the cases where the model is wrong. The design changes on what they do, not on what they say.
Specifications, tokens and components handed to engineering, with the prompt set and the acceptance criteria attached.
In practice · failure-state catalogue
Every way the assistant can fail gets a designed state: what triggers it, what the person sees, how they recover, and the test that proves it on the live model.
| State | Trigger | What the person sees | Recovery | Tested |
|---|---|---|---|---|
| No source found | Retrieval under threshold | Says so, offers search | One tap to a person | 41 / 41 |
| Low confidence | Grader score < 0.7 | Answer with caveat + sources | Ask a follow-up | 38 / 40 |
| Action needs approval | Refund over limit | Plan shown, approve or edit | Named approver | 25 / 25 |
| Prompt injection attempt | Instruction in a document | Ignores it, flags the file | Logged for review | 18 / 20 |
| Out of scope | Legal or medical ask | Declines, explains why | Hand-off with context | 9 / 12 |
Illustrative Taken from a prototype. Your catalogue is written against your own prompt set and policies.
What changes
Three measures taken on the live-model prototype, baselined before design starts. Typical targets, not promises; yours are set per engagement.
01
Every screen has a state for uncertain, slow and incorrect. People recover from a bad answer instead of losing trust in the product.
Recovery after a wrong answer
Measured by Moderated sessions on the prompt set
02
Plans are visible, consequential actions are approved, and anything the system did can be undone.
Agent actions approved first time
Measured by Approval events in the prototype log
03
Decisions tested with real users on a live model, so the build starts from what already worked.
Task success on the prompt set
Measured by Tasks completed ÷ tasks attempted
What you keep
Conversation design, the failure-state catalogue, the prompt set and the tested prototype live in your repositories and design files from the first week.

Manifest · AI Apps7 items
| # | Deliverable | Format |
|---|---|---|
| 01 | AI experience brief & capability map | Doc · board |
| 02 | Interaction design & flows | Figma |
| 03 | Conversation & prompt scripts | Doc · prompt set |
| 04 | Working prototype on a live model | Prototype · repo |
| 05 | AI interface pattern library | Figma · Storybook |
| 06 | Experience evaluation report | Report · recordings |
| 07 | Design specification & acceptance criteria | Doc · tickets |
Where it sits
AI Application Design owns what people meet. The acceptance criteria it writes become the tests engineering builds against.
Which use case, and why this one first.
The conversation, controls and failure states, tested on a live model.
Agents, retrieval and infrastructure built to the specification.
The same prompt set, re-run on every release.
The prompt set written in week one is the thread through all four.
Technologies we work with
The design, model and prototyping tools AI application work runs on, most used first. Chosen on your prompt set; nothing locks you to a vendor.
Also in use
Frameworks we build to
The frameworks that shape the guardrails, the failure states and the records we hand over. We build to them; they are not certifications we hold.
Perceivable, operable, understandable and robust content, including 2.2 criteria such as focus not obscured, target size and accessible authentication.
A risk-based regime: prohibited practices, obligations for high-risk systems, transparency duties and rules for general-purpose AI models.
Four functions for trustworthy AI: Govern, Map, Measure and Manage, with a companion profile for generative AI (NIST AI 600-1).
Risks specific to LLM systems, including prompt injection (LLM01), sensitive information disclosure (LLM02), excessive agency (LLM06) and vector and embedding weaknesses (LLM08).
Lawful basis, data-subject rights, data protection by design and by default, breach notification and DPIAs for high-risk processing.
Notice and consent, duties of data fiduciaries, rights of data principals, breach intimation and added duties for significant data fiduciaries, with the DPDP Rules.
Services & packages
Buy the design of one AI surface, such as an assistant, an agentic flow or a voice interface, or the whole experience from concept through to a specification your engineers can build. Everything is prototyped on a live model and tested with real users.
How to buy
Start a project
Three ways in, from a two-minute question to a formal RFQ. Each is read in full by the lead for the work, and anything already in your brief goes with it.
Or book a thirty-minute call01
For a first conversation, a press request, or anything that does not need a scope yet.
You get A reply from a lead, not a sales queue
02Most useful
Goals, audiences, a budget band and timing. Enough for us to come back with a shape, not only questions.
You get Options and a first scope after one call
03
Your pack, your deadlines, and the procurement and security rules the work must meet.
You get Receipt confirmed and a named bid lead
How it is priced
Each package shows how it is priced. Every engagement starts with a written scope and a quote agreed before work begins.
Opens a project brief with this package chosen.
In your brief
In your brief
In your brief
In your brief
| Package | Every engagement includes | Best for |
|---|---|---|
| SprintOne fixed question, answered in one to three weeks. |
|
Discovery, a diagnostic, a prototype or a decision you need to make soon |
| ProjectA defined scope, delivered for a fixed price. |
|
Work you can describe up front: an identity, a system, a set of tools |
| MilestoneA larger build, split into gated phases you approve and pay for one at a time. |
|
Programmes too big to fix in one contract, and teams that want control at each step |
| RetainerReserved monthly capacity to run, improve and extend what we built. |
|
Brands and products after launch that need a steady team without hiring one |
Questions
What buyers ask before AI Application Design work. Anything else, ask the team directly.
Ask the teamThat discipline engineers AI into business systems: agents, retrieval, infrastructure and production evaluation. AI Design shapes the experience people meet — the conversation, the controls and the way trust is earned on screen. On larger programmes the two run together, with design setting the acceptance criteria and engineering meeting them.
The interface is designed to be model-agnostic. We prototype on the models that suit the task, including Gemini, OpenAI, Anthropic, Mistral and open-weight families, and design so that changing model is a configuration decision rather than a redesign.
By treating it as a normal state rather than an edge case. Sources and confidence are shown, a correction is one action away, consequential steps need approval, and there is always a route to a person. Failure states are designed in the same sprint as the happy path.
Yes. Most engagements end in a specification, components and a prompt set your team builds against. Where you need hands, engineers can join the build through Technology & Intelligence.
It has to be. We design to WCAG 2.2 AA: streaming output announced to assistive technology, keyboard control of every agent action, no meaning carried by motion alone, and a typed equivalent for every voice interaction.
Also in AI Design
Each shares the models, evaluation sets and approval rules built here.
All of AI Design
04 / 04 · Adoption that survives the pilot
Bringing AI into the brand and marketing ecosystem through pilots and adoption roadmaps built to stick.
03 / 04 · The model layer of the brand
Custom-tuned models, via Flux, Adobe Firefly, ComfyUI and more, that keep brand delivery consistent at scale.
02 / 04 · Volume without losing the eye
AI creatives, films and edits — model selection, automation and creative direction, with tools like Runway, Veo and ElevenLabs.
Tell us what you’re building. We’ll answer straight.
Three ways to start
Every engagement starts with a written scope and a quote agreed before work begins.
Choose one of the three ways above