Illustrative imagery

B2B technology20253–5 months

An assistant that answers what sales kept repeating

Client: A B2B software company

One of the project’s screens, in a desktop frame. It is shown again with its caption under “On screen”.

Read the story
  • AI Design
  • Product & Experience Design
  • Technology & Intelligence

At a glance

The brief. Give sales an assistant that answers from approved sources, and says when it does not know.

Client
A B2B software company
Year
2025
Duration
3–5 months
Scale
Website and product

From brief to result

Evals before answers. How the assistant was built.

The brief in full, how we approached it, what we built and what it set out to change.

The brief, in full

Sales engineers answered the same product questions again and again, and the answers drifted: an old limit here, a retired feature there. Prospects wanted answers in the moment, but a wrong answer about security or pricing was worse than none.

The brief: an assistant that answers from approved sources, says when it does not know, and hands over to a person when the question needs one.

Our approach

We built the eval set before the assistant. Real questions from sales calls, each with its approved answer and the source it must cite, became the gate every model and prompt change has to pass.

Guardrails were written as policy first: what the assistant may answer, what it must refuse, and when it routes to sales. Prompt-injection tests (OWASP LLM01) run with every release.

How it ran

  1. Weeks 1–3Questions

    Questions collected from call notes, tagged by topic and by risk.

  2. Weeks 2–6Evals

    The eval set and its rubric: grounded, correct and cited — or a clean refusal.

  3. Weeks 6–14Build

    Retrieval over approved documentation, the assistant interface and the hand-off to sales.

  4. Months 4–5Release gate

    Released only when the eval set passes; injection tests on every change.

What we built

A retrieval assistant grounded in approved product documentation, gated by an eval set, with prompt-injection tests and a hand-off to a person when confidence is low.

What it left behind

  1. Eval set and scoring rubric
  2. Guardrail and escalation policy
  3. Assistant UI in site and product
  4. Conversation audit dashboard

What changed

The assistant is built to answer only from approved documentation, show the page it used, and pass pricing and security questions to a person; every conversation is logged for review.

Accuracy is read against the eval set, not anecdotes, and the figures are the client’s to publish.

How it is measured

Picture story

In the room. Where the questions came from.

The pictures on this page are illustrative: licensed stock photography and code-built mock-ups chosen to show the setting, not the client’s own work. Credits are at the foot of the page.

Three colleagues at a table; one holds up a printed page of pie charts while the others look on.
Fig. 01Questions from sales calls became the eval set.Vitaly Gariev on Unsplash (opens in a new tab)
A smiling woman works at a laptop at her desk, headphones and pens beside her, sticky notes on the wall behind.
Fig. 02Answers in the moment, with the source on screen.Vitaly Gariev on Unsplash (opens in a new tab)
A large open-plan office with rows of desks and monitors, people working in the distance.
Fig. 03Pricing, legal and security questions go to a person, every time.Arlington Research on Unsplash (opens in a new tab)
A man gestures beside a screen while colleagues take notes at white desks in a meeting room.
Fig. 04Release reviews read the eval results, not anecdotes.Vitaly Gariev on Unsplash (opens in a new tab)

On screen

An answer with its sources. And the gate every change must pass.

Illustrative screens, built in code to show the kind of screen the programme produced — not the client’s own. Open any of them full screen.

Illustration: an assistant answers a question about single sign-on with two cited sources and a confidence note, and offers to hand the pricing part of the question to a person.
DesktopAn answer with its sources — and a hand-off for the part that needs a person.
Illustration: an evaluation dashboard for a release, showing pass rates for grounded answers, refusals, escalations and injection tests against their targets.
LaptopThe release gate: every change runs against the eval set first.

Field notes

In detail. Notes from the programme.

What the eval set checks

  • Grounded: every statement traceable to an approved source
  • Cited: the source is shown to the person asking
  • Refuses cleanly when the answer is not in the sources
  • Hands pricing, legal and security questions to sales
  • Ignores instructions hidden in user input (OWASP LLM01)

Outcomes

Judged on the eval set. Not on anecdotes.

What each measure is compared with, when it is read and who reports it — and the results, once the client approves them.

The measurement plan

  1. Answer accuracy on the eval set

    Compared with
    The approved answers in the eval set
    Read
    At every release
    Result
    Awaiting approval Reported by The evaluation lead
  2. Escalation rate to sales

    Compared with
    Questions sales answered by hand before launch
    Read
    Monthly, from the conversation log
    Result
    Awaiting approval Reported by Sales operations
  3. Demo-to-opportunity rate

    Compared with
    The two quarters before launch
    Read
    Quarterly
    Result
    Awaiting approval Reported by The client’s revenue operations team

Results

Figures pending client approval

We publish results only with the client’s written approval. Until then, the plan beside this is what the programme is judged on, and we can walk you through it under NDA.

Credits

The team and the pictures.

People are named only with their agreement.

Team 6 people

  • Product lead × 1
  • AI engineers × 2
  • Conversation designer × 1
  • Front-end engineer × 1
  • Evaluation lead × 1

Credits

Strategy, design and engineering
Xterra Edze
Screens and before-and-after views
Xterra Edze (code-built illustrations)
Photography
Licensed stock from Unsplash, credited picture by picture

Pictures 6

Licensed stock photography and code-built mock-ups, chosen to show the setting — not the client’s own work.

Next case · 04 · Retail & commerce

Visible in AI answers, not only in search

Be cited in AI answers, not only ranked in search results.

Read the next case

Let’s build what happens next.

Tell us what you’re building. We’ll answer straight.

Book a discovery call

Three ways to start

  1. 01About 2 minutes

    A quick question

    You get A reply from a lead, not a sales queue

  2. 02About 8 minutesMost useful

    A project brief

    You get Options and a first scope after one call

  3. 03About 15 minutes

    A formal RFQ or RFP

    You get Receipt confirmed and a named bid lead

Every engagement starts with a written scope and a quote agreed before work begins. How each package is priced