extract.json · doc-ai v3

Posted to ERP · approved by AP clerk

Illustrative extraction of a fictional supplier invoice photographed on a clipboard. Its layout is detected and six fields are extracted with confidence scores from 0.86 to 0.99: supplier, invoice number, GSTIN, line items, total and bank account. The bank account number is redacted before logging. Line items score 0.86, below the 0.90 threshold, because one quantity was corrected in pencil, so an accounts-payable clerk checks them. Once approved, the invoice is posted to the ERP.

Capability 04 of 10Technology & Intelligence

AI features and workflow automation, tested on your own examples before release.

We add AI to your products and operations: answers from your documents, chat and voice assistants, document processing, image inspection and workflow automation. Each feature is tested on your own examples before it launches.

Typical length
8–14 weeks per feature
Scope
Knowledge · chat and voice · vision · automation
Standard
Evaluation-gated releases

04.1What we offer/families

Four kinds of AI feature, built to the same three rules.

Each gets an evaluation set before a launch date is fixed, and a cost per request known from the first prototype. A confidence threshold sends uncertain cases to a person.

  • A long library aisle lined with shelves of bound volumes

    01

    Knowledge assistants

    Answers drawn from your documents, with citations anyone can check and access limited by each user’s permissions.

    Typical uses

    • Policy and product assistants for staff
    • Help centres that answer questions directly
    • Contract search with clause-level citations

    Measured by

    Deflection rate

    Questions resolved without a ticket

    Inspect a cited answer
  • A headset with a microphone rests beside an open laptop on a desk

    02

    Conversational AI

    Chat and voice assistants that resolve requests through your systems and hand over to a person with the full context.

    Typical uses

    • Order, booking and account support on web and WhatsApp
    • Voice agents for first-line calls
    • Agent assist: suggested replies and live summaries

    Measured by

    Average handle time

    Minutes per resolved conversation

    Watch a handover
  • Hands measure a machined metal part with a digital caliper on a workshop bench, an inspection sheet beside it

    03

    Vision & documents

    Extraction, counting and inspection from scans, photos and video, with confidence thresholds and review queues.

    Typical uses

    • Invoice, form and ID extraction into your systems
    • Cap, fill and label checks on a production line
    • Shelf gaps and price-label compliance

    Measured by

    Straight-through processing

    Documents posted with no manual touch

    Compare before and after
  • Hands typing on a laptop showing a configuration table of workflow rules

    04

    Workflow automation

    Multi-step processes across email, documents, CRM and ERP, with retries, audit trails and people at the approval gates.

    Typical uses

    • Accounts payable from inbox to posting
    • Customer onboarding and KYC checks
    • Order exceptions and returns

    Measured by

    Cycle time

    Hours from trigger to done

    Run a workflow

04.2Knowledge assistants/inspector

Answers from your documents, inspected stage by stage.

Pick a question about a fictional company’s documents. The baseline is a typical first build; the tuned version is what we deliver. Open any stage to see what it found, ranked, wrote and scored.

rag-inspector · your-company-kb · 1,284 docs · 18,402 chunks Illustrative

Questions

text
Sentence not supported by any retrieved chunk
text
Supported, but by a superseded or draft source
1
Citation to a chunk the model was given

QuestionHow long do customers have to return a damaged item?

A

Baseline RAG

  • Fixed 1,000-token chunks
  • Vector search only
  • No reranker
  • Citations optional

queryHow long do customers have to return a damaged item?

No rewrite, no metadata filter.

  1. 1warranty-terms.pdf§4 Defectsoff-topic0.84Manufacturing defects are covered for 12 months from purchase…
  2. 2returns-policy-v2.pdf§3 Damaged itemssuperseded0.82…damaged items may be returned within 30 days of purchase…
  3. 3returns-policy-v4.pdf§2 · 1,000-token chunkcut mid-table0.79…14 days of delivery … exchanges … gift cards … store credit…
  4. 4shipping-faq.mdDelivery problems0.77If your parcel arrives damaged, keep the packaging and…
  5. 5careers-page.mdBenefitsoff-topic0.71Staff discount on returns and exchanges…

No reranker. All five chunks go to the model in vector-score order, off-topic chunks included.

Context sent to the model≈ 5,000 tokens

Customers can return a damaged item within 30 days of purchase.2superseded source They should keep the original packaging.4 Refunds reach the customer’s account in 3 to 5 working days.not in sources

Shown to the user as written

Evals

Faithfulness0.71
Answer relevance0.83
Context precision0.58
Context recall0.67

A

Model only

  • No retrieval
  • No sources
  • No citations

queryHow long do customers have to return a damaged item?

Retrieval switched off. The model answers from what it learned in training, with no access to your documents.

Nothing to rank.

Most retailers accept returns of damaged goods within 30 days, and many extend that for defects covered by warranty. Check the retailer’s own policy for details.ungrounded · no sources

Plausible, generic and unverifiable

Evals

Faithfulnessn/a
Answer relevance0.61
Context precisionn/a
Context recalln/a

B

Tuned RAG

  • Semantic chunks · 300–500 tokens
  • Hybrid BM25 + vector
  • Cross-encoder reranker
  • Citations enforced

rewritedamaged item return window · days from delivery · refund or replacement

filterstatus = current · 1 archived document excluded

  1. 1damaged-goods-sop.docx§1 Intake0.86Support logs the order, the photo and the damage type before…
  2. 2returns-policy-v4.pdf§2.3 Photo evidence0.84Damage reported with a photo within 48 hours skips inspection…
  3. 3returns-policy-v4.pdf§2.1 Damaged on arrival0.83Items that arrive damaged can be returned within 14 days of delivery…
  4. 4shipping-faq.mdDelivery problems0.77If your parcel arrives damaged, keep the packaging and…
  5. 5warranty-terms.pdf§4 Defectsoff-topic0.74Manufacturing defects are covered for 12 months from purchase…

  1. 1returns-policy-v4.pdf§2.1 Damaged on arrivalwas #30.97
  2. 2returns-policy-v4.pdf§2.3 Photo evidencewas #20.91
  3. 3damaged-goods-sop.docx§1 Intakewas #10.62
Context sent to the model≈ 1,100 tokens

Chunks scoring under 0.50 are dropped before generation.

Damaged items can be returned within 14 days of delivery for a replacement or a full refund.1 If the damage is reported with a photo within 48 hours, the inspection step is skipped.2

Grounded · every sentence cited · answered

Evals

Faithfulness0.94
Answer relevance0.92
Context precision0.88
Context recall0.90
Eval scores for question 1: column A against column B
EvalsAB
Faithfulness0.71n/a0.94
Answer relevance0.830.610.92
Context precision0.58n/a0.88
Context recall0.67n/a0.90

QuestionDoes the Pro plan include single sign-on?

A

Baseline RAG

  • Fixed 1,000-token chunks
  • Vector search only
  • No reranker
  • Citations optional

queryDoes the Pro plan include single sign-on?

No rewrite, no metadata filter.

  1. 1sso-setup-guide.mdConfigure SAML0.89Admins on paid plans can connect an identity provider using SAML 2.0…
  2. 2security-whitepaper.pdf§5 Identityoff-topic0.85Customer data is protected by encryption at rest and SSO support…
  3. 3pricing-plans.mdPlan table · 1,000-token chunkcut mid-table0.80| Feature | Pro | Business | Enter… (row cut at the boundary)
  4. 4release-notes-2026-04.mdSign-in0.78Google and Microsoft sign-in now available on every plan…
  5. 5blog-security.mdAnnouncementoff-topic0.73We believe good security should be available to everyone…

No reranker. All five chunks go to the model in vector-score order, off-topic chunks included.

Context sent to the model≈ 5,000 tokens

Yes, single sign-on is available on every paid plan, including Pro.1not in sources Admins connect an identity provider using SAML 2.0 from the security settings.1

Shown to the user as written

Evals

Faithfulness0.68
Answer relevance0.88
Context precision0.52
Context recall0.50

A

Model only

  • No retrieval
  • No sources
  • No citations

queryDoes the Pro plan include single sign-on?

Retrieval switched off. The model answers from what it learned in training, with no access to your documents.

Nothing to rank.

Single sign-on is often reserved for business or enterprise tiers, so the Pro plan may not include it. Contact the vendor’s sales team to confirm.ungrounded · no sources

Plausible, generic and unverifiable

Evals

Faithfulnessn/a
Answer relevance0.55
Context precisionn/a
Context recalln/a

B

Tuned RAG

  • Semantic chunks · 300–500 tokens
  • Hybrid BM25 + vector
  • Cross-encoder reranker
  • Citations enforced

rewritePro plan · SAML single sign-on · included or add-on

filterdoc_type in pricing, docs, release notes

  1. 1sso-setup-guide.mdConfigure SAML0.87Admins on paid plans can connect an identity provider using SAML 2.0…
  2. 2release-notes-2026-06.mdSSO add-on0.85SAML SSO is now available to Pro workspaces as a paid add-on…
  3. 3release-notes-2026-04.mdSign-in0.81Google and Microsoft sign-in now available on every plan…
  4. 4pricing-plans.mdPlans · identity row0.80SAML SSO: Business included · Enterprise included · Pro not included
  5. 5security-whitepaper.pdf§5 Identityoff-topic0.72Customer data is protected by encryption at rest and SSO support…

  1. 1pricing-plans.mdPlans · identity rowwas #40.96
  2. 2release-notes-2026-06.mdSSO add-onwas #20.93
  3. 3sso-setup-guide.mdConfigure SAMLwas #10.71
Context sent to the model≈ 1,100 tokens

Chunks scoring under 0.50 are dropped before generation.

Not by default: SAML single sign-on is included in the Business and Enterprise plans.1 Pro workspaces can add it as a paid add-on, available since June 2026.2

Grounded · every sentence cited · answered

Evals

Faithfulness0.96
Answer relevance0.94
Context precision0.91
Context recall0.92
Eval scores for question 2: column A against column B
EvalsAB
Faithfulness0.68n/a0.96
Answer relevance0.880.550.94
Context precision0.52n/a0.91
Context recall0.50n/a0.92

QuestionCan I carry unused annual leave into next year?

A

Baseline RAG

  • Fixed 1,000-token chunks
  • Vector search only
  • No reranker
  • Citations optional

queryCan I carry unused annual leave into next year?

No rewrite, no metadata filter.

  1. 1employee-handbook.pdf§7 · 1,000-token chunkcut mid-table0.86…annual leave … public holidays … carried forward … sick leave…
  2. 2holiday-calendar-2026.xlsxSheet 1off-topic0.8126 Jan Republic Day · 15 Aug Independence Day · 2 Oct…
  3. 3leave-policy-2025.pdf§3 Carry-oversuperseded0.80Up to 15 days may be carried forward to the next year…
  4. 4payroll-faq.mdLeave encashment0.76Encashment requests are processed with the March payroll…
  5. 5leave-policy-2026.pdf§1 Scope0.74This policy applies to all permanent employees in India…

No reranker. All five chunks go to the model in vector-score order, off-topic chunks included.

Context sent to the model≈ 5,000 tokens

Yes, you can carry forward up to 15 days of annual leave.3superseded source Any days above that are paid out in your final settlement.not in sources

Shown to the user as written

Evals

Faithfulness0.74
Answer relevance0.86
Context precision0.60
Context recall0.55

A

Model only

  • No retrieval
  • No sources
  • No citations

queryCan I carry unused annual leave into next year?

Retrieval switched off. The model answers from what it learned in training, with no access to your documents.

Nothing to rank.

Carry-over rules vary by employer and by local law. Many companies allow five to ten days to be carried forward. Check your HR policy.ungrounded · no sources

Plausible, generic and unverifiable

Evals

Faithfulnessn/a
Answer relevance0.58
Context precisionn/a
Context recalln/a

B

Tuned RAG

  • Semantic chunks · 300–500 tokens
  • Hybrid BM25 + vector
  • Cross-encoder reranker
  • Citations enforced

rewritecarry forward unused annual leave · limit · expiry date · exceptions

filtereffective_year = 2026 · 1 superseded policy excluded

  1. 1leave-policy-2026.pdf§1 Scope0.84This policy applies to all permanent employees in India…
  2. 2hr-portal-guide.mdExceptions0.83Managers can approve a carry-over exception under Requests…
  3. 3leave-policy-2026.pdf§3.2 Carry-over0.82Up to 10 unused days carry over and must be used by 31 March…
  4. 4payroll-faq.mdLeave encashment0.75Encashment requests are processed with the March payroll…
  5. 5holiday-calendar-2026.xlsxSheet 1off-topic0.6926 Jan Republic Day · 15 Aug Independence Day · 2 Oct…

  1. 1leave-policy-2026.pdf§3.2 Carry-overwas #30.98
  2. 2hr-portal-guide.mdExceptionswas #20.88
  3. 3leave-policy-2026.pdf§1 Scopewas #10.57
Context sent to the model≈ 1,100 tokens

Chunks scoring under 0.50 are dropped before generation.

Yes, up to 10 unused days carry over into the next year.1 Carried-over days must be used by 31 March, after which they lapse.1 Your manager can approve an exception in the HR portal.2

Grounded · every sentence cited · answered

Evals

Faithfulness0.95
Answer relevance0.93
Context precision0.89
Context recall0.94
Eval scores for question 3: column A against column B
EvalsAB
Faithfulness0.74n/a0.95
Answer relevance0.860.580.93
Context precision0.60n/a0.89
Context recall0.55n/a0.94

QuestionWhat is the API rate limit on the Team plan?

A

Baseline RAG

  • Fixed 1,000-token chunks
  • Vector search only
  • No reranker
  • Citations optional

queryWhat is the API rate limit on the Team plan?

No rewrite, no metadata filter.

  1. 1rate-limits.mdLimits by plan · 2025superseded0.88Starter 60 · Team 300 · Enterprise custom, requests per minute…
  2. 2api-reference.mdErrors0.84429 Too Many Requests is returned when a limit is exceeded…
  3. 3status-incident-0712.mdPostmortemoff-topic0.80Elevated 429 responses between 14:02 and 14:31 IST…
  4. 4pricing-plans.mdPlan table · 1,000-token chunkcut mid-table0.77| API access | Team | Enterprise | … (row cut at the boundary)
  5. 5sdk-readme.mdRetries0.74The SDK retries idempotent requests with exponential backoff…

No reranker. All five chunks go to the model in vector-score order, off-topic chunks included.

Context sent to the model≈ 5,000 tokens

The Team plan allows 300 requests per minute.1superseded source Requests over the limit return a 429 error.2 Limits reset at the start of every hour.not in sources

Shown to the user as written

Evals

Faithfulness0.70
Answer relevance0.90
Context precision0.62
Context recall0.48

A

Model only

  • No retrieval
  • No sources
  • No citations

queryWhat is the API rate limit on the Team plan?

Retrieval switched off. The model answers from what it learned in training, with no access to your documents.

Nothing to rank.

API rate limits typically range from 60 to 1,000 requests per minute depending on the plan. The provider’s documentation lists the exact figures.ungrounded · no sources

Plausible, generic and unverifiable

Evals

Faithfulnessn/a
Answer relevance0.52
Context precisionn/a
Context recalln/a

B

Tuned RAG

  • Semantic chunks · 300–500 tokens
  • Hybrid BM25 + vector
  • Cross-encoder reranker
  • Citations enforced

rewriteTeam plan API rate limit · requests per minute · burst · 429 Retry-After

filterlatest version per document · changelog included

  1. 1api-reference.mdErrors0.86429 Too Many Requests is returned with a Retry-After header…
  2. 2sdk-readme.mdRetries0.80The SDK retries idempotent requests with exponential backoff…
  3. 3rate-limits.mdLimits by plan · 20260.79Team: 600 requests per minute per workspace, bursts to 1,000 over 10 s…
  4. 4changelog-2026-08.md4 August0.78Team rate limit raised from 300 to 600 requests per minute…
  5. 5status-incident-0712.mdPostmortemoff-topic0.66Elevated 429 responses between 14:02 and 14:31 IST…

  1. 1rate-limits.mdLimits by plan · 2026was #30.97
  2. 2changelog-2026-08.md4 Augustwas #40.94
  3. 3api-reference.mdErrorswas #10.90
Context sent to the model≈ 1,100 tokens

Chunks scoring under 0.50 are dropped before generation.

The Team plan allows 600 requests per minute per workspace, raised from 300 on 4 August 2026.12 Short bursts of up to 1,000 requests over 10 seconds are allowed.1 Requests over the limit return HTTP 429 with a Retry-After header.3

Grounded · every sentence cited · answered

Evals

Faithfulness0.97
Answer relevance0.95
Context precision0.86
Context recall0.93
Eval scores for question 4: column A against column B
EvalsAB
Faithfulness0.70n/a0.97
Answer relevance0.900.520.95
Context precision0.62n/a0.86
Context recall0.48n/a0.93

QuestionWho approves an expense claim over ₹50,000?

A

Baseline RAG

  • Fixed 1,000-token chunks
  • Vector search only
  • No reranker
  • Citations optional

queryWho approves an expense claim over ₹50,000?

No rewrite, no metadata filter.

  1. 1expense-policy-v4-draft.docx§2 Approvals · draftdraft0.87Claims above ₹50,000 need department head sign-off…
  2. 2travel-policy.pdf§6 Hotelsoff-topic0.82Hotel stays above the city cap need prior approval…
  3. 3expense-policy-v3.pdf§2 Approvals0.81Claims above ₹50,000 are approved by the department head…
  4. 4finance-faq.mdReimbursement0.78Approved claims are paid in the next payroll cycle…
  5. 5onboarding-checklist.mdWeek 1off-topic0.72Set up your expense account and corporate card…

No reranker. All five chunks go to the model in vector-score order, off-topic chunks included.

Context sent to the model≈ 5,000 tokens

Expense claims above ₹50,000 are approved by the department head.1draft source Approvals are usually completed within two working days.not in sources

Shown to the user as written

Evals

Faithfulness0.66
Answer relevance0.84
Context precision0.55
Context recall0.46

A

Model only

  • No retrieval
  • No sources
  • No citations

queryWho approves an expense claim over ₹50,000?

Retrieval switched off. The model answers from what it learned in training, with no access to your documents.

Nothing to rank.

Large expense claims are usually approved by a senior manager or the finance department. Your company’s expense policy will name the approver.ungrounded · no sources

Plausible, generic and unverifiable

Evals

Faithfulnessn/a
Answer relevance0.49
Context precisionn/a
Context recalln/a

B

Tuned RAG

  • Semantic chunks · 300–500 tokens
  • Hybrid BM25 + vector
  • Cross-encoder reranker
  • Citations enforced

rewriteapprover · expense claim above ₹50,000 · delegation of authority · current

filterstatus = approved · 1 draft excluded

  1. 1finance-faq.mdReimbursement0.83Approved claims are paid in the next payroll cycle…
  2. 2expense-policy-v3.pdf§2 Approvals0.82Claims above ₹50,000 are approved by the department head…
  3. 3delegation-of-authority.xlsxRow 14 · Expenses0.81₹50,000–₹1,00,000: Finance controller · above ₹1,00,000: CFO…
  4. 4travel-policy.pdf§6 Hotelsoff-topic0.70Hotel stays above the city cap need prior approval…
  5. 5onboarding-checklist.mdWeek 1off-topic0.61Set up your expense account and corporate card…

  1. 1delegation-of-authority.xlsxRow 14 · Expenseswas #30.95
  2. 2expense-policy-v3.pdf§2 Approvalswas #20.93
  3. 3finance-faq.mdReimbursementwas #10.41
Context sent to the model≈ 1,100 tokens

Chunks scoring under 0.50 are dropped before generation.

The sources disagree: the delegation-of-authority matrix names the finance controller for claims of ₹50,000 to ₹1,00,000.1 The expense policy (v3) names the department head.2 I have sent this to Finance operations to confirm, with both sources attached.

Sources conflict · confidence 0.52 < 0.70 · handed to Finance operations

Evals

Faithfulness0.93
Answer relevance0.71
Context precision0.84
Context recall0.88
Eval scores for question 5: column A against column B
EvalsAB
Faithfulness0.66n/a0.93
Answer relevance0.840.490.71
Context precision0.55n/a0.84
Context recall0.46n/a0.88

The four scores

Faithfulness
Share of the answer’s claims that the retrieved context supports.
Answer relevance
How directly the answer addresses the question asked.
Context precision
Whether the relevant passages rank above the irrelevant ones.
Context recall
Whether the search found everything the answer needed.

Metric names follow open-source evaluation tools such as Ragas. The AI models that score answers are checked against your experts’ labels before their scores decide a release.

Citations are enforced

Every sentence must cite a passage the model was given; a check removes or rewrites any that do not. Unsupported answers are never shown, and draft or superseded documents are excluded from search.

Low confidence goes to a person

When sources conflict or scores are low, the assistant says so and passes the question to the team that owns the answer. Their reply is logged, and the question joins the next release’s evaluation set.

04.3Conversational AI/conversation

Assistants that know when to hand over.

The assistant resolves the requests you allow it to, using your own systems. When confidence is low or a request needs a person, it hands over to your team.

Illustrative support conversations. In the chat, a customer reports a cracked blender jug; the assistant looks up the order and books a replacement with confidence 0.93 and 0.88. The customer then reports a duplicate card charge; confidence falls to 0.41, below the 0.60 handover threshold, and refunds need a person, so the assistant hands over to billing support with full context. Priya from billing support reverses the charge. The card number was redacted before logging. On the voice call, a caller changes a delivery address: the caller interrupts the assistant mid-sentence, verifies with a one-time code, and the assistant updates the order with confidence 0.94 and no handover. The phone number, code and address are redacted before logging.

support · conversation 7731
  1. Customer

    My order arrived today but the blender jug is cracked.

  2. Assistant

    Sorry about that. I can see order 48213, delivered this morning. I can send a replacement jug or refund ₹3,499. Which would you prefer?

    orders.lookup · 0.93

  3. Customer

    A replacement, please. Can it arrive before Saturday? It’s a gift.

  4. Assistant

    It can reach you by Friday. I’ve booked the replacement and emailed a free return label for the cracked jug.

    replacements.create · 0.88

  5. Customer

    Also, my card 4111 1111 1111 4242 was charged twice for this order.

  6. Handed to Billing support Confidence 0.41 is under the 0.60 line, and refunds need a person. Transcript, intent, tool calls and a suggested reply attached.
  7. Assistant

    I’ll pass this to our billing team with everything so far, so you won’t need to repeat it. Someone will reply here within a few minutes.

    handover · 0.41

  8. Billing support

    Hi, this is Priya from billing support. I can see the duplicate charge of ₹3,499 and I’ve reversed it. It will show on your statement within 5 to 7 working days.

    Sent by Priya · Billing support

Voice call · order support

Streaming · barge-in on · 00:19

  1. 00:00.8 CallerHi, I need to change the delivery address for my order.
  2. 00:02.9 AssistantOf course. Is that order 48213, arriving on Friday?
  3. 00:05.1 Caller · barge-inYes, that one. Send it to my office instead, the address on my account.
  4. 00:05.2 Barge-in detected · assistant audio stopped in 90 ms
  5. 00:08.4 AssistantI’ve sent a one-time code to the number ending 42. Could you read it out?
  6. 00:14.7 CallerIt’s [code redacted].
  7. 00:16.2 AssistantThank you, that’s verified. Your order will now go to your office address on Friday.
  8. 00:19.0 Resolved · no handover · call 00:19

Time to first audio Illustrative

  • Speech-to-text 190 msWhisper large-v3 turbo, streamed
  • LLM first token 310 mssmall model, routed
  • Text-to-speech first audio 220 msElevenLabs Flash
  • Total 720 mstarget under 800 ms
  • Live voice calls

    Speech-to-text, the AI model and text-to-speech stream as one pipeline, with voice detection so callers can interrupt. Our design target is first audio in under about 800 ms.

  • Handover with context

    Thresholds, permissions and topics that always need a person are set for each type of request. Whoever takes over sees the transcript, intent, actions taken and a suggested reply.

  • Redaction before logging

    Card numbers, phone numbers, one-time codes and government IDs are masked before transcripts are stored, sent to analytics or reused as evaluation data.

04.4Vision & documents/vision

Vision models that count, read and check.

Drag the handle to compare the camera frame with what the model saw. Each detection has a label and a confidence score, and anything below the review threshold goes to a person.

Illustrative detectionsReview threshold 0.70

Seven dark wine bottles with yellow capsules pass along a conveyor in front of stainless steel tanks
Line 2 · camera 3 · frame 18,204detector v2.3 · edge GPU · 14 ms

7 bottles · 1 spacing fault · 1 to re-check

Bottling line: detections per class with precision and recall on the validation set, and the action each one triggers (illustrative)
ClassCountPR
capsule_okPass60.990.98
pitch_shortSlow the infeed to re-space10.950.93
occludedRe-check at camera 4 · 0.62 < 0.7010.880.90
Precision
Of everything flagged, the share that was right. Low precision wastes people’s time.
Recall
Of all defects present, the share that was caught. Low recall lets problems through.
A chilled supermarket shelf of milk and soy milk cartons in rows above electronic price labels
Store 114 · aisle 6 · photo 09:12detector v1.8 + OCR · phone app · 0.9 s

17 facings in frame · 1 price mismatch · 1 gap

Retail shelf: detections per class with precision and recall on the validation set, and the action each one triggers (illustrative)
ClassCountPR
facingsPlanogram check170.970.95
price_tagOCR against the price file90.980.96
price_mismatchTask to store staff10.910.88
gapReplenishment request10.940.92
low_stockHuman review · 0.66 < 0.7010.790.83
Precision
Of everything flagged, the share that was right. Low precision wastes people’s time.
Recall
Of all defects present, the share that was caught. Low recall lets problems through.

Route A

Fine-tuned detector at the camera

An object detector such as YOLO, trained on your labelled images, runs on a GPU beside the camera in tens of milliseconds per frame. It suits counting, inspection and anything moving on a line.

Frames per secondPrecision and recall per classWorks offline

Route B

AI model that reads documents

The model reads layout, tables, stamps and handwriting and returns structured data, checked against a schema before anything posts. Each page takes seconds, in your cloud or with a provider under data-processing terms.

Field-level accuracyStraight-through rateSchema validation

Both routes

People review the uncertain cases

Reviewers confirm or correct each uncertain result in one click, and corrections become labelled data. A new model goes live only if it beats the current one on a held-back test set.

Review queueLabelled dataHeld-back test set

04.5Workflow automation/automation

Automation from inbox to ERP, with a person at each gate.

An invoice arrives by email and leaves as a posted ERP document. Steps retry without posting twice, failures are kept for replay, and anything over the limit waits for a person. Try approving one.

A thick, uneven stack of printed invoices, receipts and forms
The work it removesInvoices, receipts and forms that someone retypes today become structured records.Photo: Camilo Rueda Lopez / Unsplash
accounts-payable · workflow v14 · durable execution r-2291 · completed Illustrative
Accounts payable workflow: email trigger, classify, extract, validate with retries, branch on amount, human approval over ₹2,00,000, post to ERP, notify the team, audit record, and a dead-letter queue after three failed retries. yes · over limit no after 3 failed retries retry · backoff 2 s, 4 s, 8 s TRIGGER Email with PDF ap@your-company · 1 PDF CLASSIFY LLM · document type invoice · 0.98 EXTRACT Vision · 12 fields min confidence 0.95 VALIDATE Rules + ERP lookup PO matched · 3-way ok BRANCH Over ₹2,00,000? ₹3,40,000 · yes APPROVAL Finance lead approved · 3 m 02 s POST ERP document 5100042 · idempotent NOTIFY Team channel #ap-approvals · sent AUDIT Run record r-2291 · 14 events DEAD LETTER After 3 retries empty
run console · r-2291
  1. triggeremail received · ap@your-company · inv-88412.pdf
  2. classifyinvoice · 0.98 · small model
  3. extract12 fields · min confidence 0.95
  4. validateERP lookup timed out · retry 1 in 2 s
  5. validatePO 4500123 matched · 3-way match ok
  6. branch₹3,40,000 over ₹2,00,000 · approval required
  7. approvalapproved by Finance lead
  8. postERP document 5100042 · key inv-88412
  9. notify#ap-approvals · message sent
  10. auditrun r-2291 closed · 14 events

Approval gate · ₹3,40,000 · approved by Finance lead

  • 78%Straight-through this week
  • 41 sMedian cycle, no approval
  • 6 · 1Retries · dead-lettered
Recent workflow runs with amount, path, duration, retries and status (illustrative)
RunAmountPathDurationRetriesStatus
r-2291inv-88412.pdf₹3,40,000Approval3 m 08 s1Posted · approved
r-2290inv-88409.pdf₹48,200Straight-through38 s0Posted
r-2289inv-88401.pdf₹1,12,000Straight-through41 s0Posted
r-2288inv-88397.pdf₹2,75,000Approval—0Waiting for approval
r-2287inv-88390.pdf—Dead letter2 m 05 s3ERP unavailable · replayed
  • No double posting

    Every write carries a unique key, so a retry never posts twice.
  • Automatic retries

    Timeouts and rate limits retry after 2 s, 4 s and 8 s, then go to the dead-letter queue. Permanent errors fail at once.
  • Dead-letter queue

    Runs that use up their retries are parked with their data and error, ready to replay after a fix.
  • Audit trail per run

    Each run records who approved what, which model version decided, and every input and output.

Runs on tools such as

  • n8n
  • Temporal
  • SAP
  • Python
  • Zapier

AlsoPower AutomateSlackMicrosoft TeamsGmail

04.6How it works/stack

Five stages from documents to answers, each replaceable without a rebuild.

Each stage has a job, a reason and the technologies we prefer. We own the connections between stages, so a model, index or parser can be replaced when a better one is released.

pipeline · knowledge-assistant · v24 · hosted + self-hosted
  1. 01

    Ingest & parse

    Documents, tickets, tables and scans become clean text labelled with source, version, owner and who may read it. Poor parsing is the most common cause of wrong answers.

    We work with

    • Python either
    • Apache Airflow either
    • Airbyte either
    • Apache Kafka either
    • PostgreSQL either

    AlsoUnstructured

    1,284 docs · 18,402 chunks · 300–500 tokens

  2. 02

    Embed & index

    Each passage is indexed by meaning and by keyword, and stored with its permissions. Search matches both and never returns what a user could not open.

    We work with

    • PostgreSQL + pgvector either
    • Qdrant either
    • Elasticsearch either
    • Milvus either
    • Hugging Face either

    AlsoWeaviatePinecone

    1,024-d · HNSW · nightly re-index

  3. 03

    Retrieve & rank

    Keyword and meaning search fetch twenty candidates; a reranker keeps the three most relevant, so answers are better and cheaper.

    We work with

    • LangChain either
    • Elasticsearch either
    • Redis either

    AlsoLlamaIndexCohere Rerank

    hybrid · top 20 → rerank 3 · 41 ms

  4. 04

    Generate & guard

    A router sends classification to a small model and drafting or reasoning to a larger one. Automatic checks screen inputs for prompt injection and outputs for format, citations and personal data.

    We work with

    • Anthropic hosted API
    • Google Gemini hosted API
    • Mistral AI either
    • Meta Llama self-hosted
    • DeepSeek either
    • vLLM self-hosted
    • LangGraph either

    AlsoOpenAIGuardrails

    Voice in and out

    Whisper (open weights)DeepgramElevenLabsAzure AI Speech

    routed · hosted APIs · p95 1.8 s · streamed

  5. 05

    Evaluate & observe

    Every release is tested automatically against your evaluation set. Each live answer is logged with its model version, sources, cost and response time, and a daily sample is scored.

    We work with

    • MLflow either
    • OpenTelemetry either
    • Grafana either
    • Prometheus either
    • GitHub Actions either

    AlsoRagasLangfuse

    24 gates in CI · daily sample of 200

Swaps tested like releases

Swapping the parser, index, search, model or scoring judge is a configuration change and a test run on your evaluation set. New models are released monthly, and your feature can use them the same week.

Self-hosted open models

For regulated data or residency rules, Llama, Mistral or DeepSeek models run with vLLM on GPUs in your private cloud or an Indian data centre. Nothing leaves your network, and the same tests apply.

Chosen on test scores

Each candidate setup is scored on your evaluation set for quality, 95th-percentile response time and cost per answer before we choose. You keep the comparison table, and we re-run it when a new option appears.

04.7Quality gates/quality

Quality gates that every prompt, model or document change must pass.

Each change is treated as a release and tested in CI against an evaluation set your experts signed off, before any user sees it. Select a gate to see what it tests.

release · assistant v24 · prompt v23 → v24 · model upgrade Released · canary 5% → 100% Illustrative

8 gates · 1 blocked then fixed · 1 human approval12 m 40 s

Gate 01 of 8

Golden-set evals

Every question in the golden set is answered by the candidate release and scored for faithfulness, answer relevance, context precision and context recall by a judge model calibrated against your experts’ labels. A score under threshold blocks the release; the failing answers are listed with their retrieved chunks so the fix is obvious.

Catches

  • Regression on prompt or model change
  • Retrieval drift after re-indexing

Runs with

  • Ragas
  • Langfuse
  • GitHub Actions

Gate 02 of 8

Production regression

Last month’s real questions, anonymised, are replayed through the candidate. Any sentence that the citation check newly flags as unsupported, and any answer whose route changed from answered to handed over without a reason, is listed for review.

Catches

  • Silent behaviour change
  • Handover rate creeping up

Runs with

  • Replay harness
  • Langfuse

Gate 03 of 8

Red-team suite

Attack prompts cover direct and indirect prompt injection (LLM01), attempts to extract other users’ data or secrets (LLM02), outputs that would be unsafe if rendered or executed downstream (LLM05), system prompt leakage (LLM07) and confident misinformation on questions the sources do not answer (LLM09). Every attack must fail. An attack that once succeeded stays in the suite for good.

Catches

  • LLM01 prompt injection
  • LLM02 sensitive information disclosure
  • LLM05 improper output handling
  • LLM07 system prompt leakage
  • LLM09 misinformation

Runs with

  • Promptfoo
  • Garak
  • Custom suites

Gate 04 of 8

PII leakage tests

Synthetic card numbers, phone numbers and government-ID formats are planted in retrieved context and in user turns. The release passes only when none appear in answers, traces, analytics events or evaluation datasets, which proves the redaction step runs before logging.

Catches

  • Personal data in logs
  • Personal data in eval sets

Runs with

  • Presidio
  • Regex + checksum
  • Trace assertions

Gate 05 of 8

Latency and cost budget

The eval run doubles as a load run: p95 latency and cost per answer are computed from the traces and compared with the budget agreed at prototype stage. A model upgrade that is better but three times the cost is a conversation, not a surprise on the invoice.

Catches

  • Cost regression
  • Latency regression

Runs with

  • OpenTelemetry
  • Grafana
  • Cost model

Gate 06 of 8

Accessibility of the AI UI

Streamed answers are announced to screen readers as they complete, not token by token. Citations are real links with names, the stop button is reachable by keyboard, and contrast holds in both themes. Automated checks run in CI; a manual pass runs before each major release.

Catches

  • WCAG 2.2 AA
  • Screen-reader announcements

Runs with

  • axe-core
  • Playwright
  • Manual pass

Gate 07 of 8

Content credentials

Where a feature generates images, audio or video, a C2PA content credential is attached at generation time, stating that the media is AI-generated and by which model. The gate verifies the manifest survives the delivery pipeline, so the disclosure reaches the person who sees it.

Catches

  • Undisclosed synthetic media
  • Manifest stripped by resizing

Runs with

  • c2pa-rs
  • Signing key in KMS

Gate 08 of 8

Human sign-off

A person with accountability for the feature reads the eval summary, the flagged answers and the red-team report, then approves the release to a 5% canary. The rollout widens over 24 hours only if production groundedness holds. The approval, the evidence and the rollout are one audit record.

Catches

  • Unowned releases
  • Big-bang rollouts

Runs with

  • Release record
  • Feature flags

Frameworks we build to

We drew these badges for the frameworks our gates align with. They show how we work; we hold no certification for them.

  • OWASP Top 10 for LLM ApplicationsSecurity risks in generative AI applicationsRisks specific to applications built on AI models, including prompt injection (LLM01), sensitive information disclosure (LLM02), excessive agency (LLM06) and vector and embedding weaknesses (LLM08).
  • ISO/IEC 42001:2023Artificial intelligence management systemsRequirements for establishing, running and improving an AI management system: AI policy, impact assessment, data and lifecycle controls.
  • NIST AI RMF 1.0AI Risk Management FrameworkFour functions for trustworthy AI: Govern, Map, Measure and Manage, with a companion profile for generative AI (NIST AI 600-1).
  • WCAG 2.2 AAWeb Content Accessibility GuidelinesPerceivable, operable, understandable and robust content, including 2.2 criteria such as focus not obscured, target size and accessible authentication.
  • GDPRGeneral Data Protection Regulation (EU) 2016/679Lawful basis, data-subject rights, data protection by design and by default, breach notification and DPIAs for high-risk processing.
  • DPDP Act 2023Digital Personal Data Protection Act, 2023Notice and consent, duties of data fiduciaries, rights of data principals, breach intimation and added duties for significant data fiduciaries, with the DPDP Rules.
  • EU AI ActArtificial Intelligence Act (EU) 2024/1689A risk-based regime: prohibited practices, obligations for high-risk systems, transparency duties and rules for general-purpose AI models.
  • C2PA content credentialsProvenance for generated mediaAn open standard for attaching signed provenance to images, audio and video, including whether AI generated or edited them.

04.8Cost & energy/efficiency

Lower cost and energy per answer.

The same question can cost many times more depending on design. We model cost and energy per answer in the prototype, then cut both until they fit your volume, with no drop in test scores.

cost lab · knowledge-assistant · per 1,000 answers Illustrative

Cost per 1,000 answers₹90

  • Retrieval & rerank ₹28
  • Small model ₹20
  • Large model ₹36
  • Cache ₹6

Energy per 1,000 answers150 Wh

  • Retrieval & rerank 30 Wh
  • Small model 40 Wh
  • Large model 75 Wh
  • Cache 5 Wh

As B, plus a semantic cache that serves the thirty-five percent of questions asked before, and prompt caching for the static system prompt and tool schemas. ₹90 per 1,000 answers is ₹0.09 each, inside the ₹0.12 release budget.

Tokens to the model per answer
940
p95 latency
1.2 s
Monthly cost at volume
₹18,000
Monthly energy at volume
30 kWh
Golden-set faithfulness
0.93 A 0.90 · B 0.93 · C 0.93 · all clear the 0.90 gate
Against design A
−91% cost · −85% energy

Semantic caching

Repeat questions, even when worded differently, are answered from a checked cache. Cached answers keep their sources and expire when those sources change.

35% of traffic · 0 model tokens

Prompt caching

The instructions, tool definitions and policy text are the same on every call. Providers and vLLM can cache that fixed part, cutting most of its input cost.

static prefix · up to 90% off input

Shorter context via reranking

Keeping only the three most relevant passages out of twenty means the model reads a fifth of the text and answers more precisely.

5,000 → 1,100 tokens

Small models for the small jobs

Intent, language, personal-data detection and routing run on small models. The large model is reserved for the questions that need it.

70% routed small

Batch jobs in low-carbon hours

Re-indexing, embeddings and nightly tests run as batch jobs in the region’s lowest-carbon hours, timed by a grid carbon forecast and charged at off-peak rates.

carbon-aware window

SCI · ISO/IEC 21031:2024Software Carbon IntensityHow we apply it Carbon measured per request, user or inference, so efficiency work has a baseline.

SCI = ((E × I) + M) per R Energy used, times the carbon intensity of the grid it ran on, plus the hardware’s embodied emissions, per unit of work. Here R is one answered question, so tokens per answer is the number we watch.

04.9How we work with you/process

From prototype to live feature in four phases.

With your experts, we agree what a good answer looks like before building. Every change is then scored against it. You see a working prototype on your documents in the first sprint. After launch, prompt, model and search changes are released weekly, each scored before users see it.

delivery plan · one AI feature · 14 weeks · 4 exit gates Typical plan

A typical fourteen-week plan. Frame runs weeks 1 to 2, Prototype weeks 2 to 5, Build weeks 5 to 12 and Operate weeks 12 to 14, each ending at an exit gate. The golden set grows from 40 questions in week 2 to 400 by week 14. From week 4 every build is scored against it; one release in week 9 is blocked and fixed the next day, and weekly production releases start in week 12.

Phase 01 of 4Wk 01–02

Frame

Define the task, users, data sources and the failures that matter; draft success measures and an evaluation set with your experts.

  • Feature spec
  • Evaluation set v1
  • Data inventory
Exit gate G1
Your domain experts sign off evaluation set v1: 40 to 60 questions or documents from your work, each with its expected answer and source.
Every week you see
Two working sessions with the people who own the answers. The data inventory and access plan by the end of week 2.
Where AI helps
Agents group past tickets and search logs into draft questions; your experts review the draft instead of starting from nothing.

Phase 02 of 4Wk 02–05

Prototype

Document search, prompts and models compared on the evaluation set for quality, speed and cost; the design is chosen on the scores.

  • Prototype
  • Model comparison
  • Cost model
Exit gate G2
The chosen design beats the baseline on the evaluation set, within the response-time and cost budget agreed for launch.
Every week you see
A working prototype on your own documents by week 3. The model comparison table updated every Friday.
Where AI helps
Three or four search and model designs are scored overnight on the evaluation set, so the choice is made on numbers within days.

Phase 03 of 4Wk 05–12

Build

Production pipeline, automatic checks, interface, feedback capture and monitoring. Releases must pass the evaluation set in CI.

  • Production feature
  • Automatic checks
  • Evaluation gate in CI
Exit gate G3
All eight release gates pass in CI, and a pilot group of named users has used it in daily work for two weeks.
Every week you see
A demo every Friday on staging, with that week’s evaluation scores beside it. Only work that has passed the gates is shown.
Where AI helps
Coding agents draft tests, adapters and UI states; a named engineer reviews every change, and it is re-scored before merge.

Phase 04 of 4Wk 12–14

Operate

Staged rollout, live scoring and alerts when quality drops or costs rise. The evaluation set grows from live use.

  • Rollout
  • Quality dashboard
  • Improvement backlog
Exit gate G4
Live answers stay grounded and the handover rate holds for four weeks. The runbook, dashboards and ownership pass to your team.
Every week you see
Weekly releases, each passing the evaluation gate, and a monthly quality and cost review with the feature owner.
Where AI helps
A daily sample of live answers is scored automatically, and a person reviews the flagged ones.

04.10What you get/deliver

Code, prompts and tests in a repository your team can run.

Code, pipelines, prompts, evaluation sets, automatic checks, workflows, dashboards and their documentation arrive as a pull request. Your team reviews and merges it, and can run the feature without us. Open any file.

Open · ready to mergehandover/v24 → mainyour-company/ai-assistant

Handover #1: assistant v24 and accounts-payable automation

9 deliverables141 files changed6 of 6 checks passed2 of 2 approvals from your team

  1. app/The AI feature, in production Source 64 files

    API, streaming UI with citations and the handover path, deployed to your cloud.SourceTypeScriptPython

    app/answer/answer.ts+11 · excerptTS

    1. // Answer with sources, or hand over. Never an uncited guess.
    2. export async function answer(q: Question, user: User) {
    3. const found = await retrieve(q, { acl: user.groups, topK: 20 });
    4. const top = await rerank(q, found, { keep: 3, minScore: 0.5 });
    5. if (top.length === 0) return handover(q, "no-sources");
    6. ​
    7. const draft = await generate(q, top, { model: route(q) });
    8. const checked = enforceCitations(draft, top);
    9. if (checked.confidence < 0.7) return handover(q, "low-confidence", checked);
    10. ​
    11. trace.log({ q: redact(q), sources: top.map((s) => s.id), cost: draft.cost });
    12. return checked;
    13. }

    e41c09afeat: hand over when nothing passes the reranker2 days ago

  2. pipelines/ingest/Ingestion pipelines & index Code 18 files

    Parse, chunk, embed and index, nightly and on change, with permissions carried through.CodeAirflow DAGVector index

    pipelines/ingest/dag.py+10 · excerptPY

    1. # Nightly: parse, chunk, embed and index new or changed documents.
    2. @dag(schedule="0 2 * * *", start_date=datetime(2026, 1, 1), catchup=False)
    3. def ingest_knowledge_base():
    4. docs = changed_since_last_run(sources=["sharepoint", "confluence", "drive"])
    5. parsed = parse(docs, keep_tables=True, ocr="scanned-only")
    6. chunks = semantic_chunk(parsed, min_tokens=300, max_tokens=500, overlap=80)
    7. tagged = attach_meta(chunks, fields=["source", "version", "owner", "acl"])
    8. index.upsert(embed(tagged), keyword_index=True)
    9. retire(superseded(docs)) # old versions leave the index, not only the UI
    10. ​
    11. ingest_knowledge_base()

    7b2d5f0fix: retire superseded versions from the indexlast week

  3. config/Prompt & model registry YAML 9 files

    Models, routes, retrieval settings and prompts under version control. Every change is a release.YAMLVersioned

    config/assistant.yaml+13 · excerptYAML

    1. # Every change here is a release: it runs the eval gates in CI.
    2. version: 24
    3. router:
    4. classify: small-model # intent, language, PII, routing
    5. answer: large-model # drafting and reasoning
    6. fallback: open-weights-70b # served with vLLM in your VPC
    7. retrieval:
    8. search: hybrid # BM25 + vector
    9. candidates: 20
    10. rerank: { model: cross-encoder, keep: 3, min_score: 0.50 }
    11. prompts:
    12. system: prompts/assistant.v24.md
    13. handover: prompts/handover.v7.md

    c90a3e1chore: v24 routes answers to large-model, open-weights fallbackyesterday

  4. evals/Evaluation suite Golden set 12 files

    The golden set your experts signed, the calibrated judge and the thresholds that block a release.Golden setScorersCI gate

    evals/gates.yaml+12 · excerptYAML

    1. golden_set: evals/golden-set.jsonl # 400 questions · owner: Support ops
    2. judge: calibrated-judge-v3 # agreement with expert labels: 0.91
    3. gates:
    4. faithfulness: { min: 0.90 }
    5. answer_relevance: { min: 0.85 }
    6. context_precision: { min: 0.80 }
    7. context_recall: { min: 0.80 }
    8. red_team: { successful_attacks: 0 } # OWASP LLM01, 02, 05, 07, 09
    9. pii_leakage: { found: 0 }
    10. p95_latency_s: { max: 2.5 }
    11. cost_per_answer_inr: { max: 0.12 }
    12. on_fail: block_release

    3f8b6d2test: add 40 billing questions to the golden setyesterday

  5. guardrails/Guardrails & review queues Policy 7 files

    Input and output checks, redaction and the rules for when a person takes over.PolicyWorkflow

    guardrails/policy.yaml+11 · excerptYAML

    1. input:
    2. prompt_injection: { detector: classifier, action: refuse_and_log } # LLM01
    3. pii: { detect: [card, aadhaar, phone, email], action: redact_before_log }
    4. output:
    5. citations: required # unsupported sentences are removed
    6. schema: answer.schema.json # tool calls and JSON validated
    7. render: plain_text # no raw HTML or links from the model (LLM05)
    8. handover:
    9. when: [confidence_below_0.70, sources_conflict, topic_in_refunds_or_legal]
    10. queue: support-tier-2
    11. attach: [transcript, intent, tool_calls, sources, suggested_reply]

    a17e4c8feat: redact Aadhaar numbers before logging3 days ago

  6. workflows/Automation workflows Temporal 23 files

    Durable, idempotent workflows with retries, approval gates and a dead-letter queue.Temporaln8n

    workflows/invoice_to_erp.py+18 · excerptPY

    1. # Retries after 2 s, 4 s and 8 s; a fourth failure goes to the dead-letter queue.
    2. RETRY = RetryPolicy(initial_interval=timedelta(seconds=2), backoff_coefficient=2.0, maximum_attempts=4)
    3. ​
    4. @workflow.defn
    5. class InvoiceToERP:
    6. def __init__(self) -> None:
    7. self.decision: str | None = None
    8. ​
    9. @workflow.signal
    10. def decide(self, decision: str) -> None: # "approve" or "reject", from the finance lead
    11. self.decision = decision
    12. ​
    13. @workflow.run
    14. async def run(self, email_id: str) -> str:
    15. doc = await workflow.execute_activity(extract_invoice, email_id, start_to_close_timeout=timedelta(seconds=60), retry_policy=RETRY)
    16. await workflow.execute_activity(validate_against_erp, doc, start_to_close_timeout=timedelta(seconds=30), retry_policy=RETRY)
    17. if doc.total > APPROVAL_LIMIT:
    18. await workflow.wait_condition(lambda: self.decision is not None)
    19. if self.decision == "reject":
    20. return await workflow.execute_activity(notify_rejected, doc, start_to_close_timeout=timedelta(seconds=30))
    21. return await workflow.execute_activity(post_to_erp, doc, start_to_close_timeout=timedelta(seconds=30), retry_policy=RETRY)

    5d0c9b7fix: idempotency key on the ERP post steplast week

  7. dashboards/Quality, latency & cost dashboards Grafana 6 files

    The numbers the feature is judged on, with alerts wired to an owner.GrafanaLangfuse

    dashboards/assistant.json+11 · excerptJSON

    1. {
    2. "title": "Knowledge assistant · quality, latency, cost",
    3. "panels": [
    4. { "title": "Groundedness, daily sample", "threshold": 0.90 },
    5. { "title": "Handover rate by intent", "unit": "percent" },
    6. { "title": "p95 latency by model route", "unit": "s" },
    7. { "title": "Cost per answer", "unit": "INR" },
    8. { "title": "Top unanswered questions", "type": "table" }
    9. ],
    10. "alerts": ["groundedness_7d < 0.90", "cost_per_answer > budget"]
    11. }

    b62f1a4feat: alert when 7-day groundedness is under 0.904 days ago

  8. docs/runbook.mdRunbooks Markdown 1 file

    What to do when a number moves, written so your team can act without us.MarkdownOn-call

    docs/runbook.md+8 · excerptMD

    1. # Runbook · knowledge assistant
    2. ​
    3. ## Groundedness below 0.90 for a day
    4. 1. Open the flagged answers in Langfuse and group them by source.
    5. 2. Check index freshness: a failed nightly ingest is the usual cause.
    6. 3. Re-run ingestion for that source, then re-score the sample.
    7. 4. If it persists, roll config back to the last good version.
    8. ​
    9. ## Provider outage or rate limits
    10. Traffic fails over to the fallback route. Answers stay cited.

    91e7c3ddocs: provider outage and rate-limit stepslast week

  9. docs/model-card.mdModel cards Markdown 1 file

    Intended use, limits, evaluation and owner for each release, aligned with ISO/IEC 42001 documentation.MarkdownPer release

    docs/model-card.md+8 · excerptMD

    1. # Model card · assistant v24
    2. ​
    3. Intended use Staff questions answered from approved policy documents
    4. Out of scope Legal or medical advice; decisions about individuals
    5. Models small-model (routing) · large-model (answers)
    6. Data 1,284 documents, permission-filtered at query time
    7. Evaluation Golden set of 400 · faithfulness 0.93 · relevance 0.94
    8. Known limits Tables in scanned PDFs; questions spanning 5+ documents
    9. Owner Head of Support operations · reviewed each release

    0c4a8f6docs: model card for assistant v24yesterday

04.11What changes/outcomes

Agreed measures that decide if the feature stays on.

In week one we agree the measures, their baseline and the floor that switches the feature off. After launch they sit on one dashboard your team owns. Typical measures are shown below.

production · assistant and AP automation · 12 weeks after launch liveIllustrative
  • Deflection rate

    41%+23 pts

    was 18% via searchtarget ≥ 40%

    Questions resolved with no ticket. Twelve weeks: 22, 26, 29, 31, 33, 35, 36, 37, 38, 39, 40, 41.

  • Average handle time

    6.1 min−35%

    was 9.4 mintarget ≤ 7 min

    Handed-over conversations. Twelve weeks: 9.1, 8.6, 8.2, 7.8, 7.5, 7.1, 6.9, 6.7, 6.5, 6.3, 6.2, 6.1.

  • Straight-through processing

    78%+78 pts

    was 0%, all manualtarget ≥ 75%

    Invoices posted with no touch. Twelve weeks: 52, 58, 63, 66, 69, 71, 72, 74, 75, 76, 77, 78.

  • Cost per resolution

    ₹0.55−50%

    was ₹1.10 in week 1budget ≤ ₹0.70

    Model and infrastructure, about six answers at ₹0.09. Twelve weeks: 1.1, 1.02, 0.95, 0.84, 0.76, 0.72, 0.7, 0.64, 0.6, 0.58, 0.56, 0.55.

  • Groundedness

    0.96+0.05

    floor 0.90target ≥ 0.93

    Daily sample of 200 answers. Twelve weeks: 0.91, 0.92, 0.93, 0.93, 0.94, 0.93, 0.92, 0.93, 0.94, 0.95, 0.95, 0.96.

Answers you can trace

Responses cite their sources and are scored for accuracy, so people can check them.

Groundedness · weekly mean of the daily sample0.96
0.80 0.85 0.90 0.95 1.00 1 2 3 4 5 6 7 8 9 10 11 12 floor 0.90 W7

Week 7: a policy upload failed to index and answers cited a stale version. One day’s sample fell to 0.87 (daily low; weekly mean 0.92). The alert fired the same day; fixed in 26 hours.

Every answer carries its sources; 0 uncited answers shown to users

Exceptions handled by people

Unusual cases reach a person with full context; more work is automated as accuracy is shown.

Straight-through rate · invoices posted with no touch78%
40% 50% 60% 70% 80% 1 2 3 4 5 6 7 8 9 10 11 12 target 75% W6

Week 6: the review threshold moved from 0.95 to 0.92 after 1,200 checked cases showed no loss in accuracy.

Exceptions reach a named person with the document, the fields and the reason

Running costs known up front

Cost per request and per resolved task is tracked from the prototype onwards.

Cost per resolution · model and infrastructure₹0.55
₹0.40 ₹0.60 ₹0.80 ₹1.00 ₹1.20 1 2 3 4 5 6 7 8 9 10 11 12 budget ₹0.70 W4

Week 4: the semantic cache went live and now serves about a third of questions with no model call.

Cost per answer tracked from the prototype, not found on the first invoice

Services & packages

AI features your customers use, and automation your team can trust.

From a knowledge assistant to automation for critical processes, each service lists what is included, typical timing and the teams it suits.

Categories
05
Services
17
Packages
04
Not sure what you need? Describe the problem

How to buy

  1. 01Pick services. Enquire about one, or add several to a brief.
  2. 02Choose a package. A sprint, a fixed project or an ongoing team.
  3. 03Send the brief. We reply within one working day.

Browse by category

Timelines are typical. Every quote follows a written scope.

02Conversational AI

4 services
Typical timeline: 6–10 weeks

Customer support assistant

A chat assistant for your site or app that resolves order, return and booking requests, and hands harder cases to a person with full context.

What’s included

  • Intent design from your support tickets
  • Integration with order, CRM or booking systems
  • Handover to your helpdesk with context
  • Resolution, handover and satisfaction tracking
  • Chat
  • Handover
  • Resolution rate
  • OpenAI

Best forSupport teams handling high volumes of repeat questions.

Typical timeline: 4–8 weeks

WhatsApp & messaging assistant

An assistant on WhatsApp and other messaging channels that handles enquiries, bookings, reminders and updates.

What’s included

  • WhatsApp Business set-up and templates submitted to Meta
  • Conversation flows with AI for open questions
  • Opt-in and consent handling
  • Every conversation logged in your CRM
  • WhatsApp Business
  • Templates
  • Opt-in
  • Twilio

Best forBusinesses in markets such as India where customers prefer WhatsApp.

Typical timeline: 6–12 weeks

Voice AI agent

A phone agent that understands natural speech and books appointments or routes calls, with transfer to a person when needed.

What’s included

  • Speech recognition and a natural voice
  • Call flows and telephony integration
  • Transfer to your staff with a call summary
  • Recordings and transcripts, with caller consent
  • Voice
  • Telephony
  • Call summaries
  • Twilio
  • OpenAI

Best forClinics, service businesses and contact centres with high call volumes.

Typical timeline: 4–8 weeks

Lead qualification assistant

An assistant that answers product questions, qualifies inbound enquiries and books meetings for your sales team, at any hour.

What’s included

  • Qualification questions from your sales process
  • Product answers from approved content only
  • Calendar booking
  • CRM record with a conversation summary
  • Lead qualification
  • Booking
  • CRM
  • Salesforce
  • OpenAI

Best forB2B teams losing leads that arrive outside working hours.

03Document & vision AI

3 services
Typical timeline: 6–10 weeks

Document processing

Read invoices, forms, contracts and ID documents automatically and pull out the fields you need; low-confidence cases go to a person to check.

What’s included

  • OCR and field extraction
  • Classification by document type
  • Confidence thresholds and a review queue
  • Export to your ERP, CRM or accounting system
  • OCR
  • Extraction
  • Human review
  • AWS

Best forFinance, insurance, logistics and onboarding teams keying data in by hand.

Typical timeline: 8–16 weeks

Visual inspection & detection

Cameras and models that spot defects, count items or check compliance in images and video.

What’s included

  • Image dataset and labelling plan
  • Model training and evaluation
  • Edge or cloud deployment
  • Alerts and a review dashboard
  • Computer vision
  • Edge deployment

Best forManufacturing, retail and field operations with visual checks.

Typical timeline: 6–10 weeks

Contract & document review assistant

Summarise long documents, compare clauses with your standard positions and flag what needs a lawyer’s attention.

What’s included

  • Clause library and review playbook
  • Clause extraction and comparison
  • Risk flags with references to the text
  • Reviewer workflow
  • Summaries
  • Clause comparison
  • Reviewer-led
  • OpenAI
  • LlamaIndex

Best forLegal, procurement and compliance teams reviewing high volumes.

04Workflow automation

3 services
Typical timeline: 2–6 weeks

No-code workflow automation

Connect everyday tools such as email, forms, CRM and spreadsheets with n8n, Make or Zapier so routine steps run automatically.

What’s included

  • Process mapping and an automation shortlist
  • Workflows with error handling and alerts
  • AI steps for sorting, drafting and summaries
  • Documentation and team handover
  • n8n
  • Make
  • Zapier

Best forSmall and mid-size teams losing hours to copy-and-paste work.

Typical timeline: 8–14 weeks

Automation for critical processes

Custom automation for business-critical, high-volume work that recovers from failures, retries safely and records every step, built on Temporal or similar.

What’s included

  • Workflow design with exception routing
  • Failure recovery with retries and timeouts
  • Steps a named person approves
  • Monitoring and an audit trail
  • Temporal
  • Failure recovery
  • Exceptions to people

Best forOperations where a dropped step costs money: payments, claims, fulfilment.

Typical timeline: 4–8 weeks

Inbox & ticket triage

AI that reads incoming emails and tickets, sorts them by topic and urgency, drafts replies and routes each one to the right team.

What’s included

  • Classification by topic, urgency and sentiment
  • Draft replies for a person to review
  • Routing rules into your helpdesk
  • Accuracy reporting
  • Triage
  • Draft replies
  • Routing
  • OpenAI

Best forShared inboxes and support desks with slow first responses.

05AI in your product & model operations

4 services
Typical timeline: 6–12 weeks

AI features in your product

Summaries, drafting, recommendations and natural-language search designed into your product, with a fallback when the model is unsure.

What’s included

  • Feature discovery with your product team
  • Interface patterns for AI output and feedback
  • Model and prompt set-up with cost limits
  • Evaluation set and release gate
  • In-product AI
  • Fallbacks
  • Cost per request
  • OpenAI

Best forSoftware products adding AI features their customers will use.

Typical timeline: 8–12 weeks

Recommendations & personalisation

Suggest the next product, article or action for each user, based on what they and similar users do.

What’s included

  • Data review and feature design
  • Recommendation models
  • A/B testing against the current experience
  • Performance monitoring
  • Recommendations
  • A/B tested

Best forCommerce, media and learning platforms.

Typical timeline: 4–8 weeks

Model fine-tuning

Adapt a smaller open-weight or hosted model to your tone, format or narrow task, often matching a larger model at lower cost.

What’s included

  • Dataset preparation and cleaning
  • Fine-tuning and comparison against a baseline
  • Evaluation report
  • Deployment and versioning
  • Fine-tuning
  • Open-weight
  • Lower cost

Best forHigh-volume tasks where cost, speed or format consistency matter.

Typical timeline: 3–6 weeks to set up, then monthly

Evaluation & monitoring

Keep AI features accurate after launch with continuous testing and alerts when quality, speed or cost goes off target. Prompts and models are versioned.

What’s included

  • Pre-release tests for accuracy, relevance and safety
  • Monitoring of live usage
  • Prompt and model registry
  • Cost and response-time dashboards
  • Evaluation
  • Quality alerts
  • Versioning

Best forTeams running AI in production without quality tracking.

Your brief

Tick “Add to brief” on any service, choose a package, then continue. Or enquire about one service directly.

How we work with you

Ways to engage, from a question to an RFQ.

Ask a quick question, send a project brief or issue a formal RFQ. The lead for the work reads each one in full, and any services already in your brief go with it.

Or book a thirty-minute call

What are you sending?

  1. 01

    About 2 minutes4 required answers

    For a first conversation, a press request or anything that does not need a scope yet.

    You get A reply from a named lead

  2. 02Recommended

    About 8 minutes5 short steps

    Goals, audiences, a budget band and timing, so our first reply can outline the work.

    You get Options and a first scope after one call

  3. 03

    About 15 minutesYour documents attached

    Your documents, deadlines and the procurement and security rules the work must meet.

    You get Receipt confirmed and a named bid lead

How it is priced

Each package shows its pricing model. Work starts once a written scope and quote are agreed.

  • Typical length
    1–3 weeks
    Pricing
    Fixed fee
  • Typical length
    4–12 weeks
    Pricing
    Fixed price
  • Typical length
    3–9 months
    Pricing
    Fixed price per milestone
  • Typical length
    Ongoing · 6-month minimum
    Pricing
    Monthly fee
Compare what each package includes
What each package includes and who it suits
PackageEvery engagement includesBest for
SprintA short, fixed-scope engagement that answers one defined question.
  • Scope and outcome agreed before day one
  • A senior lead plus the specialists needed
  • A working review every week
  • A decision-ready answer or prototype
Discovery, a diagnostic, a prototype or a decision you need to make soon
ProjectA defined scope, delivered for a fixed price.
  • Statement of work with deliverables and acceptance criteria
  • A named project lead and a fixed team
  • A shared plan with dated checkpoints
  • Source files, yours once paid for
Work you can describe up front: an identity, a platform or a set of tools
MilestoneA larger build in phases you approve and pay for one at a time.
  • Phases with their own scope, output and sign-off
  • A go or no-go review at every gate
  • Re-planning between phases as you learn
  • Payment tied to accepted milestones
Programmes too large for one contract, where you want control at each step
RetainerReserved monthly capacity to run, improve and extend what we built.
  • A reserved block of team time every month
  • Agreed response times for requests and fixes
  • A monthly review and a rolling backlog
  • Planned improvements as well as upkeep
Live brands and products that need a steady team without hiring one

04.12Questions/faq

AI Product & Automation questions, answered.

Each answer links to the section of this page where you can see it working.

Not answered here?

Ask it on a call, and we will answer using your own documents and data.

Book a call

Any AI model can, so we design for it. Answers come only from retrieved sources, every sentence must cite one, and a check removes sentences that do not. A release that scores under 0.90 for faithfulness on your evaluation set is blocked. When sources are missing or disagree, the assistant says so and hands over to a person.

Sources on this page §04.2RAG inspector §04.7Quality gates

Yes, and it can cite them. Retrieval-augmented generation fetches the relevant passages before each answer. You need it when answers must follow your documents, policies or data rather than the model’s training.

Sources on this page §04.2RAG inspector §04.6Stack

Yes. Documents are indexed with their permissions, so the model never sees a passage the person asking could not open. Data stays in your cloud or with providers whose enterprise terms exclude training on it. Personal data is redacted before logging, and every answer is traced for audit.

Sources on this page §04.3Conversation §04.6Stack

Whichever scores best on your own examples for quality, speed and cost. It is rarely one model. A small model classifies and routes, and a larger one drafts. An open-weight model can be the fallback, or run everything when data must stay in-house. Swapping one is a configuration change and a test run.

Sources on this page §04.6Stack §04.8Cost & energy

Yes. A growing number of providers host models in Indian cloud regions. Open-weight models can also run on GPUs in an Indian data centre or your private cloud, so documents, prompts and logs stay in the country. Before building, we map the design against the DPDP Act 2023 and sector rules such as RBI requirements for payment data.

Sources on this page §04.6Stack

Answer from your documents when knowledge changes and must be cited. Fine-tune for style, format or narrow tasks, where a smaller model can match a larger one at lower cost. Many systems use both.

Sources on this page §04.6Stack §04.8Cost & energy

We measure it against an evaluation set built with your experts. Each answer is scored for accuracy against its sources, relevance, safety and whether the right passages were found. Automatic scores are checked against human review. After launch, the same measures run on a daily sample of live answers, beside the business numbers the feature should move.

Sources on this page §04.2RAG inspector §04.7Quality gates §04.11Outcomes

It is treated as a release and passes the full gates first. Versions are pinned, never “latest”. A provider update, a deprecation or a better open model then goes to 5% of traffic, widening only while live answers stay grounded. If a number slips, the flag rolls back to the last good version automatically.

Sources on this page §04.7Quality gates §04.11Outcomes

It depends on volume, the model and how much text each request carries. We estimate cost per request during the prototype, then reduce it before launch with caching, batching and smaller models for simple steps.

Sources on this page §04.8Cost & energy

Yes. Open-weight models can run in your cloud or on your own GPUs, served with vLLM or similar. Document search can run in PostgreSQL with pgvector or a dedicated vector database.

Sources on this page §04.6Stack

Tell us what you need built.

You will speak to a lead who would run the work, and get a straight answer on fit.

Book a call

Three ways to start

Every engagement starts with a written scope and a quote agreed before work begins.

Choose one of the three ways above