Capability 05 of 10 · Technology & Intelligence
AI infrastructure that answers fast, stays up and costs what you planned.
We build and run the infrastructure your products and AI depend on: cloud foundations, model serving and data pipelines. SRE and FinOps practices keep latency, uptime and spend within targets agreed with you.
A live telemetry dashboard for an example cluster called your-platform-prod across three regions, with illustrative values: GPU utilisation across six serving nodes averaging 68 percent with two idle warm-pool nodes, p95 latency of about 842 ms against a 1.5 second objective, output throughput of 18.4 thousand tokens per second, a blended cost of $1.21 per thousand requests against a $1.45 budget, and three regions with their share of traffic and grid carbon intensity.
02Cost and latencyboth are counted in tokens
Where AI bills and latency surprise teams.
AI model calls are billed and timed by the token. Most overruns come from waste nobody chose: prompts that keep growing, uncapped retries and idle GPUs.
Anatomy of one request
A support answer grounded in your documents
- Cost per 1k of this request
- $22.20
- baseline
- Time to first token
- 0.62 s
- baseline
- Monthly · 1.5M req
- $33,300
- baseline
| Part | Tokens | Rate per 1M | Cost |
|---|---|---|---|
| System prompt & tools | 1,200 | $3.00 | $0.0036 |
| Retrieved context | 3,800 | $3.00 | $0.0114 |
| User input | 150 | $3.00 | $0.0005 |
| Output | 450 | $15.00 | $0.0068 |
| Per request | 5,600 | $0.0222 |
Time to first token is prefill: the model reads every input token before replying, so long prompts wait longer.
Tokens per second is decode: the answer streams at a steady rate, so long answers take longer to finish.
Illustrative list prices: $3 per million input tokens, $0.30 per million cached input tokens and $15 per million output tokens. Rates vary by model, provider and commitment. Illustrative
One expensive call that answers from your documents. Averaged across every feature, the platform costs $1.21 per thousand requests; the FinOps console below shows the mix.
Showing the request as shipped: 5,600 tokens, $22.20 per thousand of this request, 0.62 seconds to first token.
-
01
Growing prompts
Chat history and retrieved passages are resent each turn, so turn 12 can carry seven times the input tokens of turn 1.
FixSummarise history, cap retrieval and cache the fixed part of the prompt.
-
02
Uncapped retries
Three blind retries on a slow provider can bill one answer up to four times and add load when capacity is short.
FixCapped retries with randomised backoff, idempotency keys and a fallback route.
-
03
Idle GPUs between peaks
Self-hosted GPUs sized for evening peaks run near a tenth of capacity overnight, and idle hours cost the same as busy ones.
FixAutoscale on queue length, scale to zero off-peak and run batch work in quiet hours.
03Reference architectureclients → gateway → router → serving → data
A reference architecture that scales and fails safely.
Every model call goes through one gateway you own. Routing, fallback, caching, token budgets and evaluation all sit there, so any model can be monitored, costed and replaced without changing product code.
- 01
Clients
Web and mobile apps, internal tools, agents and scheduled jobs call one internal AI API.
- 02
API gateway
Authenticates and logs every call (OIDC, JWT) and enforces rate limits and quotas per tenant.
- 03
AI gateway and router
Routes each request by cost, latency, quality and data class; falls back when a provider slows or fails; caches prompts and answers; caps tokens per team.
- 04
Model serving
Managed model APIs from more than one provider, plus open-weight models self-hosted on vLLM in Kubernetes GPU pools that scale with the queue.
- 05
Evaluation and canary
A new model takes shadow traffic, then a small canary share, and is promoted only if evaluation scores, p95 latency and cost hold.
- 06
Data and storage
A vector database for retrieval, a feature store and object storage for model weights, documents and logs.
- 07
Controls across every tier
OpenTelemetry traces, metrics and logs; secrets in Vault or a cloud KMS with short-lived credentials; policy as code for IAM, network and data residency.
-
Fastest to start
Managed model APIs
Choose when
- Volume is modest or spiky
- You need the most capable models today
- Nobody on the team runs GPUs
WatchPer-token cost at high volume, provider rate limits and data-processing terms.
-
Most control
Self-hosted open weights
Choose when
- Steady volume makes GPU-hours cheaper than tokens
- Data must stay in your account or in India
- You need tight latency or a fine-tuned model
WatchGPU capacity, on-call, model and driver upgrades.
-
Our default
Both, behind one gateway
Choose when
- One API for every product team
- Routes by cost, latency and data class
- Fallback when a provider degrades
Governed asRouter configuration is code, reviewed and tested like any other.
Model swaps
Swap a model in days, without a rebuild.
Because every call goes through the gateway, a new or cheaper model is a configuration change, tested on your own examples before release. Infrastructure changes go out weekly through reviewed Terraform and Helm pipelines.
- Day 1Candidate added to router config · tested offline on your evaluation set
- Day 2Shadow traffic · responses scored, never shown
- Day 35% canary · p95, cost and error budget watched
- Day 5Promoted, or rolled back in one commit
Technologies we work with
- AWS
- Microsoft Azure
04Failure drillchaos engineering in a controlled window
Failure drills that test recovery before an outage does.
Pick a failure and inject it into ten minutes of production-like traffic. Run it with the safeguards off, then on, and compare latency, errors and the error budget.

Interactive failure drill. Choose one of three injected failures, turn the resilience patterns on or off, then press Play or move the timeline slider. Four charts show successful requests per second, p95 latency, error rate and the error budget left for the month; the architecture map shows which component failed and where traffic goes; the incident log lists what happened and when. All values are illustrative.
10:00/ 10:00
Successful requests req/s
39
p95 latency ms
888 ms
Error rate %
0.08%
Error budget left % of month
62.0%
Architecturegpu-node-3 ejected · replica added
Incident log8 events
- Drill started · 40 req/s at peak · p95 860 ms
- gpu-node-3 stops responding (injected)
- Health check fails 3 × 3 s → node ejected from the pool
- In-flight requests retried once on healthy nodes
- Queue depth 180 · backpressure on · p95 1.3 s
- Autoscaler: +1 replica from the warm pool
- Replica ready · weights loaded from local cache
- Recovered without a page · budget used 0.03%
- Peak client-visible errors
- 4.0%
- Peak p95 latency
- 1.38 s
- Error budget used
- 0.03% of the month
- Fast-burn page
- Not triggered
-
The objective
99.9% monthly availability
This leaves room for 0.1% of requests to fail, equal to 43.2 minutes of full downtime in a 30-day month. That allowance is the error budget, and the team decides how to spend it.
-
The alert
Paging on budget burn rate
On-call is paged when the budget burns at 14.4× the sustainable rate over both 1 hour and 5 minutes, which spends 2% of the month in an hour. Slower burns (6× over 6 hours) page too; gentle ones open a ticket.
-
The safeguards
What “on” switches on
- Health checks that eject failed nodes
- A router with circuit breakers and fallback providers
- A request queue with backpressure and capped, jittered retries
- A semantic cache and an autoscaler with a warm pool
- Automated failover to a warm standby region
05Inference engineeringbatching · caching · quantisation · speculation
Four inference techniques that cut latency and GPU cost.
Each technique serves the same model on fewer GPU-seconds. It goes live behind the gateway only after an evaluation run shows quality has held.
-
01
Continuous batching
throughput ×2–4 per GPU
A scheduler with six GPU lanes, each decoding a request at about 75 tokens a second, 450 tokens a second between them, and a queue of three waiting requests. As a lane finishes, the next queued request joins the batch at once; a fixed batch would hold the queue for another 5.8 seconds.
Requests join the running batch as soon as a slot frees, so none waits for the slowest request in a fixed batch.
Quality gate evaluation set 0.94 → 0.94 · p95 −38% · pass
-
02
KV and prefix caching
time to first token −53%
Two requests shown as rows of twelve token blocks. The first request computes every block. The second request shares the same seven-block prefix, which is served from the cache, and only computes the five new blocks.
The system prompt, tools and examples repeat on every call, so their cache is kept and only new tokens are processed.
Quality gate evaluation set 0.94 → 0.94 · TTFT 0.62 s → 0.29 s · pass
-
03
Quantisation
memory −50% to −75%
8B model · one 24 GB GPU
- Weights
- 16 GB
- 4k-token seqs
- ~12
- Decode
- 1.0×
- Golden set
- 0.94
- Weights
- 8 GB
- 4k-token seqs
- ~28
- Decode
- 1.6×
- Golden set
- 0.93
- Weights
- 4.5 GB
- 4k-token seqs
- ~35
- Decode
- 2.1×
- Golden set
- 0.90
FP8 selected: 8 bits per weight, 8 gigabytes of weights on a 24 gigabyte GPU, room for about 28 concurrent 4,000-token sequences, 1.6 times decode speed, golden-set score 0.93, inside tolerance, ships.
Storing weights in fewer bits (FP8 or INT4) frees GPU memory for more requests per batch; the quality gate measures any loss.
Quality gate baseline · ships−0.01 · inside the 0.02 tolerance · ships−0.04 · below tolerance · held for review
-
04
Speculative decoding
decode latency −45% to −60%
A draft model proposes four tokens at a time. The target model checks all four in one pass, keeps the accepted ones and adds one token of its own; the rest is discarded. The output line grows by the accepted tokens: twelve tokens take three target passes instead of twelve.
A draft model proposes several tokens and the main model verifies them in one pass, so output is unchanged but arrives sooner.
Quality gate output unchanged by construction · decode −56% · pass
Serving engines we work with
Serving engines release new versions monthly. Each is load-tested and passes the quality gate on a canary pool before the router sends it live traffic.
- vLLM PagedAttention · continuous batching · prefix caching
- NVIDIA TensorRT-LLM in-flight batching · FP8 · speculative decoding
- SGLang RadixAttention prefix cache · structured output
- Ray Serve multi-model serving · autoscaling replicas
- Hugging Face open-weight models · TGI
Gains are typical ranges for chat and retrieval workloads and depend on the model, prompt shape and traffic. Illustrative
06The stackeight layers, from GPUs at the bottom to security at the top
The AI stack, from GPUs to security.
Eight layers, each with one job, an owner and the tools we use. Open any layer to see what it holds and what we decide with you there.
U08 · tray open0 critical CVEs
Security & policy
Vault or the cloud KMS issues each workload short-lived credentials, so no secret lives in code or images. Images are scanned in CI, running workloads are monitored and pods that run as root or from unsigned images are refused.
Technologies we work with here
- Vault
- Trivy
- Falco
- Snyk
- OPA Gatekeeper
- Kyverno
Decided with you at this layer
- Workload identity instead of static keys
- Signed images only in production namespaces
- Model weights, prompts and logs treated as sensitive data
U07 · tray open14 SLO alerts
Observability & FinOps
One OpenTelemetry pipeline carries traces, metrics and logs from gateway to GPU. Each trace records tokens, cache hits and cost, so a slow or costly request leads to its feature, team and model.
Technologies we work with here
- OpenTelemetry
- Prometheus
- Grafana
- Datadog
- PagerDuty
- OpenCost
Decided with you at this layer
- GenAI semantic conventions on every model call
- Alerts on error-budget burn rate
- Spend allocated per team and feature, daily
U06 · tray open31% cache hits
Gateway & routing
Every model call passes through it for authentication, quotas, routing by cost and data class, provider fallback, caching and token budgets per team. Its configuration is versioned code.
Technologies we work with here
- Kong
- Cloudflare
- Redis
- LiteLLM
- Envoy AI Gateway
Decided with you at this layer
- Which data classes may leave the country
- Fallback order and timeouts per route
- Cache lifetimes and similarity thresholds
U05 · tray open12.4M vectors
Data & vectors
Retrieval quality depends on infrastructure too: vector indexes sized for recall within the latency budget, embeddings versioned with their model and change streams that keep the index minutes behind the source.
Technologies we work with here
- Qdrant
- PostgreSQL
- Milvus
- Apache Kafka
- pgvector
Decided with you at this layer
- Index parameters tuned against recall@10
- Re-embed only what changed
- Tenant isolation inside the index
U04 · tray open3 models live
Model serving
Open-weight models run on vLLM or TensorRT-LLM behind an OpenAI-compatible API, so the router treats self-hosted and managed models alike. Replicas scale on queue length and time to first token.
Technologies we work with here
- vLLM
- NVIDIA
- Hugging Face
- Ray
- SGLang
- Triton
Decided with you at this layer
- Model, precision and context length per route
- Prefix caching and batch limits
- Autoscaling on vllm:num_requests_waiting
U03 · tray open6 node pools
Kubernetes & IaC
GPU node pools are reserved for inference pods. Terraform builds the cloud, Helm packages workloads and Argo CD keeps every cluster in line with Git. Nothing in production is changed by hand.
Technologies we work with here
- Kubernetes
- Helm
- Terraform
- Argo
- Docker
- Karpenter
- KEDA
Decided with you at this layer
- GPU pools with a warm spare node
- Scale to zero for dev and preview environments
- Every change through a reviewed pull request
U02 · tray open3 regions · 2 in India
Cloud & regions
We build on the cloud you already use, in the regions your data rules allow. A landing zone sets identity, networking, logging and security policies once for every workload. Edge caching and a web application firewall sit in front.
Technologies we work with here
- AWS
- Microsoft Azure
- Google Cloud
- Cloudflare
Decided with you at this layer
- Primary and standby regions
- A mix of committed, on-demand and spot capacity
- Private networking to model providers where offered
U01 · tray open8 × L40S · 68%
Accelerators
We choose the smallest accelerator that fits the model and its cache at your load. An 8B model at FP8 suits an L4 or L40S; 70B-class models and long contexts need an H100 or H200. Inferentia and TPU are tested on your traffic before you commit.
Technologies we work with here
- NVIDIA
- L4 · 24 GB
- L40S · 48 GB
- H100 · 80 GB
- H200 · 141 GB
- AWS Inferentia2
- Google Cloud TPU
Decided with you at this layer
- GPU class per model and context length
- Reserved baseline, on-demand peaks, spot for batch
- A 60–80% utilisation band on serving pools

Data residency
India regions we deploy to
Residency is set as policy at the gateway: requests tagged as personal data cannot go to a model or region outside India. Batch jobs on non-personal data may run in lower-carbon or cheaper regions.
| Cloud | Region | Location | Code |
|---|---|---|---|
| AWS | Asia Pacific (Mumbai) | Maharashtra | ap-south-1 |
| AWS | Asia Pacific (Hyderabad) | Telangana | ap-south-2 |
| Microsoft Azure | Central India | Pune | centralindia |
| Microsoft Azure | South India | Chennai | southindia |
| Google Cloud | Mumbai | Maharashtra | asia-south1 |
| Google Cloud | Delhi | Delhi NCR | asia-south2 |
GPU types and capacity differ by region and change often; we confirm availability before a design is committed.
07FinOpscost per unit · showback · forecast · anomalies
Cost per request, visible before the invoice.
Every model call is tagged with its feature, team and route, so spend can be read per request. An agent checks the figures daily; a named person approves any change it proposes.
- Spend to date
- $26.3k
- day 19 of 30
- Forecast
- $41.8k
- ± $2.9k at 80%
- Budget
- $45.0k
- 7% headroom
- Cost per 1k requests
- $1.21
- −14% vs August
- Unallocated
- 3%
- of spend without a tag
Unit cost
Cost per 1k requests against the cost model
| Name | Against target | Actual | Target | Trend |
|---|---|---|---|---|
| Support assistant412k requests | $1.42 | $1.60 | −12% | |
| Document search1.1M requests | $0.38 | $0.50 | −4% | |
| Sales email drafts58k requests | $2.94 | $2.20 | +31% | |
| Invoice extraction96k documents | $0.71 | $0.90 | −9% | |
| Internal copilot203k requests | $1.18 | $1.20 | +2% |
Spend forecast
$k · cumulative
Month-to-date spend of 26.3 thousand dollars on day 19, forecast to reach 41.8 thousand dollars by day 30, with an 80 percent range of 38.9 to 44.6 thousand, under a budget of 45 thousand.
GPUs in use against reserved capacity
last 7 days · 2 h buckets
Over seven days, serving used an average of 5.4 of 8 reserved GPUs, about 68 percent. 4 afternoon peaks rose above the reservation onto on-demand GPUs, and nightly batch jobs ran on 4 spot GPUs for 168 GPU-hours.
- 5.4 avgserving on reserved · 68% of the reservation
- 4 peaksburst to on-demand
- 168 GPU-hbatch on spot, about 65% cheaper
Anomaly · yesterday 01:00–03:10
+$1,240 vs forecast
Embedding job re-ran 3×
embed-docs-nightly timed out listing the object store and retried the whole job instead of the failed batch, three times over.
Proposed by finops-agent · confidence 0.92
jobs/embed-docs.yaml - retries: 3 + retries: 1 + retryScope: batch + idempotencyKey: "{{ .source.version }}"
Awaiting your approval · the agent cannot merge
- finops-agentAnomaly: embedding spend 3.1× forecast (z = 4.2)
- finops-agentCause: 3 full re-runs after an object-store timeout
- finops-agentFix proposed · awaiting human approval
- youApproved · PR #482 opened · change CHG-1182 logged
- youHeld for review · assigned to platform on-call
-
Cost reports per team
Tags set at the gateway assign every token and GPU-hour to a feature and a team. Each team sees its own costs weekly, before finance sees the total.
-
Reserved, on-demand and spot
Steady serving runs on reserved or committed capacity; peaks use on-demand; batch jobs, evaluation and re-embedding run on spot with checkpoints, so an interruption costs minutes.
-
Scale to zero
Development, preview and test environments scale to zero outside working hours. The first request afterwards waits for start-up, so production is never scaled to zero.
Practices we align delivery with
- FinOps FrameworkFinOps Foundation · inform, optimise, operate
- FOCUSFinOps Open Cost and Usage Specification · one billing schema across clouds
08Sustainabilitysoftware carbon intensity, per unit of work
AI carbon emissions, cut by scheduling.
The same job emits more or less carbon depending on when and where it runs. Live requests stay close to users; batch work moves to lower-carbon hours and regions within your residency rules.
A chart of grid carbon intensity across one day for Mumbai and Chennai, both highest overnight and in the evening and lowest around midday, when solar output peaks; Chennai sits lower throughout. The embedding refresh is scheduled in the shaded four-hour window.
Inferencelatency-sensitive · stays near usersall day
Embedding refresh4 h · 22 kWh · drag the block or use the slider11:00–15:00
Backup snapshotfixed · keeps the RPO02:00 · 14:00
- Grid intensity, I
- 465g/kWh
- Job emissions
- 12.1kgCO₂e
- SCI per 1k documents
- 30.3gCO₂e
- −33% vs the 01:00 Mumbai cron
SCI = ((E × I) + M) per RE 22 kWh · I 465 g/kWh · M 1.9 kg · R 1k documents
Both regions are in India, so the documents stay in the country. Intensities are illustrative; in production, grid intensity comes from a grid-data API and energy from GPU telemetry. Illustrative
09Reliabilityservice level objectives · error budgets · on-call
Reliability targets we set with you, and an error budget we track.
Uptime is a target you choose and pay for. We agree the measures and targets with you, publish the error budget and let it decide when to release features or fix reliability.
Availability
Successful responses ÷ all responses, measured at the gateway, excluding client 4xx.
- Objective
- 99.9%
- over 30 days
- Observed
- 99.96%
- this window
- Budget
- 43.2 min
- total for 30 days
- Remaining
- 25.8 min
- 59.7% of budget
An error-budget burn-down for Availability across thirty days, falling from the full budget to 59.7% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom and one step drop where an incident burned budget quickly.
One provider incident on day 12 burned 6.4 minutes. The multi-provider router carried the rest of the traffic, so the budget survived it.
p95 latency
Share of non-streaming requests completed end to end within the target, per endpoint class.
- Objective
- 99% < 1.5 s
- p95 observed 1.21 s
- Observed
- 99.38%
- this window
- Budget
- 345k req
- total for 30 days
- Remaining
- 131k req
- 37.9% of budget
An error-budget burn-down for p95 latency across thirty days, falling from the full budget to 37.9% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom and one step drop where an incident burned budget quickly.
A day-20 release raised p95 by 340 ms. Fast burn paged on-call inside the hour, the release was rolled back, and the budget is being repaid before the next feature ships.
Time to first token
Share of streaming responses that put their first token on the wire within the target.
- Objective
- 95% < 800 ms
- p95 observed 612 ms
- Observed
- 96.8%
- this window
- Budget
- 965k streams
- total for 30 days
- Remaining
- 623k streams
- 64.6% of budget
An error-budget burn-down for Time to first token across thirty days, falling from the full budget to 64.6% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom.
Prompt caching and continuous batching hold first token steady through the evening peak, so the budget burns close to the ideal line.
Burn-rate alerting
A fast burn (14.4× the budget rate for one hour) pages on-call. A slow burn (6× over six hours) opens a ticket for the next working day. Alerts follow the service indicator, never a single failing host.
The freeze rule
With under 25% of the budget left, feature releases stop and the team spends the sprint on reliability. Everyone agrees this in advance, so nobody argues it during an incident.
Illustrative figures
The window, targets and thresholds are the ones we use; the plotted values are sample data for a sample platform. Illustrative
-
Runbooks that match the alerts
Written during the build, each runbook gives the first five commands, the dashboards to open and the rollback.
-
A shared on-call rotation
A primary and a secondary engineer, an escalation policy and handover notes at each shift change. Your engineers can share the rotation with ours.
-
Blameless post-incident reviews
Within five working days: timeline, contributing causes, what monitoring missed and actions with owners, scheduled into the next sprint.
-
Tested disaster recovery
Recovery targets (RTO, RPO) are agreed per system, then tested each quarter with a timed, documented rebuild from code and backups.

Frameworks we align delivery with
Controls, continuity and hardening baselines the platform is built to. Certifying your own environment is a separate project we can prepare you for.
- ISO/IEC 27001:2022Information security management systems
- SOC 2Trust Services Criteria
- CIS Controls v8.1Critical Security Controls and Benchmarks
- NIST CSF 2.0Cybersecurity Framework
- ISO 22301:2019Business continuity management systems
Support hours are agreed per engagement (business hours, extended hours or 24×7) and written into the runbook with the escalation path.
10How we workbaseline · redesign · operate
Measure first, then improve latency, uptime and cost..
We record latency, availability, spend and carbon at the start, so you can check every improvement against them. Nothing is rebuilt for looking old. Work starts where measurement shows the latency, spend or risk sits.
-
01
Baseline
Wk 01–02
We map today's set-up, measure traffic, latency, incidents, spend and carbon, then agree targets and SLOs with you.
- Baseline report
- SLO targets
- Cost breakdown
-
02
Design
Wk 02–04
We design the target architecture, serving approach, capacity plan and migration path, recording each decision against the baseline.
- Target architecture
- Capacity model
- Decision records
-
03
Build
Wk 04–10
We build and load-test the infrastructure code, pipelines and serving stack. Monitoring and cost tags start with the first resource.
- Infrastructure code
- Load-test results
- Dashboards
-
04
Run
Wk 10–12
We switch traffic over and write on-call runbooks. A monthly review of SLOs, spend and capacity decides the next improvements.
- Runbooks
- SLO reviews
- FinOps report
The three numbers a phase is judged on
Measured on your traffic before the first change and after each release. Each bar shows the reduction from the starting figure, so the three can be compared. One illustrative engagement shown. Illustrative
-
p95 latency, chat endpoint
2.9 s 1.21 s
−58%
Continuous batching, a prompt cache and a router that sends short requests to a smaller model.
-
Cost per 1,000 requests
$2.95 $1.21
−59%
Right-sized GPU pools, quantised weights and caching, with batch work moved to committed capacity.
-
Idle GPU-hours a week, non-production
96 h 12 h
−88%
Non-production environments scale to zero outside working hours.
After go-live, reviews continue: a monthly SLO review of the error budget and incident actions; a FinOps review of unit cost and commitments; a capacity check before seasonal peaks.
11What you getinfrastructure as code, in your repository
Everything we build lands in a repository you own.
What we make for you is yours once it is paid for. Tools we already had stay ours, and you get a free, permanent licence to use them. The platform is code, rebuildable from an empty account.
A browsable listing of the infrastructure repository handed over at the end of the engagement, with 9 files you can open: Terraform for the GPU pool and the residency policy, Helm values for model serving and the AI gateway, a Grafana cost dashboard, the SLO objectives with their burn-rate alert rules, the GPU node-failure runbook, a k6 peak-load profile and the disaster-recovery drill script.
Arrow keys move, → opens a directory, ← closes it, Enter shows a file.
terraform/envs/prod/main.tf
01module "gpu_pool" {02 source = "../../modules/gpu-pool"03 04 region = "ap-south-1" # Mumbai, primary05 instance_type = "g6e.2xlarge" # 1 × L40S, 48 GB06 min_nodes = 607 max_nodes = 10 # 6 serving + 2 warm + 2 burst08 warm_nodes = 209 taints = ["nvidia.com/gpu=present:NoSchedule"]10 tags = local.cost_tags # feature, team, model route11}
11 linesPR #482 · platform · plan attached, 26 policy checks passed
terraform/policy/region.rego
01package terraform.region02 03# Personal data never leaves India. Any resource planned04# outside the approved regions fails the policy check in CI.05approved := {"ap-south-1", "ap-south-2"}06 07deny[msg] {08 r := input.resource_changes[_]09 r.change.after.region10 not approved[r.change.after.region]11 msg := sprintf("%s targets %s, outside the residency set",12 [r.address, r.change.after.region])13}
13 linesPR #455 · security · ap-south-2 added to the residency set
helm/vllm-serving/values.yaml
01model:02 name: your-platform/chat-8b03 quantization: fp8 # eval gate: −0.01, inside the 0.02 tolerance04 maxModelLen: 819205 gpuMemoryUtilization: 0.9206 enablePrefixCaching: true07 08autoscaling:09 metric: vllm:num_requests_waiting # queue depth, never CPU10 target: "4"11 minReplicas: 612 maxReplicas: 1013 14podDisruptionBudget:15 minAvailable: 5
15 linesPR #501 · serving · FP8 after the eval gate came back −0.01
helm/ai-gateway/values.yaml
01routes:02 - match: { dataClass: personal }03 provider: self-hosted # residency rule, enforced here04 - match: { tokensIn: "<800" }05 provider: small-model06 - default: provider-a07 08retries:09 budget: 10% # of a route's requests per minute10 backoff: exponential-jitter11 idempotencyKey: required12 13cache:14 prompt: { ttl: 1h }15 semantic: { threshold: 0.93, ttl: 15m }
15 linesPR #498 · platform · retry budget capped at 10% of the route
dashboards/finops.json
01{02 "title": "Cost per 1,000 requests",03 "datasource": "prometheus",04 "targets": [05 {06 "expr": "sum by (feature) (rate(gen_ai_cost_usd_total[1h]))07 / sum by (feature) (rate(gen_ai_requests_total[1h])) * 1000",08 "legendFormat": "{{feature}}"09 }10 ],11 "thresholds": [{ "value": 1.43, "colorMode": "text" }],12 "description": "Cost tags are set at the gateway. Spend without a13 tag shows up as {feature=\"\"} and gets chased."14}
14 linesPR #471 · finops · cost tag added to the gateway span
slo/objectives.yaml
01- name: availability02 sli: successful_responses / all_responses # at the gateway, 4xx excluded03 objective: 99.904 window: 30d # 43.2 minutes of budget05 burnRateAlerts:06 - { factor: 14.4, long: 1h, short: 5m, severity: page }07 - { factor: 6, long: 6h, short: 30m, severity: ticket }08 freezeBelowBudget: 25 # % remaining — feature releases stop09 10- name: time_to_first_token11 sli: streams_first_token_under_800ms / streams12 objective: 9513 window: 30d
13 linesPR #440 · SRE · burn-rate pair agreed with you, then merged
runbooks/gpu-node-failure.md
01# GPU node failure02 03Alert: vllm_node_unhealthy · Dashboard: SLO / serving04Owner: platform on-call · Expected impact: none, the warm pool absorbs it05 06## First five commands07 081. kubectl get nodes -l nvidia.com/gpu=present -o wide092. kubectl describe node $NODE | sed -n '/Conditions/,/Events/p'103. kubectl drain $NODE --ignore-daemonsets --delete-emptydir-data114. kubectl -n serving get hpa vllm-serving125. k6 run loadtest/k6/smoke.js13 14## If p95 stays above 1.5 s for ten minutes15 16Shift the chat route to provider A in the gateway, page the model17owner, and open an incident. The drill for this is in section 04.
17 linesPR #466 · on-call · the five commands verified in a drill
loadtest/k6/peak.js
01import http from 'k6/http';02import { check } from 'k6';03 04export const options = {05 scenarios: {06 peak: { executor: 'constant-arrival-rate', rate: 40,07 timeUnit: '1s', duration: '20m', preAllocatedVUs: 300 },08 },09 thresholds: {10 'http_req_duration{endpoint:chat}': ['p(95)<1500', 'p(99)<3000'],11 'http_req_failed': ['rate<0.01'],12 },13};14 15export default function () {16 const res = http.post(`${__ENV.BASE}/v1/chat`, PAYLOAD,17 { tags: { endpoint: 'chat' } });18 check(res, { 'first token under 800 ms': (r) => r.timings.waiting < 800 });19}
19 linesPR #489 · quality · peak profile raised to 40 req/s
dr/restore-drill.sh
01#!/usr/bin/env bash02set -euo pipefail03 04# Quarterly drill: rebuild production in the standby region from05# code and backups only, and time it against the agreed RTO.06START=$(date +%s)07 08terraform -chdir=terraform/envs/dr apply -auto-approve09helm upgrade --install platform helm/platform -f helm/platform/dr.yaml10./dr/restore-vectors.sh --snapshot "$(date -u +%F)" # RPO: hourly11k6 run loadtest/k6/smoke.js12 13echo "RTO $(( $(date +%s) - START ))s against a 3600s target."14echo "Evidence written to dr/evidence/$(date -u +%F)/."
14 linesPR #430 · platform · last quarter's RTO evidence attached
Open any file to see its contents. Paths and values are illustrative; the structure and the arithmetic behind the thresholds are what we hand over. Illustrative
Handover pack
- Infrastructure as code modulesTerraform · Helm
- Reference architecture & decision recordsDiagrams · ADRs
- Model serving stackContainers · config
- CI/CD & GitOps pipelinesPipelines · Argo CD
- SLOs, dashboards & alertsGrafana · alert rules
- Load & failover test reportsk6 · report
- Cost & carbon allocation reportDashboard · spreadsheet
- Change goes through reviewEvery environment change is a pull request with a plan attached, a policy check and a named approver. Nothing reaches production from a laptop.
- Your accounts, your keysThe cloud accounts, registries, secret stores and provider keys are yours from day one. We work inside them with scoped, auditable access.
- Decisions are written downArchitecture decision records explain what was chosen, what was rejected and what would change the answer, so a future team has the context.
Built and operated with
- Terraform
- Helm
- Kubernetes
- Argo
- GitHub Actions
- Vault
- Grafana
- k6
12Outcomesreviewed monthly, against the same definitions
Five numbers, measured every month.
We agree these targets and their definitions at the start. Each monthly review shows every reading with its evidence, including any that are out of band.
-
p95 latency, chat endpoint
Ninety-nine per cent of chat requests answered inside the agreed p95, measured per endpoint class.
1.21 s was 1.34 s Inside the objective
0 s agreed band 0–1.5 s 3 s
Set per endpoint class and held under load tests at forecast peak, rather than assumed from average traffic.
-
Cost per 1,000 requests
Within ten per cent either side of the cost model agreed at forecast volume.
$1.21 was $1.41 On the cost model
$0.00 agreed band $1.17–$1.43 $2.40
Tracked per feature and per team, so a price rise or a prompt that grew shows up as a line, not a surprise invoice.
-
GPU utilisation at peak
Between 55 and 75 per cent at the daily peak — busy enough to be economic, with headroom left.
81% was 68% Above the band
0% agreed band 55–75% 100%
Below the band you are paying for idle accelerators. Above it, queues build and tail latency runs away.
Why it is out, and what happens next. A launch in week three tripled one route’s traffic. The pool is being resized and the autoscaler’s queue-depth target lowered — owner and date are in the review log.
-
SCI per 1,000 requests
Lower each quarter against a ceiling agreed at the start of it, measured the same way each time.
26.4 g was 29.8 g Trending down
0 g agreed band 0–32 g 60 g
Carbon falls with the same levers as cost: batching, caching, quantisation, and batch work scheduled into cleaner hours.
-
Error budget spent
No month where the availability budget is exhausted without a decision behind it.
40.3% was 52% Inside the allowance
0% agreed band 0–60% 100%
A planned burn during a migration is a decision. An unplanned one is a design or a monitoring gap, and it gets a review.
-
01
Latency on target
Each endpoint gets 95th and 99th percentile latency targets, held under load with autoscaling tested.
-
02
Reliability you can quote
SLOs with error budgets that decide when to release features and when to fix reliability.
-
03
Spend each team can see
Cost per request, user and model is visible to the team that controls it, and trends down.
Bands and readings are illustrative sample data, including the one out of band. Your targets are set against your own baseline in the first phase and reported in the same review each month. Illustrative
Services & packages
Cloud and AI infrastructure that stays fast, reliable and affordable.
Cloud, platform, AI serving and reliability services. Buy one on its own or combine several in a single brief.
- Categories
- 04
- Services
- 16
- Packages
- 04
How to buy
- 01Pick services. Enquire about one, or add several to a brief.
- 02Choose a package. A sprint, a fixed project or an ongoing team.
- 03Send the brief. We reply within one working day.
How we work with you
Ways to engage, from a question to an RFQ.
Ask a quick question, send a project brief or issue a formal RFQ. The lead for the work reads each one in full, and any services already in your brief go with it.
Or book a thirty-minute call-
01
For a first conversation, a press request or anything that does not need a scope yet.
You get A reply from a named lead
-
02Recommended
Goals, audiences, a budget band and timing, so our first reply can outline the work.
You get Options and a first scope after one call
-
03
Your documents, deadlines and the procurement and security rules the work must meet.
You get Receipt confirmed and a named bid lead
How it is priced
Each package shows its pricing model. Work starts once a written scope and quote are agreed.
Opens a project brief with this package chosen.
-
In your brief
- Typical length
- 1–3 weeks
- Pricing
- Fixed fee
-
In your brief
- Typical length
- 3–9 months
- Pricing
- Fixed price per milestone
-
In your brief
- Typical length
- Ongoing · 6-month minimum
- Pricing
- Monthly fee
-
In your brief
- Typical length
- 6–18 months
- Pricing
- Programme fee · by statement of work
Compare what each package includes
| Package | Every engagement includes | Best for |
|---|---|---|
| SprintA short, fixed-scope engagement that answers one defined question. |
|
Discovery, a diagnostic, a prototype or a decision you need to make soon |
| MilestoneA larger build in phases you approve and pay for one at a time. |
|
Programmes too large for one contract, where you want control at each step |
| RetainerReserved monthly capacity to run, improve and extend what we built. |
|
Live brands and products that need a steady team without hiring one |
| EnterpriseA multi-workstream programme with a dedicated team, governance and agreed service levels. |
|
Large organisations running change across markets, portfolios or business units |
13Questionsanswered plainly
AI infrastructure, your questions answered.
Short answers on hosting, reliability, cloud choice, capacity, cost and carbon for AI in production.
-
01 Should we self-host models or use APIs?
APIs are fastest to start and hard to beat at low volume. Self-hosting open-weight models pays off with steady high volume, strict data residency or tight latency needs. We model the break-even on your traffic before recommending either.
-
02 What does an SLO and error budget look like?
An SLO is a target such as 99.9% of requests succeeding within 800 ms over 30 days. The error budget is the remaining 0.1%, about 43 minutes of full downtime in a 30-day month. When it is spent, reliability work takes priority over new features.
-
03 Which cloud do you recommend?
The one that fits your existing contracts, data location, skills and the AI services you need. We work across AWS, Azure and Google Cloud and design for portability where it is worth the cost.
-
04 How do you reduce GPU costs?
By matching capacity to demand and doing less work per request. We right-size instances, batch requests, cache prompts and responses, quantise models and send simple requests to smaller models. Where the workload allows, we scale to zero or use spot or committed capacity.
-
05 Do you measure carbon?
Yes, using the Green Software Foundation’s Software Carbon Intensity (SCI) method. We report emissions per request and reduce them alongside cost.
-
06 Do you plan Kubernetes and GPU capacity?
Yes. We design the cluster, node pools per workload type and safe GPU sharing, sized from your own traffic. The plan shows what happens at forecast peak, what the warm pool costs and where a queue absorbs a spike instead of a new node.
Tell us what you need built.
You will speak to a lead who would run the work, and get a straight answer on fit.
Three ways to start
Every engagement starts with a written scope and a quote agreed before work begins.
Choose one of the three ways above



