Capability 05 of 10 · Technology & Intelligence

AI infrastructure that answers fast, stays up and costs what you planned.

We build and run the infrastructure your products and AI depend on: cloud foundations, model serving and data pipelines. SRE and FinOps practices keep latency, uptime and spend within targets agreed with you.

A live telemetry dashboard for an example cluster called your-platform-prod across three regions, with illustrative values: GPU utilisation across six serving nodes averaging 68 percent with two idle warm-pool nodes, p95 latency of about 842 ms against a 1.5 second objective, output throughput of 18.4 thousand tokens per second, a blended cost of $1.21 per thousand requests against a $1.45 budget, and three regions with their share of traffic and grid carbon intensity.

Typical length
4–12 weeks per phase
Clouds
AWS · Azure · Google Cloud
Standard
SLOs and budgets from day one

02Cost and latencyboth are counted in tokens

Where AI bills and latency surprise teams.

AI model calls are billed and timed by the token. Most overruns come from waste nobody chose: prompts that keep growing, uncapped retries and idle GPUs.

Anatomy of one request

A support answer grounded in your documents

Cost per 1k of this request
$22.20
baseline
Time to first token
0.62 s
baseline
Monthly · 1.5M req
$33,300
baseline
Tokens, rate and cost for each part of one request
PartTokensRate per 1MCost
System prompt & tools 1,200 $3.00 $0.0036
Retrieved context 3,800 $3.00 $0.0114
User input 150 $3.00 $0.0005
Output 450 $15.00 $0.0068
Per request5,600$0.0222

Time to first token is prefill: the model reads every input token before replying, so long prompts wait longer.

Tokens per second is decode: the answer streams at a steady rate, so long answers take longer to finish.

Illustrative list prices: $3 per million input tokens, $0.30 per million cached input tokens and $15 per million output tokens. Rates vary by model, provider and commitment. Illustrative

One expensive call that answers from your documents. Averaged across every feature, the platform costs $1.21 per thousand requests; the FinOps console below shows the mix.

Showing the request as shipped: 5,600 tokens, $22.20 per thousand of this request, 0.62 seconds to first token.

  • 01

    Growing prompts

    Chat history and retrieved passages are resent each turn, so turn 12 can carry seven times the input tokens of turn 1.

    FixSummarise history, cap retrieval and cache the fixed part of the prompt.

  • 02

    Uncapped retries

    Three blind retries on a slow provider can bill one answer up to four times and add load when capacity is short.

    FixCapped retries with randomised backoff, idempotency keys and a fallback route.

  • 03

    Idle GPUs between peaks

    Self-hosted GPUs sized for evening peaks run near a tenth of capacity overnight, and idle hours cost the same as busy ones.

    FixAutoscale on queue length, scale to zero off-peak and run batch work in quiet hours.

03Reference architectureclients → gateway → router → serving → data

A reference architecture that scales and fails safely.

Every model call goes through one gateway you own. Routing, fallback, caching, token budgets and evaluation all sit there, so any model can be monitored, costed and replaced without changing product code.

  1. 01

    Clients

    Web and mobile apps, internal tools, agents and scheduled jobs call one internal AI API.

  2. 02

    API gateway

    Authenticates and logs every call (OIDC, JWT) and enforces rate limits and quotas per tenant.

  3. 03

    AI gateway and router

    Routes each request by cost, latency, quality and data class; falls back when a provider slows or fails; caches prompts and answers; caps tokens per team.

  4. 04

    Model serving

    Managed model APIs from more than one provider, plus open-weight models self-hosted on vLLM in Kubernetes GPU pools that scale with the queue.

  5. 05

    Evaluation and canary

    A new model takes shadow traffic, then a small canary share, and is promoted only if evaluation scores, p95 latency and cost hold.

  6. 06

    Data and storage

    A vector database for retrieval, a feature store and object storage for model weights, documents and logs.

  7. 07

    Controls across every tier

    OpenTelemetry traces, metrics and logs; secrets in Vault or a cloud KMS with short-lived credentials; policy as code for IAM, network and data residency.

  • Fastest to start

    Managed model APIs

    Choose when

    • Volume is modest or spiky
    • You need the most capable models today
    • Nobody on the team runs GPUs

    WatchPer-token cost at high volume, provider rate limits and data-processing terms.

  • Most control

    Self-hosted open weights

    Choose when

    • Steady volume makes GPU-hours cheaper than tokens
    • Data must stay in your account or in India
    • You need tight latency or a fine-tuned model

    WatchGPU capacity, on-call, model and driver upgrades.

  • Our default

    Both, behind one gateway

    Choose when

    • One API for every product team
    • Routes by cost, latency and data class
    • Fallback when a provider degrades

    Governed asRouter configuration is code, reviewed and tested like any other.

Model swaps

Swap a model in days, without a rebuild.

Because every call goes through the gateway, a new or cheaper model is a configuration change, tested on your own examples before release. Infrastructure changes go out weekly through reviewed Terraform and Helm pipelines.

  1. Day 1Candidate added to router config · tested offline on your evaluation set
  2. Day 2Shadow traffic · responses scored, never shown
  3. Day 35% canary · p95, cost and error budget watched
  4. Day 5Promoted, or rolled back in one commit

Technologies we work with

04Failure drillchaos engineering in a controlled window

Failure drills that test recovery before an outage does.

Pick a failure and inject it into ten minutes of production-like traffic. Run it with the safeguards off, then on, and compare latency, errors and the error budget.

Source code on a dark monitor in a dim room lit by a red and blue glow
Rehearsed in advanceDrills run in working hours from a written runbook, so the night call is one the team has rehearsed.Photo: Jakub Żerdzicki / Unsplash

Interactive failure drill. Choose one of three injected failures, turn the resilience patterns on or off, then press Play or move the timeline slider. Four charts show successful requests per second, p95 latency, error rate and the error budget left for the month; the architecture map shows which component failed and where traffic goes; the incident log lists what happened and when. All values are illustrative.

Inject

10:00/ 10:00

Successful requests req/s

39

60 0

p95 latency ms

888 ms

SLO 1.5 s ≥ 4 s 0 ms

Error rate %

0.08%

100.0% 0.00%

Error budget left % of month

62.0%

65.0% 40.0%

Architecturegpu-node-3 ejected · replica added

Incident log8 events

  1. Drill started · 40 req/s at peak · p95 860 ms
  2. gpu-node-3 stops responding (injected)
  3. Health check fails 3 × 3 s → node ejected from the pool
  4. In-flight requests retried once on healthy nodes
  5. Queue depth 180 · backpressure on · p95 1.3 s
  6. Autoscaler: +1 replica from the warm pool
  7. Replica ready · weights loaded from local cache
  8. Recovered without a page · budget used 0.03%
Peak client-visible errors
4.0%
Peak p95 latency
1.38 s
Error budget used
0.03% of the month
Fast-burn page
Not triggered

  • The objective

    99.9% monthly availability

    This leaves room for 0.1% of requests to fail, equal to 43.2 minutes of full downtime in a 30-day month. That allowance is the error budget, and the team decides how to spend it.

  • The alert

    Paging on budget burn rate

    On-call is paged when the budget burns at 14.4× the sustainable rate over both 1 hour and 5 minutes, which spends 2% of the month in an hour. Slower burns (6× over 6 hours) page too; gentle ones open a ticket.

  • The safeguards

    What “on” switches on

    • Health checks that eject failed nodes
    • A router with circuit breakers and fallback providers
    • A request queue with backpressure and capped, jittered retries
    • A semantic cache and an autoscaler with a warm pool
    • Automated failover to a warm standby region

05Inference engineeringbatching · caching · quantisation · speculation

Four inference techniques that cut latency and GPU cost.

Each technique serves the same model on fewer GPU-seconds. It goes live behind the gateway only after an evaluation run shows quality has held.

  • 01

    Continuous batching

    throughput ×2–4 per GPU

    A scheduler with six GPU lanes, each decoding a request at about 75 tokens a second, 450 tokens a second between them, and a queue of three waiting requests. As a lane finishes, the next queued request joins the batch at once; a fixed batch would hold the queue for another 5.8 seconds.

    Requests join the running batch as soon as a slot frees, so none waits for the slowest request in a fixed batch.

    Quality gate evaluation set 0.94 → 0.94 · p95 −38% · pass

  • 02

    KV and prefix caching

    time to first token −53%

    Two requests shown as rows of twelve token blocks. The first request computes every block. The second request shares the same seven-block prefix, which is served from the cache, and only computes the five new blocks.

    The system prompt, tools and examples repeat on every call, so their cache is kept and only new tokens are processed.

    Quality gate evaluation set 0.94 → 0.94 · TTFT 0.62 s → 0.29 s · pass

  • 03

    Quantisation

    memory −50% to −75%

    8B model · one 24 GB GPU

    Weights
    8 GB
    4k-token seqs
    ~28
    Decode
    1.6×
    Golden set
    0.93

    FP8 selected: 8 bits per weight, 8 gigabytes of weights on a 24 gigabyte GPU, room for about 28 concurrent 4,000-token sequences, 1.6 times decode speed, golden-set score 0.93, inside tolerance, ships.

    Storing weights in fewer bits (FP8 or INT4) frees GPU memory for more requests per batch; the quality gate measures any loss.

    Quality gate −0.01 · inside the 0.02 tolerance · ships

  • 04

    Speculative decoding

    decode latency −45% to −60%

    A draft model proposes four tokens at a time. The target model checks all four in one pass, keeps the accepted ones and adds one token of its own; the rest is discarded. The output line grows by the accepted tokens: twelve tokens take three target passes instead of twelve.

    A draft model proposes several tokens and the main model verifies them in one pass, so output is unchanged but arrives sooner.

    Quality gate output unchanged by construction · decode −56% · pass

Serving engines we work with

Serving engines release new versions monthly. Each is load-tested and passes the quality gate on a canary pool before the router sends it live traffic.

  • vLLM PagedAttention · continuous batching · prefix caching
  • NVIDIA TensorRT-LLM in-flight batching · FP8 · speculative decoding
  • SGLang RadixAttention prefix cache · structured output
  • Ray Serve multi-model serving · autoscaling replicas
  • Hugging Face open-weight models · TGI

Gains are typical ranges for chat and retrieval workloads and depend on the model, prompt shape and traffic. Illustrative

06The stackeight layers, from GPUs at the bottom to security at the top

The AI stack, from GPUs to security.

Eight layers, each with one job, an owner and the tools we use. Open any layer to see what it holds and what we decide with you there.

U08 · tray open0 critical CVEs

Security & policy

Vault or the cloud KMS issues each workload short-lived credentials, so no secret lives in code or images. Images are scanned in CI, running workloads are monitored and pods that run as root or from unsigned images are refused.

Technologies we work with here

  • Vault
  • Trivy
  • Falco
  • Snyk
  • OPA Gatekeeper
  • Kyverno

Decided with you at this layer

  • Workload identity instead of static keys
  • Signed images only in production namespaces
  • Model weights, prompts and logs treated as sensitive data

U07 · tray open14 SLO alerts

Observability & FinOps

One OpenTelemetry pipeline carries traces, metrics and logs from gateway to GPU. Each trace records tokens, cache hits and cost, so a slow or costly request leads to its feature, team and model.

Technologies we work with here

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Datadog
  • PagerDuty
  • OpenCost

Decided with you at this layer

  • GenAI semantic conventions on every model call
  • Alerts on error-budget burn rate
  • Spend allocated per team and feature, daily

U06 · tray open31% cache hits

Gateway & routing

Every model call passes through it for authentication, quotas, routing by cost and data class, provider fallback, caching and token budgets per team. Its configuration is versioned code.

Technologies we work with here

  • Kong
  • Cloudflare
  • Redis
  • LiteLLM
  • Envoy AI Gateway

Decided with you at this layer

  • Which data classes may leave the country
  • Fallback order and timeouts per route
  • Cache lifetimes and similarity thresholds

U05 · tray open12.4M vectors

Data & vectors

Retrieval quality depends on infrastructure too: vector indexes sized for recall within the latency budget, embeddings versioned with their model and change streams that keep the index minutes behind the source.

Technologies we work with here

  • Qdrant
  • PostgreSQL
  • Milvus
  • Apache Kafka
  • pgvector

Decided with you at this layer

  • Index parameters tuned against recall@10
  • Re-embed only what changed
  • Tenant isolation inside the index

U04 · tray open3 models live

Model serving

Open-weight models run on vLLM or TensorRT-LLM behind an OpenAI-compatible API, so the router treats self-hosted and managed models alike. Replicas scale on queue length and time to first token.

Technologies we work with here

  • vLLM
  • NVIDIA
  • Hugging Face
  • Ray
  • SGLang
  • Triton

Decided with you at this layer

  • Model, precision and context length per route
  • Prefix caching and batch limits
  • Autoscaling on vllm:num_requests_waiting

U03 · tray open6 node pools

Kubernetes & IaC

GPU node pools are reserved for inference pods. Terraform builds the cloud, Helm packages workloads and Argo CD keeps every cluster in line with Git. Nothing in production is changed by hand.

Technologies we work with here

  • Kubernetes
  • Helm
  • Terraform
  • Argo
  • Docker
  • Karpenter
  • KEDA

Decided with you at this layer

  • GPU pools with a warm spare node
  • Scale to zero for dev and preview environments
  • Every change through a reviewed pull request

U02 · tray open3 regions · 2 in India

Cloud & regions

We build on the cloud you already use, in the regions your data rules allow. A landing zone sets identity, networking, logging and security policies once for every workload. Edge caching and a web application firewall sit in front.

Technologies we work with here

  • AWS
  • Microsoft Azure
  • Google Cloud
  • Cloudflare

Decided with you at this layer

  • Primary and standby regions
  • A mix of committed, on-demand and spot capacity
  • Private networking to model providers where offered

U01 · tray open8 × L40S · 68%

Accelerators

We choose the smallest accelerator that fits the model and its cache at your load. An 8B model at FP8 suits an L4 or L40S; 70B-class models and long contexts need an H100 or H200. Inferentia and TPU are tested on your traffic before you commit.

Technologies we work with here

  • NVIDIA
  • L4 · 24 GB
  • L40S · 48 GB
  • H100 · 80 GB
  • H200 · 141 GB
  • AWS Inferentia2
  • Google Cloud TPU

Decided with you at this layer

  • GPU class per model and context length
  • Reserved baseline, on-demand peaks, spot for batch
  • A 60–80% utilisation band on serving pools
A data-centre switch with direct-attach cables plugged into its high-speed ports
U01 · the networkGPU nodes exchange data over high-bandwidth links; a slow network leaves fast GPUs waiting.

Data residency

India regions we deploy to

Residency is set as policy at the gateway: requests tagged as personal data cannot go to a model or region outside India. Batch jobs on non-personal data may run in lower-carbon or cheaper regions.

Cloud regions in India used for data residency
CloudRegionLocationCode
AWSAsia Pacific (Mumbai)Maharashtraap-south-1
AWSAsia Pacific (Hyderabad)Telanganaap-south-2
Microsoft AzureCentral IndiaPunecentralindia
Microsoft AzureSouth IndiaChennaisouthindia
Google CloudMumbaiMaharashtraasia-south1
Google CloudDelhiDelhi NCRasia-south2

GPU types and capacity differ by region and change often; we confirm availability before a design is committed.

07FinOpscost per unit · showback · forecast · anomalies

Cost per request, visible before the invoice.

Every model call is tagged with its feature, team and route, so spend can be read per request. An agent checks the figures daily; a named person approves any change it proposes.

finops · your-platform · September · month to date

Illustrative
Spend to date
$26.3k
day 19 of 30
Forecast
$41.8k
± $2.9k at 80%
Budget
$45.0k
7% headroom
Cost per 1k requests
$1.21
−14% vs August
Unallocated
3%
of spend without a tag

Unit cost

Cost per 1k requests against the cost model

Cost per 1k requests against the cost model, with target and change against last month
NameAgainst targetActualTargetTrend
Support assistant412k requests $1.42 $1.60 −12%
Document search1.1M requests $0.38 $0.50 −4%
Sales email drafts58k requests $2.94 $2.20 +31%
Invoice extraction96k documents $0.71 $0.90 −9%
Internal copilot203k requests $1.18 $1.20 +2%

Spend forecast

$k · cumulative

Month-to-date spend of 26.3 thousand dollars on day 19, forecast to reach 41.8 thousand dollars by day 30, with an 80 percent range of 38.9 to 44.6 thousand, under a budget of 45 thousand.

GPUs in use against reserved capacity

last 7 days · 2 h buckets

Over seven days, serving used an average of 5.4 of 8 reserved GPUs, about 68 percent. 4 afternoon peaks rose above the reservation onto on-demand GPUs, and nightly batch jobs ran on 4 spot GPUs for 168 GPU-hours.

  • 5.4 avgserving on reserved · 68% of the reservation
  • 4 peaksburst to on-demand
  • 168 GPU-hbatch on spot, about 65% cheaper

Anomaly · yesterday 01:00–03:10

+$1,240 vs forecast

Embedding job re-ran 3×

embed-docs-nightly timed out listing the object store and retried the whole job instead of the failed batch, three times over.

Proposed by finops-agent · confidence 0.92

jobs/embed-docs.yaml
-  retries: 3
+  retries: 1
+  retryScope: batch
+  idempotencyKey: "{{ .source.version }}"

Awaiting your approval · the agent cannot merge

  1. finops-agentAnomaly: embedding spend 3.1× forecast (z = 4.2)
  2. finops-agentCause: 3 full re-runs after an object-store timeout
  3. finops-agentFix proposed · awaiting human approval
  • Cost reports per team

    Tags set at the gateway assign every token and GPU-hour to a feature and a team. Each team sees its own costs weekly, before finance sees the total.

  • Reserved, on-demand and spot

    Steady serving runs on reserved or committed capacity; peaks use on-demand; batch jobs, evaluation and re-embedding run on spot with checkpoints, so an interruption costs minutes.

  • Scale to zero

    Development, preview and test environments scale to zero outside working hours. The first request afterwards waits for start-up, so production is never scaled to zero.

Practices we align delivery with

  • FinOps FrameworkFinOps Foundation · inform, optimise, operate
  • FOCUSFinOps Open Cost and Usage Specification · one billing schema across clouds

08Sustainabilitysoftware carbon intensity, per unit of work

AI carbon emissions, cut by scheduling.

The same job emits more or less carbon depending on when and where it runs. Live requests stay close to users; batch work moves to lower-carbon hours and regions within your residency rules.

scheduler · your-platform · tomorrow

A chart of grid carbon intensity across one day for Mumbai and Chennai, both highest overnight and in the evening and lowest around midday, when solar output peaks; Chennai sits lower throughout. The embedding refresh is scheduled in the shaded four-hour window.

Inferencelatency-sensitive · stays near usersall day

Embedding refresh4 h · 22 kWh · drag the block or use the slider11:00–15:00

Backup snapshotfixed · keeps the RPO02:00 · 14:00

Grid intensity, I
465g/kWh
Job emissions
12.1kgCO₂e
SCI per 1k documents
30.3gCO₂e
−33% vs the 01:00 Mumbai cron

SCI = ((E × I) + M) per RE 22 kWh · I 465 g/kWh · M 1.9 kg · R 1k documents

Both regions are in India, so the documents stay in the country. Intensities are illustrative; in production, grid intensity comes from a grid-data API and energy from GPU telemetry. Illustrative

09Reliabilityservice level objectives · error budgets · on-call

Reliability targets we set with you, and an error budget we track.

Uptime is a target you choose and pay for. We agree the measures and targets with you, publish the error budget and let it decide when to release features or fix reliability.

slo · your-platform-prod · rolling 30 days

Availability

Successful responses ÷ all responses, measured at the gateway, excluding client 4xx.

Objective
99.9%
over 30 days
Observed
99.96%
this window
Budget
43.2 min
total for 30 days
Remaining
25.8 min
59.7% of budget

An error-budget burn-down for Availability across thirty days, falling from the full budget to 59.7% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom and one step drop where an incident burned budget quickly.

One provider incident on day 12 burned 6.4 minutes. The multi-provider router carried the rest of the traffic, so the budget survived it.

p95 latency

Share of non-streaming requests completed end to end within the target, per endpoint class.

Objective
99% < 1.5 s
p95 observed 1.21 s
Observed
99.38%
this window
Budget
345k req
total for 30 days
Remaining
131k req
37.9% of budget

An error-budget burn-down for p95 latency across thirty days, falling from the full budget to 37.9% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom and one step drop where an incident burned budget quickly.

A day-20 release raised p95 by 340 ms. Fast burn paged on-call inside the hour, the release was rolled back, and the budget is being repaid before the next feature ships.

Time to first token

Share of streaming responses that put their first token on the wire within the target.

Objective
95% < 800 ms
p95 observed 612 ms
Observed
96.8%
this window
Budget
965k streams
total for 30 days
Remaining
623k streams
64.6% of budget

An error-budget burn-down for Time to first token across thirty days, falling from the full budget to 64.6% remaining, plotted against a straight ideal-burn line, with deployment markers along the bottom.

Prompt caching and continuous batching hold first token steady through the evening peak, so the budget burns close to the ideal line.

Burn-rate alerting

A fast burn (14.4× the budget rate for one hour) pages on-call. A slow burn (6× over six hours) opens a ticket for the next working day. Alerts follow the service indicator, never a single failing host.

The freeze rule

With under 25% of the budget left, feature releases stop and the team spends the sprint on reliability. Everyone agrees this in advance, so nobody argues it during an incident.

Illustrative figures

The window, targets and thresholds are the ones we use; the plotted values are sample data for a sample platform. Illustrative

  • Runbooks that match the alerts

    Written during the build, each runbook gives the first five commands, the dashboards to open and the rollback.

  • A shared on-call rotation

    A primary and a secondary engineer, an escalation policy and handover notes at each shift change. Your engineers can share the rotation with ours.

  • Blameless post-incident reviews

    Within five working days: timeline, contributing causes, what monitoring missed and actions with owners, scheduled into the next sprint.

  • Tested disaster recovery

    Recovery targets (RTO, RPO) are agreed per system, then tested each quarter with a timed, documented rebuild from code and backups.

A network patch panel in close-up, its ports numbered and grey and blue cables plugged into them
02:40 · on-callEach alert links to its runbook, so on-call starts from a written plan.

Frameworks we align delivery with

Controls, continuity and hardening baselines the platform is built to. Certifying your own environment is a separate project we can prepare you for.

  • ISO/IEC 27001:2022Information security management systems
  • SOC 2Trust Services Criteria
  • CIS Controls v8.1Critical Security Controls and Benchmarks
  • NIST CSF 2.0Cybersecurity Framework
  • ISO 22301:2019Business continuity management systems

Support hours are agreed per engagement (business hours, extended hours or 24×7) and written into the runbook with the escalation path.

10How we workbaseline · redesign · operate

Measure first, then improve latency, uptime and cost..

We record latency, availability, spend and carbon at the start, so you can check every improvement against them. Nothing is rebuilt for looking old. Work starts where measurement shows the latency, spend or risk sits.

  1. 01

    Baseline

    Wk 01–02

    We map today's set-up, measure traffic, latency, incidents, spend and carbon, then agree targets and SLOs with you.

    • Baseline report
    • SLO targets
    • Cost breakdown
  2. 02

    Design

    Wk 02–04

    We design the target architecture, serving approach, capacity plan and migration path, recording each decision against the baseline.

    • Target architecture
    • Capacity model
    • Decision records
  3. 03

    Build

    Wk 04–10

    We build and load-test the infrastructure code, pipelines and serving stack. Monitoring and cost tags start with the first resource.

    • Infrastructure code
    • Load-test results
    • Dashboards
  4. 04

    Run

    Wk 10–12

    We switch traffic over and write on-call runbooks. A monthly review of SLOs, spend and capacity decides the next improvements.

    • Runbooks
    • SLO reviews
    • FinOps report

The three numbers a phase is judged on

Measured on your traffic before the first change and after each release. Each bar shows the reduction from the starting figure, so the three can be compared. One illustrative engagement shown. Illustrative

  • p95 latency, chat endpoint

    2.9 s 1.21 s

    −58%

    Continuous batching, a prompt cache and a router that sends short requests to a smaller model.

  • Cost per 1,000 requests

    $2.95 $1.21

    −59%

    Right-sized GPU pools, quantised weights and caching, with batch work moved to committed capacity.

  • Idle GPU-hours a week, non-production

    96 h 12 h

    −88%

    Non-production environments scale to zero outside working hours.

After go-live, reviews continue: a monthly SLO review of the error budget and incident actions; a FinOps review of unit cost and commitments; a capacity check before seasonal peaks.

11What you getinfrastructure as code, in your repository

Everything we build lands in a repository you own.

What we make for you is yours once it is paid for. Tools we already had stay ours, and you get a free, permanent licence to use them. The platform is code, rebuildable from an empty account.

A browsable listing of the infrastructure repository handed over at the end of the engagement, with 9 files you can open: Terraform for the GPU pool and the residency policy, Helm values for model serving and the AI gateway, a Grafana cost dashboard, the SLO objectives with their burn-rate alert rules, the GPU node-failure runbook, a k6 peak-load profile and the disaster-recovery drill script.

your-platform-infra — main

Arrow keys move, → opens a directory, ← closes it, Enter shows a file.

terraform/envs/prod/main.tf

module "gpu_pool" {  source = "../../modules/gpu-pool"   region        = "ap-south-1"    # Mumbai, primary  instance_type = "g6e.2xlarge"   # 1 × L40S, 48 GB  min_nodes     = 6  max_nodes     = 10              # 6 serving + 2 warm + 2 burst  warm_nodes    = 2  taints        = ["nvidia.com/gpu=present:NoSchedule"]  tags          = local.cost_tags # feature, team, model route}

11 linesPR #482 · platform · plan attached, 26 policy checks passed

Open any file to see its contents. Paths and values are illustrative; the structure and the arithmetic behind the thresholds are what we hand over. Illustrative

Handover pack

  • Infrastructure as code modulesTerraform · Helm
  • Reference architecture & decision recordsDiagrams · ADRs
  • Model serving stackContainers · config
  • CI/CD & GitOps pipelinesPipelines · Argo CD
  • SLOs, dashboards & alertsGrafana · alert rules
  • Load & failover test reportsk6 · report
  • Cost & carbon allocation reportDashboard · spreadsheet
  • Change goes through reviewEvery environment change is a pull request with a plan attached, a policy check and a named approver. Nothing reaches production from a laptop.
  • Your accounts, your keysThe cloud accounts, registries, secret stores and provider keys are yours from day one. We work inside them with scoped, auditable access.
  • Decisions are written downArchitecture decision records explain what was chosen, what was rejected and what would change the answer, so a future team has the context.

Built and operated with

  • Terraform
  • Helm
  • Kubernetes
  • Argo
  • GitHub Actions
  • Vault
  • Grafana
  • k6

12Outcomesreviewed monthly, against the same definitions

Five numbers, measured every month.

We agree these targets and their definitions at the start. Each monthly review shows every reading with its evidence, including any that are out of band.

monthly review · your-platform · September 4 of 5 inside the agreed band · 1 out, with an owner

  • p95 latency, chat endpoint

    Ninety-nine per cent of chat requests answered inside the agreed p95, measured per endpoint class.

    1.21 s was 1.34 s Inside the objective

    0 s agreed band 0–1.5 s 3 s

    Set per endpoint class and held under load tests at forecast peak, rather than assumed from average traffic.

  • Cost per 1,000 requests

    Within ten per cent either side of the cost model agreed at forecast volume.

    $1.21 was $1.41 On the cost model

    $0.00 agreed band $1.17–$1.43 $2.40

    Tracked per feature and per team, so a price rise or a prompt that grew shows up as a line, not a surprise invoice.

  • GPU utilisation at peak

    Between 55 and 75 per cent at the daily peak — busy enough to be economic, with headroom left.

    81% was 68% Above the band

    0% agreed band 55–75% 100%

    Below the band you are paying for idle accelerators. Above it, queues build and tail latency runs away.

    Why it is out, and what happens next. A launch in week three tripled one route’s traffic. The pool is being resized and the autoscaler’s queue-depth target lowered — owner and date are in the review log.

  • SCI per 1,000 requests

    Lower each quarter against a ceiling agreed at the start of it, measured the same way each time.

    26.4 g was 29.8 g Trending down

    0 g agreed band 0–32 g 60 g

    Carbon falls with the same levers as cost: batching, caching, quantisation, and batch work scheduled into cleaner hours.

  • Error budget spent

    No month where the availability budget is exhausted without a decision behind it.

    40.3% was 52% Inside the allowance

    0% agreed band 0–60% 100%

    A planned burn during a migration is a decision. An unplanned one is a design or a monitoring gap, and it gets a review.

  • 01

    Latency on target

    Each endpoint gets 95th and 99th percentile latency targets, held under load with autoscaling tested.

  • 02

    Reliability you can quote

    SLOs with error budgets that decide when to release features and when to fix reliability.

  • 03

    Spend each team can see

    Cost per request, user and model is visible to the team that controls it, and trends down.

Bands and readings are illustrative sample data, including the one out of band. Your targets are set against your own baseline in the first phase and reported in the same review each month. Illustrative

Services & packages

Cloud and AI infrastructure that stays fast, reliable and affordable.

Cloud, platform, AI serving and reliability services. Buy one on its own or combine several in a single brief.

Categories
04
Services
16
Packages
04
Not sure what you need? Describe the problem

How to buy

  1. 01Pick services. Enquire about one, or add several to a brief.
  2. 02Choose a package. A sprint, a fixed project or an ongoing team.
  3. 03Send the brief. We reply within one working day.

Browse by category

Timelines are typical. Every quote follows a written scope.

01Cloud architecture & migration

4 services
Typical timeline: 8–20 weeks

Cloud migration

Move applications and data from your own servers or another provider to AWS, Azure or Google Cloud, in planned waves with minimal downtime.

What’s included

  • Discovery and dependency mapping
  • Strategy per application: rehost, replatform or refactor
  • Wave plan with rollback points
  • Cutover and post-launch support
  • AWS
  • Azure
  • Google Cloud
  • AWS
  • Microsoft Azure

Best forOrganisations leaving a data centre or consolidating cloud providers.

Typical timeline: 4–8 weeks

Cloud foundations (landing zone)

A secure starting point in the cloud: accounts, networking, identity, logging and policies set up as code before workloads arrive.

What’s included

  • Account and network structure
  • Identity, SSO and least-privilege roles
  • Policy as code with CIS-aligned baselines
  • Central logging and cost tagging
  • Landing zone
  • Policy as code
  • CIS Benchmarks
  • AWS
  • Microsoft Azure

Best forTeams starting in the cloud, or tidying an estate that grew without a plan.

Typical timeline: 2–3 weeks

Cloud architecture review

A review of your current set-up for reliability, security, performance and cost, with a ranked list of improvements.

What’s included

  • Review against well-architected principles
  • Risk and cost findings
  • Target architecture
  • Prioritised roadmap
  • Well-architected
  • Roadmap
  • AWS
  • Microsoft Azure

Best forTeams preparing for growth, an audit or a funding round.

Typical timeline: 2–6 weeks

Edge & CDN delivery

Serve your site, content and APIs from locations close to your users, with caching, edge functions and a web application firewall.

What’s included

  • CDN and caching rules
  • Edge functions
  • WAF and bot protection
  • Latency monitoring by region
  • CDN
  • WAF
  • Edge

Best forSites and apps with users spread across regions.

02Kubernetes & platform engineering

3 services
Typical timeline: 6–12 weeks

Kubernetes platform

A production-ready Kubernetes set-up for running your services, with autoscaling, security policies and deployments managed through Git.

What’s included

  • Managed cluster set-up on EKS, AKS or GKE
  • GitOps delivery with Argo CD
  • Autoscaling and resource policies
  • Secrets, network policies and image scanning
  • Kubernetes
  • GitOps
  • Autoscaling

Best forTeams running many services, or outgrowing simple hosting.

Typical timeline: 8–16 weeks

Internal developer platform

Self-service templates and pipelines so engineers can create, test and deploy services the approved way without waiting on another team.

What’s included

  • Approved service templates
  • CI/CD pipelines with security checks built in
  • Preview environments
  • Developer portal and documentation
  • Service templates
  • Self-service
  • DORA metrics

Best forEngineering organisations where releases are slowed by manual steps.

Typical timeline: 3–8 weeks

CI/CD & infrastructure as code

Automate how code is tested and released, and describe your infrastructure in code so every environment can be rebuilt the same way.

What’s included

  • Build, test and deploy pipelines
  • Terraform or Pulumi modules
  • Environment promotion with approvals
  • Rollback procedures
  • CI/CD
  • IaC
  • Repeatable

Best forTeams deploying by hand or with fragile scripts.

03AI serving & MLOps

4 services
Typical timeline: 4–10 weeks

Model serving & inference optimisation

Run language models on your own infrastructure or through managed endpoints, tuned so answers come back fast and each request costs less.

What’s included

  • Serving with vLLM, Triton or managed endpoints
  • Batching, quantisation and caching
  • Routing between models by cost and quality
  • Load tests against latency targets
  • vLLM
  • Latency
  • Cost per 1k requests

Best forProducts with steady AI traffic, strict data residency or tight latency needs.

Typical timeline: 3–8 weeks

GPU infrastructure

GPU capacity sized to the workload, with scheduling and scale-to-zero, so expensive hardware is not left idle.

What’s included

  • Capacity and instance selection
  • GPU node pools and scheduling
  • Spot and committed capacity plan
  • Utilisation dashboards
  • GPU
  • Scale to zero
  • Utilisation

Best forTeams training or serving models with rising GPU bills.

Typical timeline: 6–12 weeks

MLOps pipelines

Repeatable pipelines to train, evaluate, version and deploy models and prompts, so every release can be traced and rolled back.

What’s included

  • Experiment tracking and a model registry
  • Training and embedding pipelines
  • Evaluation gates before deployment
  • Model and prompt versioning with rollback
  • Model registry
  • Lineage
  • Reproducible

Best forData science teams moving models from notebooks to production.

Typical timeline: 4–8 weeks

Private & self-hosted AI

Run open-weight models such as Llama, Mistral or DeepSeek inside your own cloud account or data centre, so sensitive data stays under your control.

What’s included

  • Model selection and licence check
  • Deployment in your VPC or on-premises
  • Access control and logging
  • Break-even analysis against API pricing
  • Open-weight
  • Data residency
  • In your VPC

Best forRegulated sectors and teams with strict data-location rules.

04Reliability, cost & sustainability

5 services
Typical timeline: 4–8 weeks

Observability & SRE

See how your systems behave in production and fix problems before customers notice, with service level objectives, tracing and alerts worth acting on.

What’s included

  • SLOs and error budgets per service
  • Traces, metrics and logs on OpenTelemetry
  • Alert tuning to cut noise
  • On-call runbooks and incident reviews
  • SLOs
  • OpenTelemetry
  • Error budgets

Best forTeams who hear about outages from customers first.

Typical timeline: 3–6 weeks, then monthly

FinOps & cloud cost reduction

Find and remove wasted cloud spend, and show each team what its services cost, without hurting performance.

What’s included

  • Cost allocation and tagging
  • Rightsizing and idle-resource clean-up
  • Savings plans, reserved and spot capacity
  • Monthly cost report per product or team
  • Rightsizing
  • Unit cost
  • Commitments
  • AWS
  • Microsoft Azure

Best forCompanies whose cloud bill grows faster than their revenue.

Typical timeline: 2–4 weeks

Performance & load testing

Find out how much traffic your system can handle, and what breaks first, before a launch or peak.

What’s included

  • Load models based on your traffic
  • Load, spike and soak tests
  • Bottleneck analysis
  • Capacity plan
  • k6
  • Capacity
  • Bottlenecks

Best forTeams expecting a launch, sale or seasonal peak.

Typical timeline: 4–8 weeks

Backup, disaster recovery & resilience

Make sure you can recover from outages, deletions or ransomware, with recovery targets set per service and tested in drills.

What’s included

  • Recovery time and point objectives per service
  • Backups and cross-region replication
  • Failover design
  • Restore and failover drills
  • RTO · RPO
  • ISO 22301-aligned
  • Drills
  • AWS
  • Microsoft Azure

Best forBusinesses where downtime or data loss would be costly.

Typical timeline: 3–6 weeks

Green & carbon-aware compute

Measure the carbon footprint of your software and reduce it alongside cost, using the Green Software Foundation’s Software Carbon Intensity (SCI) method.

What’s included

  • SCI baseline per request or workload
  • Region and scheduling choices by grid carbon intensity
  • Efficiency fixes that cut cost and emissions together
  • Reporting for sustainability teams
  • SCI
  • Carbon per request
  • GreenOps

Best forOrganisations with sustainability targets or reporting duties.

Your brief

Tick “Add to brief” on any service, choose a package, then continue. Or enquire about one service directly.

How we work with you

Ways to engage, from a question to an RFQ.

Ask a quick question, send a project brief or issue a formal RFQ. The lead for the work reads each one in full, and any services already in your brief go with it.

Or book a thirty-minute call

What are you sending?

  1. 01

    About 2 minutes4 required answers

    For a first conversation, a press request or anything that does not need a scope yet.

    You get A reply from a named lead

  2. 02Recommended

    About 8 minutes5 short steps

    Goals, audiences, a budget band and timing, so our first reply can outline the work.

    You get Options and a first scope after one call

  3. 03

    About 15 minutesYour documents attached

    Your documents, deadlines and the procurement and security rules the work must meet.

    You get Receipt confirmed and a named bid lead

How it is priced

Each package shows its pricing model. Work starts once a written scope and quote are agreed.

  • Typical length
    1–3 weeks
    Pricing
    Fixed fee
  • Typical length
    3–9 months
    Pricing
    Fixed price per milestone
  • Typical length
    Ongoing · 6-month minimum
    Pricing
    Monthly fee
  • Typical length
    6–18 months
    Pricing
    Programme fee · by statement of work
Compare what each package includes
What each package includes and who it suits
PackageEvery engagement includesBest for
SprintA short, fixed-scope engagement that answers one defined question.
  • Scope and outcome agreed before day one
  • A senior lead plus the specialists needed
  • A working review every week
  • A decision-ready answer or prototype
Discovery, a diagnostic, a prototype or a decision you need to make soon
MilestoneA larger build in phases you approve and pay for one at a time.
  • Phases with their own scope, output and sign-off
  • A go or no-go review at every gate
  • Re-planning between phases as you learn
  • Payment tied to accepted milestones
Programmes too large for one contract, where you want control at each step
RetainerReserved monthly capacity to run, improve and extend what we built.
  • A reserved block of team time every month
  • Agreed response times for requests and fixes
  • A monthly review and a rolling backlog
  • Planned improvements as well as upkeep
Live brands and products that need a steady team without hiring one
EnterpriseA multi-workstream programme with a dedicated team, governance and agreed service levels.
  • An engagement director and a steering group
  • A dedicated team across several workstreams
  • Service levels, reporting and a risk register
  • Security, legal and procurement reviews in the plan
Large organisations running change across markets, portfolios or business units

13Questionsanswered plainly

AI infrastructure, your questions answered.

Short answers on hosting, reliability, cloud choice, capacity, cost and carbon for AI in production.

  • 01 Should we self-host models or use APIs?

    APIs are fastest to start and hard to beat at low volume. Self-hosting open-weight models pays off with steady high volume, strict data residency or tight latency needs. We model the break-even on your traffic before recommending either.

  • 02 What does an SLO and error budget look like?

    An SLO is a target such as 99.9% of requests succeeding within 800 ms over 30 days. The error budget is the remaining 0.1%, about 43 minutes of full downtime in a 30-day month. When it is spent, reliability work takes priority over new features.

  • 03 Which cloud do you recommend?

    The one that fits your existing contracts, data location, skills and the AI services you need. We work across AWS, Azure and Google Cloud and design for portability where it is worth the cost.

  • 04 How do you reduce GPU costs?

    By matching capacity to demand and doing less work per request. We right-size instances, batch requests, cache prompts and responses, quantise models and send simple requests to smaller models. Where the workload allows, we scale to zero or use spot or committed capacity.

  • 05 Do you measure carbon?

    Yes, using the Green Software Foundation’s Software Carbon Intensity (SCI) method. We report emissions per request and reduce them alongside cost.

  • 06 Do you plan Kubernetes and GPU capacity?

    Yes. We design the cluster, node pools per workload type and safe GPU sharing, sized from your own traffic. The plan shows what happens at forecast peak, what the warm pool costs and where a queue absorbs a spike instead of a new node.

Tell us what you need built.

You will speak to a lead who would run the work, and get a straight answer on fit.

Book a call

Three ways to start

Every engagement starts with a written scope and a quote agreed before work begins.

Choose one of the three ways above