Technology & Intelligence Ref XE-TEC-203 Open

AI Engineer, Agents & RAG

Build agents and retrieval systems that do real work for clients, with the evals, guardrails and logs to trust them.

Illustrative opening — hiring confirms the details before it opens.

Share Email it
Two engineers at a desk discussing code on a monitor, one leaning over the other's shoulder.
The practiceTechnology & Intelligence
Practice
Technology & Intelligence
Location
New Delhi
Work pattern
Hybrid
Level
Mid–Senior
Type
Full-time
Experience
3+ years
Salary band
₹22–36 lakh a year Indicative
Openings
2
Posted
Closes
57 days left

About the role

You will build the retrieval pipelines, tools and agent workflows behind our clients' AI products — and the evidence that they work.

An agent is not done when the demo works. It is done when it passes an eval suite for grounding, accuracy and safety, when its guardrails hold against prompt injection (OWASP LLM01) and data leakage, when a human approves the actions that matter, and when its cost, latency and quality are visible in production. That is the job.

What you will do

  1. Design retrieval pipelines, tools and agent workflows
  2. Write eval suites for grounding, accuracy and safety, and run them on every change
  3. Implement guardrails for prompt injection and data leakage
  4. Put human approval where an agent takes an action that matters
  5. Instrument cost, latency and quality in production

A typical week

Where the hours go in an ordinary week, and which parts agents carry. Whatever an agent drafts or checks, a person decides what happens next.

AI Engineer, Agents & RAG 40 hours in a typical week
  • You14 h · 35%
  • You, with an agent drafting20 h · 50%
  • An agent runs it, you review6 h · 15%
  1. Retrieval pipelines and agent workflows

    A coding assistant drafts; you design, test and decide.

    You, with an agent drafting12 h

  2. Writing and running evals

    You write the cases; agents run the suites on every change and summarise what failed.

    You, with an agent drafting8 h

  3. Guardrails and red-team checks

    Prompt injection (OWASP LLM01), data leakage, and the approval steps for actions that matter.

    You6 h

  4. Production monitoring

    Dashboards and alerts watch cost, latency and quality; you investigate what they flag.

    An agent runs it, you review6 h

  5. Design reviews and client sessions

    Reviewing designs and code, and explaining results to clients.

    You8 h

An illustrative split, not a timesheet. The hiring lead walks you through the real one in your first conversation.

What you bring

  • 3+ years of software engineering, including 1+ year shipping LLM features
  • Python and one vector store in production
  • A working understanding of the OWASP Top 10 for LLM Applications
  • Evidence that you measure quality rather than eyeball it

Nice to have

  • Fine-tuning or small-model deployment
  • Experience in regulated industries

Never a reason to rule anyone out. If you have most of the list above and none of this, apply.

Tools you will use

Grouped by what they are for. Nobody knows all of them on day one.

Languages & runtimes
  • Python
Backend & APIs
  • FastAPI
Model providers
  • Anthropic
  • OpenAI
Agents, retrieval & evals
  • LangGraph
  • LlamaIndex
  • pgvector
  • Qdrant
Containers, IaC & CI/CD
  • Docker
Observability
  • OpenTelemetry

Technologies we work with and license — not partnerships.

Your first 90 days

  1. Days 1–30

    Extend the eval suite of a live agent and fix the first failure it finds

  2. Days 31–60

    Ship a retrieval improvement with before-and-after numbers from the evals

  3. Days 61–90

    Own an agent workflow in production: its guardrails, its dashboards and its approval steps

A hand-drawn goal review on dot-grid paper: tallies for weeks one to eleven beside an assess-and-plan column, two pens resting on the page.
The plan is written down before your first day, and you review it with your lead as the weeks go by.

Where you would work

The places this role can be based. Each team agrees its own studio days; remote roles meet in a studio for a team week each quarter.

  • Auto rickshaws on a wide avenue in New Delhi, the dome of Rashtrapati Bhavan at the end of the road in evening haze.

    Studio

    New Delhi

    Studio at Aerocity, ten minutes from the airport — handy for client travel.

    Pictured: New Delhi

How and where we work

What we offer

With this role

  • A monthly model-usage budget for experiments, on approved enterprise accounts

With every role

  • Health and family

    Medical insurance · Parental leave · Mental health

  • Time

    Paid leave · Winter break · Flexible days

  • Growth

    Learning budget · Growth reviews · Internal moves

  • Setup

    Your own machine · Home office · Team weeks

The full benefits schedule

The team you would join

Technology & Intelligence, the practice behind this role.

Engineers who build websites, platforms, data systems and AI products, with tests, evals and security in the pipeline on every change.

Open roles in this practice
4 on the board
Based in
Remote (India) · New Delhi · Ludhiana

Where this role sits

L3 of 6 · Mid

Owns a feature. Owns a piece of work end to end.

Each level is defined by the scope you own, not years served. Growth reviews twice a year measure you against the six standards at your level.

Ways in, and ways up
  1. L1 Intern A task
  2. L2 Associate A deliverable
  3. This role L3 Mid A feature
  4. L4 Senior A workstream
  5. L5 Lead An engagement
  6. L6 Principal A practice

The rung below

L2 · Associate — A deliverable. Ships reliably with review.

Next rung

L4 · Senior — A workstream. Sets the approach and grows others.

How we hire for this role

5 steps, and a reply after every one.

The steps for this role, set by its hiring lead. Timings are targets; when one slips, we tell you.

  1. Within 5 working days

    Step 1: Application read

    An AI engineer reads it. A link to something you built and measured helps most.

  2. Week 1 · 45 min

    Step 2: Technical conversation

    An LLM system you shipped: how you knew it worked, and what broke.

  3. Week 2 · 2 hours, paid

    Step 3: Eval exercise

    Given a small RAG pipeline and ten failing cases, find out why and fix what you can.

    Work sample · the brief
  4. Week 2 · 60 min

    Step 4: Design review

    Design the guardrails and approval steps for an agent that sends email on a client’s behalf.

  5. Week 3

    Step 5: Offer

    In writing, with the level, the band and the start date.

What helps an application

  • A link to work you made, with your own part in it named.
  • One project you could talk through for half an hour.
  • If you used AI, a line on how you checked what it produced.
  • Plain words. We read for substance, not polish.

Adjustments. Extra time, questions in advance, captions, a different format — tell us on the form or by email. Asking never counts against you.

Start the application

AI Engineer, Agents & RAGNew Delhi · Hybrid · Closes 30 Nov

Apply for this role

Your move

Not the right role yet? Tell us what you would build here.

A speculative application gets the same read as any other: a practitioner in the practice you choose, and a reply either way.