AI Engineers

--IS - India/Dev Labs.--

Location: Noida / Bengaluru, India (hybrid) · Experience: 2–4 years · Type: Full-time


About us

Instant Systems builds and scales technology ventures out of its development labs in India. One of them, InstantMarkets, is an open business search engine for procurement — it aggregates government and private-sector bids, RFPs, RFQs and contract awards from thousands of sources worldwide and makes them searchable in one place. Across the portfolio, our products share a common problem: enormous volumes of messy, unstructured, real-world documents that need to be understood rather than merely stored.

The role

You will build the AI features that make those products work — and you will build them for production, not for a demo.

That means retrieval systems over large document corpora that arrive as inconsistent PDFs, scans and HTML. Extraction pipelines that pull structured fields out of documents that were never designed to be parsed. Classification against real-world taxonomies. Agentic workflows where they genuinely beat a deterministic pipeline. And, underneath all of it, the evaluation harnesses that tell you whether any of it actually works.

Because you will work across several ventures rather than one, you will also be one of the people deciding what gets built once and reused everywhere. That leverage is the most interesting part of the job.

This is an engineering role. Most of your time goes into services, pipelines, data plumbing, tests and evals. The model is a component, not the job.

What you'll do

  • Take AI features from vague problem statement through to specified, built, evaluated and deployed.
  • Build RAG systems end to end — ingestion, parsing, chunking, embedding, retrieval, reranking, grounding and citation — over messy documents at scale.
  • Build extraction and classification pipelines that map unstructured documents onto structured taxonomies, and handle the long tail where documents don't cooperate.
  • Build evaluation harnesses, golden datasets and regression suites, and feed real production failures back into them.
  • Own cost and latency as first-class constraints: model routing, caching, batching, context economy, fallbacks when a provider degrades.
  • Ship with real observability — tracing, quality dashboards, spend tracking.
  • Identify what's common across ventures and build it once, properly, with documentation.

What we're looking for

  • 2–4 years of professional software engineering, including at least a year building with language models. Strong engineering fundamentals matter more here than exotic ML knowledge.
  • You've shipped an LLM-powered feature to real users in production and can describe it end to end — the data flow, the failure handling, what it cost, what broke after launch, and what you changed. Notebooks, demos and course projects don't count for this one.
  • Evaluation discipline — you can explain how you knew a system was good enough to ship, and how you measured whether changes made it better. The dataset, the metric, the baseline.
  • Strong Python and real production practice: version control, testing, CI, containers, and operating a service you wrote. Comfort with SQL and a cloud platform.
  • Practical command of the applied LLM toolkit — frontier model APIs, structured and schema-constrained output, tool calling, embeddings and vector search, reranking, context management. We care that you understand the mechanics, not that you've used a particular framework.
  • Judgment about when not to use a model. Tell us about the time you replaced an LLM call with a regex, a database query or a smaller model — and why that was right.
  • Clear written and spoken English.

Nice to have

  • Document-heavy AI — OCR, PDF parsing, layout-aware extraction, tables, scanned material.
  • Search and IR beyond naive vector similarity: hybrid search, BM25, reranking, query rewriting, and diagnosing why retrieval failed.
  • A real cost or latency reduction you drove, with the before and after numbers.
  • Fine-tuning or adapting smaller models where it beat prompting a large one.
  • Classical ML and NLP — classification, ranking, embeddings, evaluation metrics.
  • Enough frontend ability to build your own demo UI.
  • Public evidence of building: side projects, open source, technical writing.

Not required

  • A PhD or publications — we're not building novel models.
  • Experience training large models from scratch.
  • A specific degree. We care about what you've shipped.

Why join

  • Real production AI, at volume, on problems that don't have textbook answers.
  • Breadth: you'll work across several ventures rather than one narrow surface, and what you build gets reused.
  • A domain with genuinely hard document problems — procurement data is as messy as it gets.
  • Direct access to engineering leadership, short feedback loops, freedom in technical choices.
  • The field moves quarterly and you're expected to keep up on company time — evaluating what's new is part of the job, not something you do at midnight.

Instant Systems is an equal opportunity employer. We evaluate candidates on demonstrated ability, and we welcome applicants from all backgrounds.

To apply: send your CV along with a short description of one LLM system you built and shipped to real users — what it did, how you evaluated it, and what broke.

Instant Systems builds and scales technology ventures out of its development labs in India. One of them, InstantMarkets, is an open business search engine for procurement — it aggregates government and private-sector bids, RFPs, RFQs and contract awards from thousands of sources worldwide. Across the portfolio, our products share a common problem: enormous volumes of messy, unstructured, real-world documents that need to be understood rather than merely stored.

You will build the AI features that make those products work, and you will build them for production rather than for a demo. Retrieval over document corpora that arrive as inconsistent PDFs and scans. Extraction pipelines for documents never designed to be parsed. Agentic workflows where they genuinely beat a deterministic pipeline. And underneath all of it, the evaluation harnesses that tell you whether any of it actually works. This is an engineering role — the model is a component, not the job.

Production Engineering
Applied LLM & Retrieval
Evaluation & Measurement
Research Depth
Autonomy

Responsibilities

  • Take AI features from vague problem statement through to specified, built, evaluated and deployed
  • Build retrieval-augmented systems end to end — ingestion, parsing, chunking, embedding, retrieval, reranking, grounding and citation — over messy documents at scale
  • Build extraction and classification pipelines that map unstructured documents onto structured taxonomies such as NAICS, UNSPSC and CPV
  • Build evaluation harnesses, golden datasets and regression suites, and feed real production failures back into them
  • Own cost and latency as first-class constraints: model routing, caching, batching, context economy and provider fallbacks   
  • Identify what is common across our ventures and build it once, properly, with documentation

Must Have

  • 2–4 years of professional software engineering, including at least a year building with language models
  • An LLM-powered feature you shipped to real users in production and can describe end to end — data flow, failure handling, cost, what broke after launch
  • Evaluation discipline: you can explain the dataset, the metric and the baseline you used to know a system was good enough to ship
  • Strong Python, plus real production practice — version control, testing, CI, containers, and operating a service you wrote   
  • Practical command of the applied LLM toolkit: frontier model APIs, structured output, tool calling, embeddings and vector search, reranking, context management
  • Judgment about when not to use a model — you have replaced an LLM call with simpler code and can explain why
  • Clear written and spoken English

Nice to have

  • Document-heavy AI — OCR, PDF parsing, layout-aware extraction, scanned and low-quality source material
  • Search and information retrieval beyond naive vector similarity: hybrid search, BM25, ELSER, reranking, query rewriting
  • A real cost or latency reduction you drove, with the before and after numbers
  • Fine-tuning or adapting smaller models where it beat prompting a large one
  • Classical ML and NLP — classification, ranking, embeddings, evaluation metrics
  • Public evidence of building: side projects, open source, technical writing

What's great in the job?


  • Real production AI, at volume, on problems that do not have textbook answers
  • Breadth — you work across several ventures rather than one narrow surface, and what you build gets reused
  • Procurement data is as messy as document problems get, which makes the work genuinely hard and genuinely interesting
  • Direct access to engineering leadership, short feedback loops, and real freedom in technical choices
  • Keeping current is part of the job, on company time — this field moves quarterly and we expect you to move with it
  • A small, technical team that ships
Our Product

Discover our ventures and starrtups.

READ

What We Offer


Each employee has a chance to see the impact of his work. You can make a real contribution to the success of the company.
Several activities are often organized all over the year, such as weekly sports sessions, team building events, monthly drink, and much more

Perks

A full-time position
Attractive salary package.

Trainings

12 days / year, including
6 of your choice.

Sport Activity

Play any sport with colleagues,
the bill is covered.

Eat & Drink

Fruit, coffee and
snacks provided.