Senior AI & Data Engineer India

Vamstar
Vamstar

Software Engineering, Data Science · Full-time

India

Posted on Sep 19, 2026

Title: Senior AI & Data Engineer

Location: India only | Remote Full time Permanent

Type: Full-time

Core responsibilities & objectives

  • Design, build, and maintain batch/streaming data pipelines, ingestion, cleaning, normalisation, enrichment, deduplication.
  • Build and own ML/LLM pipelines end-to-end: document and log parsing, chunking, embeddings generation, vector indexing, agentic tool calling, multi-step workflows, retries, fallbacks, and state handling.
  • Turn raw execution logs and documents into reliable, versioned training and evaluation datasets, with leakage-aware splits and measurable quality gates.
  • Build reproducible evaluation harnesses: measure tool-selection accuracy, hallucination rates, and failure slices — not just aggregate metrics.
  • Write production-grade, well-tested Python that processes large volumes of data and documents reliably.
  • Own pipeline health: if data is stale, broken, or wrong, it's on you.
  • Work autonomously to project deadlines with minimal hand-holding.

Key qualifications & skills (non-negotiable)

  • 7+ years in backend data-heavy development or data engineering
  • Previously worked in Startup
  • Highly proficient in Python
  • Hands-on experience with large datasets and high-velocity data streams (Kafka, Flink, Spark).
  • Strong with pipeline orchestration tools (Airflow, MLflow, or equivalent).
  • Solid SQL skills (Postgres, BigQuery, or Snowflake) and NoSQL experience (DynamoDB, OpenSearch, Elastic).
  • Real experience with LLM workflows: RAG architectures, embeddings/vector DBs, prompt engineering, function/tool calling, observability.
  • Hands-on LLM dataset preparation and evaluation: deduplication, sampling, train/validation splitting without task or session leakage, and tracing bad metrics back to inputs and labels.
  • Applied experience with supervised fine-tuning or model benchmarking, and an understanding of quality-vs-cost trade-offs.
  • Deep understanding of ETL/ELT patterns and data processing at scale.

Preferred background (strong signals)

  • Experience with AWS data stack at scale (S3, Glue, EC2/GPU instances).
  • Exposure to enterprise or regulated environments where data governance matters.
  • Built and shipped data, ML and LLM-powered pipelines in production.
  • Experience with knowledge distillation, parameter-efficient fine-tuning (LoRA/QLoRA), or long-context evaluation.
  • Familiarity with model serving concerns (vLLM, quantisation, latency/cost trade-offs), even if you haven't owned inference infrastructure.
  • Has debugged a pipeline and knows why observability matters.
  • Worked in a fast-moving startup where "that's not my job" doesn't exist.

What will get you rejected

  • "I set up the pipeline, someone else monitors it" mindset.
  • Tutorials and side projects but no production experience at scale.
  • Prompt-only AI experience — you've never built or evaluated the data and training pipeline behind the model.
  • Can't explain trade-offs between streaming vs. batch, why you chose one vector DB over another, or how you split training data to avoid leakage.
  • Needs detailed specs before writing a line of code.
  • No curiosity about what the data actually means or how the model uses it.

Interested? We're a distributed team solving hard problems in enterprise AI — turning messy logs and documents into models that reliably automate complex tool-use workflows. If you want ownership, not just tickets, we'd like to hear from you.