Senior AI & Data Engineer India
Software Engineering, Data Science · Full-time
India
Posted on Sep 19, 2026
Title: Senior AI & Data Engineer
Location: India only | Remote Full time Permanent
Type: Full-time
Core responsibilities & objectives
- Design, build, and maintain batch/streaming data pipelines, ingestion, cleaning, normalisation, enrichment, deduplication.
- Build and own ML/LLM pipelines end-to-end: document and log parsing, chunking, embeddings generation, vector indexing, agentic tool calling, multi-step workflows, retries, fallbacks, and state handling.
- Turn raw execution logs and documents into reliable, versioned training and evaluation datasets, with leakage-aware splits and measurable quality gates.
- Build reproducible evaluation harnesses: measure tool-selection accuracy, hallucination rates, and failure slices — not just aggregate metrics.
- Write production-grade, well-tested Python that processes large volumes of data and documents reliably.
- Own pipeline health: if data is stale, broken, or wrong, it's on you.
- Work autonomously to project deadlines with minimal hand-holding.
Key qualifications & skills (non-negotiable)
- 7+ years in backend data-heavy development or data engineering
- Previously worked in Startup
- Highly proficient in Python
- Hands-on experience with large datasets and high-velocity data streams (Kafka, Flink, Spark).
- Strong with pipeline orchestration tools (Airflow, MLflow, or equivalent).
- Solid SQL skills (Postgres, BigQuery, or Snowflake) and NoSQL experience (DynamoDB, OpenSearch, Elastic).
- Real experience with LLM workflows: RAG architectures, embeddings/vector DBs, prompt engineering, function/tool calling, observability.
- Hands-on LLM dataset preparation and evaluation: deduplication, sampling, train/validation splitting without task or session leakage, and tracing bad metrics back to inputs and labels.
- Applied experience with supervised fine-tuning or model benchmarking, and an understanding of quality-vs-cost trade-offs.
- Deep understanding of ETL/ELT patterns and data processing at scale.
Preferred background (strong signals)
- Experience with AWS data stack at scale (S3, Glue, EC2/GPU instances).
- Exposure to enterprise or regulated environments where data governance matters.
- Built and shipped data, ML and LLM-powered pipelines in production.
- Experience with knowledge distillation, parameter-efficient fine-tuning (LoRA/QLoRA), or long-context evaluation.
- Familiarity with model serving concerns (vLLM, quantisation, latency/cost trade-offs), even if you haven't owned inference infrastructure.
- Has debugged a pipeline and knows why observability matters.
- Worked in a fast-moving startup where "that's not my job" doesn't exist.
What will get you rejected
- "I set up the pipeline, someone else monitors it" mindset.
- Tutorials and side projects but no production experience at scale.
- Prompt-only AI experience — you've never built or evaluated the data and training pipeline behind the model.
- Can't explain trade-offs between streaming vs. batch, why you chose one vector DB over another, or how you split training data to avoid leakage.
- Needs detailed specs before writing a line of code.
- No curiosity about what the data actually means or how the model uses it.
Interested? We're a distributed team solving hard problems in enterprise AI — turning messy logs and documents into models that reliably automate complex tool-use workflows. If you want ownership, not just tickets, we'd like to hear from you.