Skip to main content
AVAILABLE FOR 2026 ROLES

MILAN SONI

[01]AI ENGINEER · FULL STACK · RAG SPECIALIST

Engineering intelligent systems, multi-agent AI pipelines, and production RAG architectures.

Scroll

[03]AI Systems Lab

How I actually build these systems.

Not a list of tools — the pipeline itself. It walks itself stage by stage; select any node to take over and see what runs there, why it was built that way, and what it is measured on.

User query
Final response
Paused·Stage 01 / 07

Stage 01 · Query Understanding

Bound the input before it reaches anything expensive

The cheapest place to stop a bad request is before it touches a model. Everything here runs ahead of retrieval.

  • 23 prompt-injection guard patterns — instruction override, role-play, and system-prompt extraction attempts are rejected up front.
  • A hard 1,500-character query cap. Anything longer returns a 422 before it reaches retrieval, rather than quietly costing tokens.
  • Rate limiting at 120 requests/minute per IP via slowapi, applied as a default limit rather than per-route opt-in.
FastAPIPydantic v2slowapi
MiningNiti

[04]Evidence

Credentials that can be checked.

Three claims, each with a primary source behind it — an official result, a peer-reviewed paper, and merged code in someone else's repository.

SIH 2023
National Winner
Ministry of Coal · recognised by Coal India Limited & CMPDI

Built MiningNiti — AI document intelligence with 5 specialised agents, hybrid retrieval scored by a CI-gated eval, and per-clause compliance auditing.

Scopus indexed
PiCET-2026
IET Conference Proceedings · paper ID PU/PiCET26/COP/327

Co-authored research on the Hybrid Attention Temporal Framework for early dropout prediction — then built the half a paper leaves out: explanations, calibrated uncertainty, and a published fairness audit.

Open source
OmniRoute · 50k★
Universal AI gateway · 230+ LLM providers · 21,000+ tests

5+ PRs merged and shipped in v3.8.44 / v3.8.45, plus a provider-flag schema design adopted by the maintainer into their own fix.

[05]Experience

Shipping inside real teams.

Four teams, one open-source project, and a national hackathon I helped run — with the numbers each one moved.

Jul 2026

OmniRoute

Open Source Contributor

Remote

  • Diagnosed an HTTP 400 regression in the memory-injection pipeline affecting strict LLM providers (Xiaomi MiMo); proposed a declarative systemMessageMustBeFirst Zod schema flag adopted by the maintainer into the shipped fix — 25/25 Vitest + 30/30 Node test-runner coverage (PR #6225)
  • Built an accessible 'Configured Only' filter toggle for the provider-rankings dashboard, mapping live connection state to a filterable data grid — 168 additions across 4 files, 9/9 tests passing, shipped into v3.8.45 (PR #6245)
  • Integrated Claude 5 Sonnet into the claude_web provider registry with regression test coverage, shipping a verified signed commit within hours of the model's release (PR #6209, v3.8.45)
  • Audited and normalized 9 core docs + 20+ localized READMEs across 42 locales, removing untranslated Portuguese/Chinese prose and correcting stale architecture metrics (routing strategies 13→17, service modules 36→134) — passed docs-sync-strict CI gate with zero regressions (PR #6105, v3.8.44)
  • Proposed a schema extension (Issue #6241) to standardize effort and thinking parameters across the multi-provider routing platform for next-gen reasoning models

Oct – Dec 2025

nTheta Works Pvt. Ltd.

Full Stack Developer Intern

Remote

  • Built a two-stage semantic retrieval pipeline (Ollama embeddings → Redis HNSW → FlashRank cross-encoder reranking) for NLPForge, an enterprise LLM API testing platform — improved template matching accuracy by 40% and reduced manual QA effort by 60%
  • Shipped async FastAPI microservices + Next.js/TypeScript dashboards with Docker Compose orchestration and CI/CD on Linux servers

Jul 2025 – Aug 2025

Freelance Client

AI & Full Stack Developer

Remote

  • Built SmartLearnX, an AI-powered LMS, with a dropout prediction model (Logistic Regression, 91.4% accuracy) and academic performance forecasting (Random Forest, R² = 0.89) deployed as a FastAPI microservice alongside a React/Node.js frontend
  • Integrated NLP features (BERT for quiz generation, spaCy chatbot) serving 24/7 student support with sub-2-second response times under load

May – Jul 2025

OBG Outsourcing Pvt. Ltd.

Full Stack Developer Intern

Jaipur

  • Led FinSageAI360, a multi-tenant financial intelligence SaaS — cut monthly close reporting time by 45% and reduced manual operational effort by 30% through AI-driven anomaly detection and real-time KPI dashboards
  • Designed a JWT-authenticated REST API (Node.js/Express/MongoDB) with granular RBAC for multi-tenant data isolation

Jun – Aug 2024

Om Logistics Ltd.

Software Developer Intern

Delhi

  • Optimized enterprise document search by implementing LangChain + FAISS vector embeddings — reduced query latency by 70% across 10,000+ documents and improved retrieval accuracy by 40%
  • Built RESTful APIs (Node.js) to automate logistics workflows, eliminating 20% of manual data-entry tasks

2024 – 2025

CodeFiesta

Organising Team & Sponsorship Head

GIT Jaipur

Official event report
  • Sponsors: 18+ — outreach, partnership negotiation and sponsor relationships across editions 3.0 and 4.0
  • 2025 edition: 200 teams from 30+ colleges, 24-hour onsite build, 13 in the final round
  • Handled logistics, scaling and technical operations across two editions of the event

Foundation

B.Tech, Computer Science & Engineering

Global Institute of Technology, Jaipur · Oct 2022 – May 2026

8.10Cumulative GPA
2026Graduating
Milan Soni
Milan Soni · Churu (Rajasthan)

·Behind the work

Most of the value isn't in the model — it's in the system around it.

Retrieval, evaluation, guardrails, UI, latency, auth, deploys. I build all of it — in production, at three companies.

Open to 2026 roles

Let's talk.

Looking for someone to ship product end to end — from applied GenAI features to the full-stack platform around them? I'm interested in AI engineering and SDE roles.