Skip to main content
All work
🏆 SIH 2023 National Winner (Ministry of Coal) · Independent solo rebuild, June 2025 → present

MiningNiti

A document-intelligence platform for coal mining: four AI agents analyse every uploaded regulation, a fifth audits compliance on demand, and a hybrid-retrieval chat answers questions with page-level citations. Retrieval quality is scored against a labelled golden set on every CI run, and the score blocks the build. Two builds, four years apart — the SIH 2023 winner was a team prototype; this is an independent, ground-up rebuild, solo since June 2025, sharing none of that code.

The backend runs on a free HuggingFace Space that sleeps when idle — a cold first request can take up to a minute.

MiningNiti UI

The Problem

Coal mining generates thousands of critical documents — MSHA regulations, equipment manuals, safety protocols, environmental impact assessments, incident investigations — scattered across PDFs and siloed systems. Finding a specific clause across 500 pages takes hours, and a missed regulation update means violations, fines, or lives.

The Solution

Four agents run on every upload: a classifier whose category feeds a safety analyzer, entity extractor and summarizer running concurrently under asyncio.gather(). A fifth audits compliance on demand, cross-referencing operational documents against regulations into a per-clause Pass / Fail / Not Addressed matrix. The safety analyzer is skipped entirely for non-safety categories, so an equipment manual never pays for a hazard screen it does not need. Questions go through a measured retrieval pipeline and every answer cites its document and page.

Architecture

Interactive Node Graph

MiningNiti Multi-Agent Pipeline.

System idle. Ready for request.
  • 01Frontend: Next.js 16 App Router (Turbopack) + React 19 across 15 routes — Clerk auth, TanStack Query, Recharts analytics, react-pdf viewer.
  • 02API: FastAPI 0.128 with 11 routers under /api/v1, Clerk JWT verification via JWKS (PyJWT) with RS256 pinned, JWKS cached and the azp claim validated, a DNS-resolving SSRF guard on URL ingestion, slowapi rate limiting at 120 req/min per IP, Pydantic v2 validation and audit logging on every mutation.
  • 03Agents: Classifier (Groq gpt-oss-120b) runs first because its category feeds the rest; Safety Analyzer (Mistral magistral-small), Entity Extractor and Summarizer (Cerebras gpt-oss-120b) run concurrently; Compliance Auditor (Groq) runs on demand.
  • 04Orchestration: no agent framework. Coordination is hand-written on asyncio.gather() with per-agent error isolation, quota-aware provider failover and a content-addressed Redis cache, so a re-upload of identical content costs zero LLM calls. Failure is loud: a section that cannot be produced is recorded on Document.processing_error and metadata.degraded_sections and shown in the UI, rather than laundered into a plausible-looking empty result.
  • 05Ingestion: pdfplumber for layout and tables with a Tesseract OCR fallback for pages yielding almost no extractable text (capped at 50 pages/doc at 200 DPI); chunks of ~1000 words with 200 overlap under a hard 4,000-character ceiling, embedded 100 at a time.
  • 06Retrieval: 23 prompt-injection guard patterns and a 1,500-character query cap → Gemini gemini-embedding-001 (768-dim) → pgvector cosine (HNSW) fused with PostgreSQL full-text ts_rank_cd over a GIN tsvector via Reciprocal Rank Fusion (k=60) → over-fetch 20 → ms-marco-MiniLM-L-6-v2 cross-encoder rerank to top 5 → generation streamed over SSE with inline [Document, Page X] citations.
  • 07Data: Supabase PostgreSQL 16 with pgvector HNSW, Supabase Storage for uploads, Upstash Redis caching completed analyses keyed by SHA-256(extracted text) + PIPELINE_VERSION — failed runs are never cached.
  • 08CI: five jobs — backend lint, a bandit security scan, pytest with coverage against real PostgreSQL and Redis services, the retrieval quality gate, and a frontend lint-and-build. Alembic migrations run before the server binds, so a failed migration stops the deploy rather than serving a half-built schema.
  • 09Evaluation & observability: retrieval and generation are scored separately, because a blended number hides which half failed. Retrieval runs in CI as a blocking gate on a local sentence-transformers model — deterministic, no API keys. Generation quality (faithfulness and relevancy, 0.70 thresholds) is Gemini-judged on demand. LangSmith traces the orchestrator, hybrid search and chat generation end to end.

Technical Trade-offs

  • Chose the weaker PDF library on purpose: PyMuPDF extracts more reliably but is AGPL-3.0, and this project ships MIT — so pdfplumber it was. It then turned out to recover tables that PyMuPDF flattens, and mining regulations are largely tabular, which made the licence-driven choice the better technical one too.
  • A 400-row table became a single 14,703-character chunk. Chunk size was configured in words and grouped by sentence, but a Markdown table contains no sentence-ending punctuation, so the whole table emitted as one chunk. gemini-embedding-001 truncates silently past ~2,048 tokens, so most of that table was never indexed and nothing anywhere raised an error. Fix: a hard MAX_CHUNK_CHARS = 4000 applied after sentence grouping.
  • Aggregate metrics could not prove the lexical arm was alive. Making the keyword search return nothing left every aggregate retrieval metric unchanged — at a 130-chunk corpus the cross-encoder fully compensated, so a green dashboard could not distinguish working hybrid search from half-dead hybrid search. The suite now carries direct guards that the lexical index returns rows and can tell 30 CFR 75.323 from 75.400.:
  • Two retry layers multiplied instead of adding. Backoff existed in both the orchestrator and the agent base class; composed, they reached up to nine attempts and minutes of sleep per agent on a single failure. Retry and provider fallback now live in BaseAgent._generate_json and nowhere else.:
  • Fallback routed by token budget, not preference: Groq's free tier allows 8K tokens/minute — the tightest constraint in the system — while Cerebras serves the identical gpt-oss-120b at 30K/minute. So agents fall back to Cerebras on a Groq rate limit: same model, same output, four times the headroom.

Impact & Results

Two builds, four years apart. The Smart India Hackathon 2023 entry — a team prototype against the Ministry of Coal problem statement — won the National Finale and was recognised by Coal India Limited & CMPDI. This is not that codebase: it is an independent, ground-up rebuild started June 2025 and developed solo since, and none of the 2023 code carried over. Retrieval is scored on every CI run against a labelled golden set of 12 queries over a 130-chunk mining corpus: Hit Rate@5 1.000 (floor 0.90), MRR 1.000 (floor 0.75), Recall@5 0.958 (floor 0.85), nDCG@5 0.968 (floor 0.75) — the gate blocks the build. 242 tests pass as blocking CI gates on every push (215 unit, 27 integration), out of 274 collected once the eval suites are counted, across 27.1K lines in two apps and 36 REST endpoints — all on pgvector rather than a managed vector database.

Known Limits

What this system does not do yet, stated plainly. All of it is tracked as next work.

  • 01Database access is synchronous inside async endpoints, so a query blocks the event loop and throughput per worker is bounded.
  • 02The background queue is an in-process asyncio.Queue: queued work does not survive a restart and does not scale across replicas.
  • 03Analysis agents see roughly 15K characters of head and tail rather than the whole document. Retrieval still indexes the full text — the limit applies to the agents only.
  • 04The lexical arm is PostgreSQL full-text search, not true BM25. Everyone claims 'hybrid BM25 + vector'; real BM25 needs an extension like pg_search. Worth being precise about rather than rounding up.
  • 05The retrieval gate runs against a 130-chunk golden corpus. At that scale the metrics are directional — they catch a regression, they are not a claim about production-scale behaviour.
  • 06No frontend test suite. All 215 unit tests are backend; the frontend is covered by lint and a blocking production build, nothing more.

Impact

5
Specialized AI Agents
1.000
Hit Rate@5 (CI-gated)
pgvector
No Managed Vector DB

Stack

Next.js 16React 19FastAPIPostgreSQL + pgvectorSupabaseUpstash RedisClerk AuthGroqCerebrasMistralGeminiDocker

Want something similar?

Let's build a production-grade system.