Ground the answer
RAG and contextual retrieval, hybrid search (BM25 with HNSW vectors, fused by Reciprocal Rank Fusion), rerankers, embeddings, Vespa and citation grounding.
In Qi Flash ↗AI Software Engineer Mumbai, India
Finding the signal.
Building what matters.
I build AI systems end to end: agentic LLM workflows, retrieval that shows its sources, document pipelines and the infrastructure that keeps them running in production.
Explore my workSoftware Development Engineer (G2)
63 Moons Technologies
Nine AI systems, from the first query to the last log line. Each chapter opens a diagram of how it works.
An answer a lawyer can use has to point back to the law it came from.
A legal research assistant is only useful if every claim leads back to the law it came from. I hardened the Python and FastAPI backend that decides where to look, gathers the evidence and attaches the citations, for legal research and for a lawyer's own uploaded documents.
Python · FastAPI · Vespa · LiteLLM · Redis
A question finds sources. The response points back to them.
Over 900,000 judgments and acts, one search box, and a migration that lost nothing.
One search across Supreme Court and High Court judgments, central and state acts. I built the search engine behind it and moved the whole index onto production hardware.
Vespa · Docker · AWS S3 · CUDA 12
Next chapterDocGen & CLMA 300K-token brief, drafted one section at a time, with the lawyer holding the pen.
Long legal drafts, written section by section with a lawyer in control: they set the structure first, then the system writes one section at a time and carries context between them.
FastAPI · Gemini · React · Next.js · Docker
Next chapterUnified Legacy SearchFour collections, four shapes of data, one box to search them all.
Reworked search across Supreme Court judgments, High Court judgments, central acts and state acts: four collections with different shapes, searched from one box.
Separate MongoDB aggregation paths per court and act scope, compound indexes, timeouts on every query, and tolerant matching for judge names and case titles (v, vs, v.). Dates and legal entities are normalized, and the React interface gained mixed-source pagination, filters, highlights and search history.
MongoDB · Python · TypeScript · React
Next chapterGenAI PlatformA whole GenAI product: search, chat, outcome analysis and the servers underneath.
At a small product company I led the architecture and built across the whole stack of a public AI platform for legal users: retrieval-augmented chat, case-outcome analysis, document generation with a human in the loop, data ingestion and the deployment underneath. Outcome scores come from a local classification model with structured reasoning; they are model outputs, not guaranteed outcomes.
FastAPI · Svelte · Vespa · LiteLLM · PostgreSQL
Next chapterGranular OCR & Document ExtractionSome legal PDFs are half typed, half scanned. The old pipeline only read half.
Re-engineered the extraction pipeline to read hybrid PDFs page by page, fixing a legacy bug where native text let scanned pages skip recognition. Only scanned pages go through OCR, in English, Hindi and Marathi.
Encrypted PDFs are supported, and one extraction interface covers PDF, DOCX, PPTX, XLSX, EML, EPUB and HTML, with temporary-file cleanup for large legal documents.
OCRmyPDF · Tesseract · MarkItDown · Python
Next chapterAgent Runtime & Model ServingAn agent is a loop. Making the loop trustworthy is the engineering.
Behind the research assistant I engineered the runtime itself rather than relying on framework defaults: a custom agent loop that plans, calls tools in parallel where it is safe, streams its answer and keeps the model’s reasoning out of what the user reads. Underneath, models are served in explicit tiers, so a workflow can move between hosted and self-hosted models without being rewritten.
Python · LiteLLM · LangChain · Agents · Embeddings · Docker
Next chapterSupply Chain Security & RemediationCredential-stealing releases of two popular packages hit our stack. I hunted them down.
When credential-stealing releases of litellm (v1.82.7/1.82.8) and axios (v1.14.1/0.30.4) appeared, I found them in our stack and removed them. I wrote the organization-wide remediation guide, hardened dependency lockfiles and moved build pipelines to secure baselines.
Incident Response · Security Audit · LiteLLM · Axios · Dependency Hardening
Next chapterIXCAN-ATeaching a cancer-risk model to listen to the clinicians who use it.
IXCAN-A predicts colon-cancer risk to support clinical decisions. My part was the data and the feedback loop: cleaning and preparing training data, validating behaviour against expected clinical outcomes, and turning clinicians' feedback into model refinements.
A contribution to a team's model. Not a medical guarantee.
Python · Machine learning · Data preparation · Clinical feedback
Next chapterQi FlashGrounding the answer. Orchestrating the models. Keeping it running.
RAG and contextual retrieval, hybrid search (BM25 with HNSW vectors, fused by Reciprocal Rank Fusion), rerankers, embeddings, Vespa and citation grounding.
In Qi Flash ↗Custom agentic loops, parallel tool calling and multi-provider routing through LiteLLM, streamed over SSE with reasoning kept out of the answer. Model fallback, token budgets and access controls, in Python, FastAPI and async services.
In the agent runtime ↗Docker, Linux, GPU inference and embedding servers, Redis, PostgreSQL, MongoDB, AWS S3 and MinIO, pytest and regression testing. Timeouts, logs without private data and dependency hygiene.
In enterprise search ↗Also in my toolkit
TypeScript · React · Next.js · Svelte · SQL · PyTorch · LangChain · Hugging Face · Milvus · ChromaDB · Gemini · Claude · OpenAI

I’m Ibrahim Sultan Abdulaziz, an AI Software Engineer based in Mumbai.
For about three years I have built AI products end to end: the retrieval layer, the agent loop and the models behind it, the APIs and screens around them, and the servers underneath. The companies I have worked with build for law and healthcare, two fields where an answer is worthless unless you can trace it back to its source. That shaped how I build everything.
I care about what happens after the demo: when context is messy, a provider fails mid-stream, or someone needs to know exactly why the system said what it said.
At Pillai University I led the technical committee’s workshops and talks, and mentored the college theatre troupe and music band.
Download resume63 Moons Technologies · Mumbai
Agentic LLM workflows, grounded retrieval, streaming APIs and local AI infrastructure for an enterprise research platform.
October 2025 – presentTechpeek Magnus Solutions · Bangalore
Led the architecture of a GenAI product end to end: RAG over Vespa, seven LLM providers behind one router, OCR ingestion and 40+ enterprise connectors. Mentored developers on system design and deployment.
August 2024 – September 2025ILM UX · Navi Mumbai
Healthcare machine learning: data preparation and clinician-informed refinements for IXCAN-A, a colon-cancer prediction model.
June 2023 – March 2024Pillai University · New Panvel
CGPA 8.42 / 10
January 2021 – June 2024A Python scraper that collected and structured 10,000+ Income Tax Appellate Tribunal orders into a research dataset.
Hiring an AI software engineer, or building something with LLMs that has to work in production? Email me.
ibrahimsultan4705Answered by an open model (Llama 4 Scout on Cloudflare Workers AI) from this site’s own text: hybrid retrieval, then an answer that cites it. It can be wrong; the chapters are the record. Questions are not stored.
Mumbai, India
Open to conversations, wherever you are.
