Ibrahim SultanAI Software Engineer
Resume

AI Software Engineer Mumbai, India

IbrahimSultan.

Finding the signal.
Building what matters.

I build AI systems end to end: agentic LLM workflows, retrieval that shows its sources, document pipelines and the infrastructure that keeps them running in production.

Explore my work
Currently

Software Development Engineer (G2)
63 Moons Technologies

01 / Selected engineering

09 projects

A few ways
to make AI useful.

Nine AI systems, from the first query to the last log line. Each chapter opens a diagram of how it works.

All work

01 / Retrieval & streaming

Qi Flash

63 Moons Technologies · 2025 to now

An answer a lawyer can use has to point back to the law it came from.

A legal research assistant is only useful if every claim leads back to the law it came from. I hardened the Python and FastAPI backend that decides where to look, gathers the evidence and attaches the citations, for legal research and for a lawyer's own uploaded documents.

~10,000
Indian acts searched; every answer cites the one it relied on.

My contribution

  • Routing: each question goes to internal Vespa legal search, a bounded web search, or a request for clarification.
  • Contextual retrieval: Reciprocal Rank Fusion across several LLM-generated queries, over chunks enriched with their document summaries.
  • Gemini fallback through LiteLLM before the first model output, with cancellation handling and sanitized streaming errors that keep partial answers.
  • Per-user token budgets admitted with renewable Redis locks, so one user cannot exhaust shared capacity.
  • Subscription checks on uploads, downloads and files carried through chat history.
  • AI web tools restricted to validated public HTTPS URLs and bounded searches; request logging that keeps private queries, responses and identifiers out of routine logs.

Python · FastAPI · Vespa · LiteLLM · Redis

A sketch of the workflow
  1. Question
  2. Evidence
  3. Cited response

A question finds sources. The response points back to them.

Next chapterQiLegal Enterprise
All work

02 / Enterprise search

QiLegal Enterprise

63 Moons Technologies

Over 900,000 judgments and acts, one search box, and a migration that lost nothing.

One search across Supreme Court and High Court judgments, central and state acts. I built the search engine behind it and moved the whole index onto production hardware.

900,000+
legal documents indexed and searchable.
~100×
faster embedding and indexing on a GPU server.

My contribution

  • Automated ingestion and vectorization of 900,000+ Supreme Court, High Court and statutory acts from AWS S3.
  • Provisioned a Vast.ai GPU server with CUDA 12 and PyTorch, making nomic embeddings about 100 times faster.
  • Migrated the full index from the GPU machine to a public Hetzner CPU deployment with S3 volume backups and no data loss.
  • Hybrid retrieval: BM25 keyword search for exact legal terms and HNSW vector search for meaning, merged with Reciprocal Rank Fusion; Vespa schema and indexing adapted for CPU query performance.
  • Enterprise connectors (S3, Dropbox, Google Drive, MinIO) with role-based access control.

Vespa · Docker · AWS S3 · CUDA 12

Next chapterDocGen & CLM
All work

03 / Document intelligence

DocGen & CLM

DocGen and CLM Auto Drafting · 63 Moons Technologies

A 300K-token brief, drafted one section at a time, with the lawyer holding the pen.

Long legal drafts, written section by section with a lawyer in control: they set the structure first, then the system writes one section at a time and carries context between them.

300K
tokens per upload, enough for long briefs and contracts with annexures.

My contribution

  • Lawyers review, edit and reorder the title and outline before any text is generated.
  • Section-by-section generation with streamed progress, ending in a final PDF with a persistent short link.
  • Rolling context: each finished section is summarised and carried forward, so later sections stay consistent within token bounds.
  • Took over the CLM Auto Drafting front end to ship it on time, and kept in-progress drafts intact across a hard refresh.
  • Built and containerized local document-partitioning and Granite embedding services with batched inference and vector-dimension validation.

FastAPI · Gemini · React · Next.js · Docker

Next chapterUnified Legacy Search
All work

05 / Product & architecture

GenAI Platform

Techpeek Magnus Solutions · 2024 to 2025

A whole GenAI product: search, chat, outcome analysis and the servers underneath.

At a small product company I led the architecture and built across the whole stack of a public AI platform for legal users: retrieval-augmented chat, case-outcome analysis, document generation with a human in the loop, data ingestion and the deployment underneath. Outcome scores come from a local classification model with structured reasoning; they are model outputs, not guaranteed outcomes.

40%
less time spent on legal research.
$20/mo
flat plan replaced about $375 an hour of usage-based AI spend.
100+
tribunal orders ingested every day, behind case-outcome dashboards.

My contribution

  • Retrieval-augmented chat on Vespa (Sentence Transformers, keyword and vector search, reranking), with streaming responses and OCR support.
  • Seven providers behind LiteLLM (Gemini, Groq, Cerebras, Hugging Face, Ollama, Claude and OpenAI) for flexible model routing.
  • NLP layers for intent detection, classification and reranking, and GPU-based asynchronous model servers for inference and embeddings.
  • Automated ingestion of tribunal orders with Python, SQL and OCR, structured into datasets for case-outcome dashboards.
  • Enterprise search with 40+ connectors (Slack, Google Drive, GitHub and more), multi-tenant workspaces, RBAC, OAuth2 and SSO, deployed with Docker Compose, Redis, NGINX and Cloudflare Tunnels.
  • Mentored developers on system design, API development and deployment practice.

FastAPI · Svelte · Vespa · LiteLLM · PostgreSQL

Next chapterGranular OCR & Document Extraction
All work

06 / Document extraction

Granular OCR & Document Extraction

63 Moons Technologies

Some legal PDFs are half typed, half scanned. The old pipeline only read half.

Re-engineered the extraction pipeline to read hybrid PDFs page by page, fixing a legacy bug where native text let scanned pages skip recognition. Only scanned pages go through OCR, in English, Hindi and Marathi.

Encrypted PDFs are supported, and one extraction interface covers PDF, DOCX, PPTX, XLSX, EML, EPUB and HTML, with temporary-file cleanup for large legal documents.

OCRmyPDF · Tesseract · MarkItDown · Python

Next chapterAgent Runtime & Model Serving
All work

07 / LLM orchestration

Agent Runtime & Model Serving

63 Moons Technologies

An agent is a loop. Making the loop trustworthy is the engineering.

Behind the research assistant I engineered the runtime itself rather than relying on framework defaults: a custom agent loop that plans, calls tools in parallel where it is safe, streams its answer and keeps the model’s reasoning out of what the user reads. Underneath, models are served in explicit tiers, so a workflow can move between hosted and self-hosted models without being rewritten.

My contribution

  • A custom agent loop for chat state, tool orchestration, retries and streamed output, using LangChain core components only where they earn their place.
  • Independent tool calls (retrieval, search, document processing) run in parallel and merge into one synthesis, cutting latency on multi-step research.
  • A streaming parser that separates reasoning markers such as <think> from the final answer, so internal traces never reach the screen.
  • Deep Research orchestration: multi-step planning, autonomous tool calling, web search and evidence retrieval before the answer is written.
  • Model tiers behind LiteLLM, and a serving path moved from Hugging Face Inference to Google’s Generative AI SDK while staying provider-agnostic.
  • Containerized local document-partitioning and Granite embedding services: batched normalized inference, bounded request queues, thread offloading and vector-dimension validation.

Python · LiteLLM · LangChain · Agents · Embeddings · Docker

Next chapterSupply Chain Security & Remediation
All work

08 / Dependency security

Supply Chain Security & Remediation

63 Moons Technologies

Credential-stealing releases of two popular packages hit our stack. I hunted them down.

When credential-stealing releases of litellm (v1.82.7/1.82.8) and axios (v1.14.1/0.30.4) appeared, I found them in our stack and removed them. I wrote the organization-wide remediation guide, hardened dependency lockfiles and moved build pipelines to secure baselines.

Incident Response · Security Audit · LiteLLM · Axios · Dependency Hardening

Next chapterIXCAN-A
All work

09 / Applied machine learning

IXCAN-A

ILM UX · 2023 to 2024

Teaching a cancer-risk model to listen to the clinicians who use it.

IXCAN-A predicts colon-cancer risk to support clinical decisions. My part was the data and the feedback loop: cleaning and preparing training data, validating behaviour against expected clinical outcomes, and turning clinicians' feedback into model refinements.

92%
accuracy reached by the model I helped refine.
2
hospitals adopted the system.

A contribution to a team's model. Not a medical guarantee.

Python · Machine learning · Data preparation · Clinical feedback

Next chapterQi Flash

02 / Expertise

03 practices

The work between
idea & reality.

Grounding the answer. Orchestrating the models. Keeping it running.

01

Ground the answer

RAG and contextual retrieval, hybrid search (BM25 with HNSW vectors, fused by Reciprocal Rank Fusion), rerankers, embeddings, Vespa and citation grounding.

In Qi Flash ↗
02

Orchestrate the models

Custom agentic loops, parallel tool calling and multi-provider routing through LiteLLM, streamed over SSE with reasoning kept out of the answer. Model fallback, token budgets and access controls, in Python, FastAPI and async services.

In the agent runtime ↗
03

Keep it running

Docker, Linux, GPU inference and embedding servers, Redis, PostgreSQL, MongoDB, AWS S3 and MinIO, pytest and regression testing. Timeouts, logs without private data and dependency hygiene.

In enterprise search ↗

Also in my toolkit
TypeScript · React · Next.js · Svelte · SQL · PyTorch · LangChain · Hugging Face · Milvus · ChromaDB · Gemini · Claude · OpenAI

03 / The person

04 stages · 2021 – now

Curiosity,
with follow-through.

Ibrahim Sultan by a lake

I’m Ibrahim Sultan Abdulaziz, an AI Software Engineer based in Mumbai.

For about three years I have built AI products end to end: the retrieval layer, the agent loop and the models behind it, the APIs and screens around them, and the servers underneath. The companies I have worked with build for law and healthcare, two fields where an answer is worthless unless you can trace it back to its source. That shaped how I build everything.

I care about what happens after the demo: when context is messy, a provider fails mid-stream, or someone needs to know exactly why the system said what it said.

At Pillai University I led the technical committee’s workshops and talks, and mentored the college theatre troupe and music band.

Download resume

The path so far

Software Development Engineer (G2)

63 Moons Technologies · Mumbai

Agentic LLM workflows, grounded retrieval, streaming APIs and local AI infrastructure for an enterprise research platform.

October 2025 – present

AI/ML Integration Engineer

Techpeek Magnus Solutions · Bangalore

Led the architecture of a GenAI product end to end: RAG over Vespa, seven LLM providers behind one router, OCR ingestion and 40+ enterprise connectors. Mentored developers on system design and deployment.

August 2024 – September 2025

AI/ML Software Engineer

ILM UX · Navi Mumbai

Healthcare machine learning: data preparation and clinician-informed refinements for IXCAN-A, a colon-cancer prediction model.

June 2023 – March 2024

B.Tech · Electronics & Computer Science

Pillai University · New Panvel

CGPA 8.42 / 10

January 2021 – June 2024

On the side

ITAT tribunal order scraper

A Python scraper that collected and structured 10,000+ Income Tax Appellate Tribunal orders into a research dataset.

What are you
working on?

Hiring an AI software engineer, or building something with LLMs that has to work in production? Email me.

Mumbai, India
Open to conversations, wherever you are.

An illustrative black hole, with a thin luminous accretion disk curving around its shadow
Light around a black hole