Rakesh Raushan

Rakesh Raushan

ML Engineer  ·  Technical Writer

I specialize in production ML systems — RAG pipelines, inference optimization, memory management, and the gap between a model that works in a notebook and one that survives real load. I write about it too, with concrete numbers rather than hand-waving.

Areas of expertise

My work sits at the intersection of machine learning, clean Python engineering, and technical communication — building systems that hold up in production, and writing that makes them understandable.

01

ML Engineering

Inference optimization, training pipelines, and production deployment. Specializing in the part most projects underestimate: making models fast, memory-stable, and reliable under real load.

02

RAG Systems

Production retrieval-augmented generation — chunking and embedding strategies, vector search, retrieval evaluation, and grounding LLM outputs so they stay accurate at scale.

03

Python Development

Clean, well-tested Python — data pipelines, automation scripts, CLI tools, and backend services built for maintainability.

04

Technical Writing

Documentation, tutorials, and blog posts that make complex systems understandable — without losing accuracy.

05

System Design

Architecture and design of ML systems and data-intensive applications — with an eye for catching problems early.

Stack

Machine Learning
PyTorch ONNX Runtime YOLOv8 HuggingFace LLM Inference RAG Vector Search Quantization
Python & Data
Python FastAPI NumPy Pandas OpenCV Pydantic
Infrastructure
Docker AKS GitHub Actions Linux PostgreSQL

From the blog & Medium

I write about production ML, inference optimization, and Python internals — on blog.rakeshraushan.org and Medium.

Get in touch

Always happy to talk shop — whether it's an interesting open-source idea, a question about something I've written, or a problem in production ML you want to think through together. Drop me an email; I typically respond within a day.