ML Engineer · Technical Writer
I specialize in production ML systems — RAG pipelines, inference optimization, memory management, and the gap between a model that works in a notebook and one that survives real load. I write about it too, with concrete numbers rather than hand-waving.
What I do
My work sits at the intersection of machine learning, clean Python engineering, and technical communication — building systems that hold up in production, and writing that makes them understandable.
01
Inference optimization, training pipelines, and production deployment. Specializing in the part most projects underestimate: making models fast, memory-stable, and reliable under real load.
02
Production retrieval-augmented generation — chunking and embedding strategies, vector search, retrieval evaluation, and grounding LLM outputs so they stay accurate at scale.
03
Clean, well-tested Python — data pipelines, automation scripts, CLI tools, and backend services built for maintainability.
04
Documentation, tutorials, and blog posts that make complex systems understandable — without losing accuracy.
05
Architecture and design of ML systems and data-intensive applications — with an eye for catching problems early.
Technical skills
Writing
I write about production ML, inference optimization, and Python internals — on blog.rakeshraushan.org and Medium.
Let's connect
Always happy to talk shop — whether it's an interesting open-source idea, a question about something I've written, or a problem in production ML you want to think through together. Drop me an email; I typically respond within a day.