Now
A short snapshot of the ideas and experiences that shape my work.
This is a lightweight snapshot rather than a second résumé.
Currently
I’m an AI Research Scientist at Mistral AI in the San Francisco Bay Area, working on multimodal models. My current focus is OCR and document understanding, along with computer-use capabilities. I was a core member of the team behind Mistral OCR 4 and Mistral OCR 4.1.
The thread running through my work
I keep returning to one question: how can multimodal models understand the things people work with every day—documents, interfaces, and instructions—and turn that understanding into useful action? I’m especially interested in the step from recognizing what is on a page or screen to helping someone get something done.
Background
- MSCS at Carnegie Mellon University (2025) and B.Tech. in CSE at IIT Guwahati (2024), with a minor in Robotics and Artificial Intelligence.
- CMU research on LLM compression, diffusion-model memorization, and model debiasing.
- A machine-learning engineering internship at Apple, focused on multimodal LLM agents for automated data analysis.
- Earlier systems engineering at Rubrik and research spanning video anomaly detection, ML efficiency, and model behavior.
Away from the keyboard
Photography, travel, games, and the occasional rabbit hole about how something was designed or built.
If you want the longer version, start with my writing or research. For the formal version, see my CV.