Gunjan Dhanuka

AI Research Scientist at Mistral AI · Multimodal models

gd-profile.jpeg
San Francisco Bay Area · ✉️ gdhanuka192 [at] gmail.com

Hello, I’m Gunjan.

I’m an AI Research Scientist at Mistral AI, working on multimodal models. My current work spans OCR and document understanding, as well as computer-use capabilities. I was a core member of the team behind Mistral OCR 4 and Mistral OCR 4.1, state-of-the-art OCR models for document intelligence. I care about making models genuinely helpful with everyday tasks: understanding documents, navigating interfaces, and turning human intent into useful actions.

I completed an MS in Computer Science at Carnegie Mellon University and a B.Tech. in Computer Science and Engineering at IIT Guwahati. At CMU, I worked on compression-based methods for LLMs and memorization in diffusion models with Zico Kolter and Pratyush Maini, and on model debiasing with Fernando de la Torre and Shayok Chakraborty. Before that, I interned as an ML Engineer at Apple and as a Software Engineer at Rubrik.

I enjoy turning fuzzy problems into concrete experiments, building the infrastructure to test them, and writing down what I learn along the way. My work on weakly supervised video anomaly detection was accepted to WACV 2025; recent work on compression-based memorization in diffusion models was accepted to the ICML MemFM Workshop 2025.

Outside of work, I’m usually taking photos, planning a trip, playing a game, or watching an improbably specific engineering video on YouTube. I’m also starting to collect trail notes and trip reports.

A field-notes illustration of mountains, a trail, code, gaming, and football

Beyond the screen

The things that keep my curiosity well-fed.

Some problems are best solved at a desk. The rest tend to get better after a long hike, a close football match, a beautifully designed game, or an unnecessarily deep coding rabbit hole.

mountain hours trail maps football nights game worlds camera roll
Follow the trail notes

What I care about

Useful intelligence needs more than a good benchmark.

01

Document understanding

Building multimodal models that can make sense of documents beyond plain text: their structure, context, and visual information.

02

Computer use

Helping models understand interfaces and take useful steps through the tools people already use every day.

03

Everyday usefulness

Making capable models feel practical and intuitive—useful for real work, not just impressive on a benchmark.

Start here: browse research for publications, learning notes for an evolving technical garden, writing for long-form notes, travels for trail notes, or the now page for a concise snapshot.

latest posts

selected publications

  1. Shieldstral
    Shieldstral
    Mistral AI, Gunjan Dhanuka, and  others
    Technical report · arXiv preprint , 2026
  2. Voxtral TTS
    Voxtral TTS
    Mistral AI, Gunjan Dhanuka, and  others
    Technical report · arXiv preprint , 2026
  3. Voxtral Realtime
    Voxtral Realtime
    Mistral AI, Gunjan Dhanuka, and  others
    Technical report · arXiv preprint , 2026
  4. Dynamic Mitigation of Hardware Trojan Induced Black Hole Router Attack in Network-on-Chip
    Dynamic Mitigation of Hardware Trojan Induced Black Hole Router Attack in Network-on-Chip
    Gunjan Dhanuka, Syam Sankar, and John Jose
    In 2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI), 2025
  5. Distilling Aggregated Knowledge for Weakly-Supervised Video Anomaly Detection
    Distilling Aggregated Knowledge for Weakly-Supervised Video Anomaly Detection
    Jash Dalvi, Ali Dabouei, Gunjan Dhanuka, and 1 more author
    In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025