I am a Senior Applied Scientist in Microsoft's IDEAS Research group. My research leverages one of the largest productivity data infrastructures in the world to build better AI agents. Specifically, I aim to (i) understand how AI tools are used in real productivity workflows, and (ii) turn those large-scale, data-driven insights into more capable agents & more intuitive human–AI interactions. Some recent work:

Feed in your data, workflows, and expertise from M365 and beyond

Led research on workflow inference from M365 signals and telemetry that powers agentic skills in Copilot Frontier Tuning. This enables customers to bring the "content, processes, conventions, terminology, and workflows that collectively define how their business runs" into an RL environment for lower-barrier tuning of agents. Build 2026 Keynote Microsoft 365 Blog Microsoft AI Blog

I received my Ph.D. in Computer Science from Georgia Tech, where I developed robust, efficient, and adaptable multimodal AI models. My work has been recognized by leading media outlets, multiple doctoral fellowships and grants, and best paper awards, and has led to academic publications and over a dozen patents that have influenced industry products (TechCrunch).

Before my Ph.D., I was a researcher at Adobe Research (India), working on multimodal content generation. I completed my undergraduate studies at IIT Kanpur, and during my doctoral studies I interned at JPMorgan AI Research, Microsoft Research, and Adobe Research.

Adobe Experience Manager - personalized content snippets

Led research on multimodal content synthesis, showcased as a Sneak at Adobe Summit and covered in TechCrunch. The innovation enables customers to present content through AI-powered personalized image-text snippets. TechCrunch Adobe Research Blog Adobe Summit Sneak Video Patent Paper

Awards and Honors

Selected Recent Publications (Complete list on Google Scholar)

How Copilot Changed the Pace of Work in Word — And How We Measured It. article

Gaurav Verma, Siddharth Suri, Scott Counts.

The New Future of Work Reader, Microsoft, 2026.

Abstracting Cross-Domain Action Sequences into Interpretable Workflows. arXiv

Gaurav Verma, Scott Counts.

Preprint, 2026. Methodology behind the Copilot pace-of-work analysis.

Microsoft New Future of Work Report 2025.

Jenna Butler, Sonia Jaffe, Rebecca Janssen, Nancy Baym, Jake Hofman, Brent Hecht, Sean Rintel, Bahar Sarrafzadeh, Abigail Sellen, Mihaela Vorvoreanu, Jaime Teevan (editors) and other authors.

Microsoft Research Tech Report MSR-TR-2025-58 (https://aka.ms/nfw2025), 2025.

A Framework for Situating Innovations, Opportunities, and Challenges in Advancing Vertical Systems with Large AI Models. pdf webpage

Gaurav Verma, Jiawei Zhou, Mohit Chandra, Srijan Kumar, Munmun De Choudhury.

In Proceedings of the ACM/AAAI Conference on Artificial Intelligence, Ethics, and Society (AIES 2025).

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations. pdf NeurIPS 2024 Workshop on Adaptative Foundation Models Poster

Gaurav Verma, Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Tucker Balch, Manuela Veloso.

In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025).

Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space. pdf code webpage

Gaurav Verma, Minje Choi, Kartik Sharma, Jamelle Watson-Daniels, Sejoon Oh, Srijan Kumar.

In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024).

Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning.

Shivaen Ramshetty*, Gaurav Verma*, Srijan Kumar. pdf code

In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023).

Learning the Visualness of Text Using Large Vision-Language Models. pdf webpage

Gaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu, Ani Nenkova.

In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023).

Patents

  • Generating actionable insights from activity logs
  • Multimodal LLM web agents for autonomous execution of virtual actions App. 19/187,564
  • Discovering relationships across collections of electronic documents US 12,198,459
  • Automatically associating context-based sounds with text US 11,727,913
  • Identifying visual text using vision-language models App. 18/339,883
  • Personalizing videos with nonlinear playback US 11,670,085
  • Multimodal content fragments aligned to content criteria US 11,308,146
  • Stylistic text rewriting for a target author US 11,157,693
  • Goal-driven authoring assistance using causal stylistic prescriptions US 11,062,085
  • Digital document update using static and transient tags US 10,846,466
  • Digital document update US 10,489,498
  • Task-aware command recommendation and proactive help App. 16/454,683

Service

Area Chair
ACL ARR (Multimodality and Language Grounding)
Journal Reviewing
ACM Transactions on Computer-Human Interaction (TOCHI), IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), Journal of Medical Internet Research (JMIR), Data Mining & Knowledge Discovery
Conference Reviewing
(most iterations between 2021-2024): AAAI, ACL, EMNLP, ACL ARR, NeurIPS (Ethics reviewer), FAccT, CHI (2023: Special Recognition for Outstanding Reviews), COLM, KDD, TheWebConf, ICWSM (2021: Best Reviewer Award)