Harsh Trivedi

Harsh Trivedi

I am a Research Scientist at Allen Institute for Artificial Intelligence (Ai2). These days, I work on building and evaluating AI agents for complex tasks involving tool use, coding and UI interactions. If you're curious about this line of work, check out AppWorld (our recent project), and feel free to reach out if you're interested in collaborating!

Previously, I completed my PhD at Stony Brook University, where I was advised by Niranjan Balasubramanian. I worked on two core questions: (i) how to evaluate models so that we know that they are employing reliable multi-step reasoning (e.g., taking all the steps of reasoning and not taking shortcuts or not hallucinating), and (ii) how to teach models the same. In that theme, I built benchmarks (AppWorld, MuSiQue) and evaluation (DiRe), training (TeaBReaC, Multee), and prompting (IRCoT, DecomP) methods.

Publications

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

ICML 2026 Conference

Junzhi Chen*, Harsh Trivedi*, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal

Paper Website Code Copy Bib

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

ICLR 2026 Conference

Mike A. Merrill, Alexander G. Shaw, et al. (including Harsh Trivedi)

Paper Website Code Copy Bib

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

ICLR 2026 Conference

Sayash Kapoor, Benedikt Stroebl, et al. (including Harsh Trivedi)

Paper Website Copy Bib

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

ECCV 2026 Conference

Tanmay Gupta, Piper Wolters, Zixian Ma, Peter Sushko, et. al (including Harsh Trivedi)

Paper ECCV Blog Code Copy Bib

MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents

TACL 2026 Journal

Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty

Paper Website Copy Bib

Olmo 3

arXiv 2025 Tech report

Team Olmo (including Harsh Trivedi)

Paper Website Code Copy Bib

AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

ACL 2024 Conference 🏆 Best Resource Paper Award

Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, Niranjan Balasubramanian

Paper Website Tweet Talks Blog Poster Code Leaderboard Copy Bib

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

ACL 2023 Conference

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

Paper Tweet Video Code Copy Bib

Decomposed Prompting: A Modular Approach for Solving Complex Tasks

ICLR 2023 Conference

Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, Ashish Sabharwal

Paper Code Copy Bib

Two-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions

ML Safety @ NeurIPS 2022 Workshop 🏆 Best Paper Award

Alicia Parrish*, Harsh Trivedi*, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Amanpreet Singh Saimbhi, Samuel R. Bowman

Paper Copy Bib

Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard Contexts

EMNLP 2022 Conference

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

Paper Video Code Copy Bib

MuSiQue: Multihop Questions via Single-hop Question Composition

TACL, presented at NAACL 2022 Journal

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

Paper Video Code Leaderboard Copy Bib

Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions

Learning with Natural Language Supervision @ ACL 2022 Workshop

Alicia Parrish*, Harsh Trivedi*, Ethan Perez, Angelica Chen, Nikita Nangia, Jason Phang, Samuel R. Bowman

Paper Video Data Copy Bib

Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading Comprehension

EMNLP 2021 Conference

Naoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian, Kentaro Inui

Paper Video Code Copy Bib

IrEne-viz: Visualizing Energy Consumption of Transformer Models

EMNLP 2021 Demo

Yash Kumar Lal, Reetu Singh, Harsh Trivedi, Qingqing Cao, Aruna Balasubramanian, Niranjan Balasubramanian

Paper Code Copy Bib

What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?

ACL 2021 Conference

Nikita Nangia*, Saku Sugawara*, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman

Paper Video Code Copy Bib

IrEne: Interpretable Energy Prediction for Transformers

ACL 2021 Conference

Qingqing Cao, Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian

Paper Video Code Copy Bib

Is Multihop QA in DiRe Condition? Measuring and Reducing Disconnected Reasoning

EMNLP 2020 Conference

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

Paper Code Video Slides Copy Bib

DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering

ACL 2020 Conference

Qingqing Cao, Harsh Trivedi, Aruna Balasubramanian, Niranjan Balasubramanian

Paper Video Slides Code Copy Bib

Repurposing Entailment for Multi-Hop Question Answering Tasks

NAACL 2019 Conference

Harsh Trivedi, Heeyoung Kwon, Tushar Khot, Ashish Sabharwal, Niranjan Balasubramanian

Paper Code Slides Copy Bib

Controlling Information Aggregation for Complex Question Answering

ECIR 2018 Conference

Heeyoung Kwon, Harsh Trivedi, Peter Jansen, Mihai Surdeanu, Niranjan Balasubramanian

Paper Poster Copy Bib

* denotes equal contribution