Publications

Two questions run through most of my work: how much structure a language model needs in order to act reliably, and how much structure it can recover from unstructured text. The first is what I work on now, on agents and their security; the second is what led me there, through information extraction on scientific writing.

Research

Agents that use tools and catch their own mistakes

RIMRULE (ACL 2026) learns rules for tool-using agents under an MDL objective, so that what the agent learns stays compact enough to be inspected. Metacognitive Self-Correction for Multi-Agent Systems (ACL 2026) reconstructs what a multi-agent system was about to do next, and uses the mismatch as a signal for catching its own errors mid-execution.

AI agent security

This is my current focus at AWS: tool-use security, robustness, and anomaly detection for agents that take real actions in real systems. An agent that can call tools is an agent that can be made to call the wrong one — the reliability question and the security question turn out to be the same question.

Structured knowledge from text

Most of my Ph.D. work is about pulling entities, relations, and mentions out of scientific writing — and about making that extraction hold up when supervision is noisy or scarce. SciER (EMNLP 2024) is an entity and relation extraction dataset covering datasets, methods, and tasks in scientific documents; DMDD (TACL) and SciDMT (LREC-COLING 2024) target dataset and scientific-mention detection at scale. DynClean (NAACL 2025) uses training dynamics to clean labels for distantly-supervised NER, and Many-Shot In-Context Learning for NER (ACL 2026) asks how far in-context learning can substitute for annotation in the low-resource case.

Structure for scientific domains

A related thread applies this to specific scientific fields, where the vocabulary is the hard part. Taxonomy-Driven Knowledge Graph Construction (ACL 2025) builds domain knowledge graphs from a taxonomy rather than a flat schema; ClimateIE and Querying Climate Knowledge (both 2025) bring information extraction and semantic retrieval to climate science; FlowLearn (ECAI 2024) extends the question past text, evaluating how well vision-language models read flowcharts.

Publications

Underline marks equal contribution. A continuously updated list lives on Google Scholar.

2026

2025

2024

2023

2022