Skip to content

Author

Jerry Ma

We have 2 of 7 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations

LLM assistants and agentic systems log long multi-turn conversations. AI providers often scan these conversations for Personally Identifiable Information (PII) and mask the PII before storing or processing conversation data. Yet most PII detectors and benchmarks target self-contained records rather than cross-turn eval...

Kai-Yuan Zhang, Chuan Wang, Joey Zhong et al. · 0 citations
Review Aug 2026

WANDR: A Benchmark for Wide and Deep Research

WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents that replaces static gold answer sets with task-specific judges that refetch cited pages and verify each record against its evidence, allowing evaluation of current and changing facts.

Vitaliy Polshkov, Marcin Pitera, Jeremy Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.