Skip to content

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

UTP-Bench is introduced, a large-scale benchmark for uncertainty-aware travel planning and proposes three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which define the ability of generated itineraries to main- tain robustness against transit delays and crowd variability.

Abstract

Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction and ignoring whether generated plans remain robust when such uncertainties arise. To address this limitation, we introduce UTP-Bench1 , a large-scale benchmark for uncertainty-aware travel planning. The dataset integrates real-world travel data spanning 504 cities of India, including attractions, restau- rants, accommodations, and multi-modal trans- portation networks. To model realistic disrup- tions, UTP-Bench incorporates empirical delay distributions and crowd-density patterns col- lected from major cities, enabling evaluation of travel plans under stochastic conditions. We further propose three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which quan- tify the ability of generated itineraries to main- tain robustness against transit delays and crowd variability. Experiments with state-of-the-art LLMs like GPT-5, Qwen3, Mistral and Phi-4 re- veal substantial gaps between model-generated and human-authored plans, particularly in tem- poral buffering, delay-aware transportation scheduling, and crowd-sensitive planning.

View source

Similar papers

TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning

This work proposes TRIPPULSE1, a multi- agent framework for review-grounded travel planning, and introduces Review-Grounded Per- sona Alignment (RGPA), an LLM-as-a-Judge metric for evaluating alignment with human- centric travel experiences.

Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and ev...

Xiao-Fei Yuan, Yan Zhang, Shao-Bo Qiao et al. · 1 citation
Preprint Sep 2026

Map2Route: Benchmarking Compositional Language-Grounded Route Planning over Semantic Maps

Across seven representative adapted baselines, Grounding2Route substantially outperforms existing methods in all metrics, and a substantial gap to human demonstrations remains, highlighting the difficulty of Map2Route and the considerable headroom for future progress.

Mu-Yi Bao, Hang Xu, Jin Tang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner

UDAV, an Uncertainty-Driven Adaptive VLM Waypoint Planner for UAV-guided UGV navigation is presented, demonstrating that stochastic VLM predictions provide both a stronger nominal route and an actionable uncertainty signal for selectively mitigating large planning errors.

G. Farhani, Shabnam Shabani · 0 citations
#artificial intelligence Preprint Sep 2026

CityPlanner: A Sandbox Agent for Executable Urban Planning

CityPlanner is proposed, a sandbox-agent framework for executable urban planning that consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines and is verified on a real-world benchmark.

Wen-Tao Zhang, Jing-Yuan Wang, Ze-Tong Zhou et al. · 0 citations
#artificial intelligence Review Sep 2026

CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation

We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM e...

Ye-Zhou Cheng · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.