Skip to content
Open access

A Spatio-Temporal Multimodal Deep Learning Framework for Taxi Demand Prediction From Multi-Source and Mixed-Type Data

2026 · IEEE Access · Vol 14, pp. 134419-134435 · 0 citations · 53 references

Abstract

Metro systems are fast and high-capacity but cannot provide door-to-door service, so their accessibility depends on first- and last-mile connections. Taxis bridge the longer and less reliable of these trips, so predicting taxi demand at metro stations is essential for allocating feeder fleets and improving accessibility in high-demand but low-coverage urban peripheries. The task is difficult because it couples non-linear temporal dynamics with station-specific characteristics, such as proximity to commercial hubs or a role as an interchange, that models built on raw trip history omit. We propose a multimodal deep learning framework that integrates historical taxi Global Positioning System (GPS) trajectories, metro passenger flows, and categorical spatial identifiers. A Long Short-Term Memory (LSTM) module captures long-range temporal dependencies and an entity embedding layer maps discrete station IDs to dense, low-dimensional latent vectors; the fused representation feeds a Multi-Layer Perceptron (MLP) regressor. We validated the framework on approximately 700 million GPS records from the Bangkok Metropolitan Region against twelve learned baselines, spanning classical regressors such as eXtreme Gradient Boosting (XGBoost) and recurrent, attention-based, and graph-based deep architectures, plus measured persistence and seasonal-naive predictors. The proposed model is the most accurate at every tested history length at the hourly resolution and competitive with the strongest baselines at the daily resolution. The ablation study reveals that the decisive inputs are the calendar markers, day-of-week at the daily resolution (+ 3.8% RMSE when removed) and hour-of-day at the hourly (+ 1.1%), followed at both by the station identity embedding (+ 1.9% and + 0.9%), whereas removing any single flow series costs less than 1%, making multi-source integration a framework capability rather than a demonstrated accuracy gain. The embedding is only partly interpretable: a few hourly dimensions correlate moderately ( $\rho \le 0.33$ ) with transport-infrastructure indicators, and no daily correlation survives false-discovery-rate correction. Because the embedding is a lookup table trained and tested on the same stations under a chronological split, the evidence establishes forecasting for observed stations rather than generalization to unseen ones.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.