Improving Indian Address Parsing in Data-Scarce Environments Using Chunk-Based Retrieval-Augmented Transformer Models
Abstract
Accurate parsing of unstructured Indian addresses remains challenging due to the linguistic variability and the limited annotated data. While transformer-based models achieve near-perfect performance on synthetic datasets, their generalization to real-world inputs is limited, with F1-scores degrading substantially for context-dependent entities such as ROAD and SUBURB. To mitigate this gap, in this paper, a chunk-based retrieval-augmented framework that integrates transformer-based sequence labeling with a geospatial knowledge base derived from OpenStreetMap is proposed. The method performs vector-based retrieval over multi-token spans and selectively refines predictions using a confidence threshold (τ = 1.2). Across three independent training seeds on a manually reviewed test set of 500 real-world addresses, the proposed approach yields a substantial improvement for the SUBURB entity (F1 from 0.162 ± 0.016 to 0.519 ± 0.042), with smaller but statistically significant gains for ROAD (0.011 ± 0.008 to 0.151 ± 0.070), where retrieval alone is insufficient to resolve free-form descriptive spatial language. Statistical validation via McNemar's test (p = 2.8 × 10⁻²³ for SUBURB, p = 3.6 × 10⁻¹² for ROAD) and test-set bootstrapping (95% CIs) confirms the significance and stability of these improvements. The results indicate that retrieval-based augmentation is effective for grounding named geospatial entities, while its applicability to descriptive spatial expressions remains an open problem.