Skip to content
Review Open access

Improving Indian Address Parsing in Data-Scarce Environments Using Chunk-Based Retrieval-Augmented Transformer Models

Aug 2026 · Engineering, Technology & Applied Science Research · 0 citations · 27 references

Abstract

Accurate parsing of unstructured Indian addresses remains challenging due to the linguistic variability and the limited annotated data. While transformer-based models achieve near-perfect performance on synthetic datasets, their generalization to real-world inputs is limited, with F1-scores degrading substantially for context-dependent entities such as ROAD and SUBURB. To mitigate this gap, in this paper, a chunk-based retrieval-augmented framework that integrates transformer-based sequence labeling with a geospatial knowledge base derived from OpenStreetMap is proposed. The method performs vector-based retrieval over multi-token spans and selectively refines predictions using a confidence threshold (τ = 1.2). Across three independent training seeds on a manually reviewed test set of 500 real-world addresses, the proposed approach yields a substantial improvement for the SUBURB entity (F1 from 0.162 ± 0.016 to 0.519 ± 0.042), with smaller but statistically significant gains for ROAD (0.011 ± 0.008 to 0.151 ± 0.070), where retrieval alone is insufficient to resolve free-form descriptive spatial language. Statistical validation via McNemar's test (p = 2.8 × 10⁻²³ for SUBURB, p = 3.6 × 10⁻¹² for ROAD) and test-set bootstrapping (95% CIs) confirms the significance and stability of these improvements. The results indicate that retrieval-based augmentation is effective for grounding named geospatial entities, while its applicability to descriptive spatial expressions remains an open problem.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.