Improving Indian Address Parsing in Data-Scarce Environments Using Chunk-Based Retrieval-Augmented Transformer Models
Accurate parsing of unstructured Indian addresses remains challenging due to the linguistic variability and the limited annotated data. While transformer-based models achieve near-perfect performance on synthetic datasets, their generalization to real-world inputs is limited, with F1-scores degrading substantially for...