Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
SmolDocling, a compact 256M-parameter vision-language model (VLM), is fine-tune to perform end-to-end key-value extraction directly from document images, jointly solving identification, localization, and association in a single pass without OCR preprocessing.
A. Gurbuz, A. Nassar, Christoph Auer et al.
· 0 citations