We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
Agata Savary, Manon Scholivet, Carlos Ramisch et al.· International Conference on...· 1 citation
LLM-assisted sense assignment with a Serbian WordNet-based custom inventory, iterative inventory expansion, and expert validation is combined with a constrained JSON-formatted output to support the practical construction and refinement of sense-annotated resources in a low-resource setting.
Saša Petalinkar, R. Stanković, Milica Ikonić Nešić et al.· Intelligent Data Analysis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.