PARSEME corpus release 1.3

SAVARY, Agata, BEN KHELIL, Cherifa, RAMISCH, Carlos, GIOULI, Voula, BARBU MITITELU, Verginica, HADJ MOHAMED, Najet, KRSTEV, Cvetana, LIEBESKIND, Chaya, XU, Hongzhi, STYMNE, Sara, GÜNGÖR, Tunga, PICKARD, Thomas, GUILLAUME, Bruno, BEJČEK, Eduard, BHATIA, Archna, CANDITO, Marie, GANTAR, Polona, INURRIETA, Uxoa, GATT, Albert, KOVALEVSKAITE, Jolanta, LICHTE, Timm, LJUBEŠIĆ, Nikola, MONTI, Johanna, PARRA ESCARTÍN, Carla, SHAMSFARD, Mehrnoush, STOYANOVA, Ivelina, VINCZE, Veronika and WALSH, Abigail (2023). PARSEME corpus release 1.3. In: BHATIA, Archna, EVANG, Kilian, GARCIA, Marcos, GIOULI, Voula, HAN, Lifeng and TASLIMIPOOR, Shiva, (eds.) Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023). Dubrovnik, Croatia, Association for Computational Linguistics, 24-35. [Book Section]

Documents
37456:1289506
[thumbnail of PARSEME.pdf]
Preview
PDF
PARSEME.pdf - Published Version
Available under License Creative Commons Attribution.

Download (359kB) | Preview
Abstract
We present version 1.3 of the PARSEME multilingual corpus annotated with verbal multiword expressions. Since the previous version, new languages have joined the undertaking of creating such a resource, some of the already existing corpora have been enriched with new annotated texts, while others have been enhanced in various ways. The PARSEME multilingual corpus represents 26 languages now. All monolingual corpora therein use Universal Dependencies v.2 tagset. They are (re-)split observing the PARSEME v.1.2 standard, which puts impact on unseen VMWEs. With the current iteration, the corpus release process has been detached from shared tasks; instead, a process for continuous improvement and systematic releases has been introduced.
More Information
Statistics

Downloads

Downloads per month over past year

View more statistics

Metrics

Altmetric Badge

Dimensions Badge

Share
Add to AnyAdd to TwitterAdd to FacebookAdd to LinkedinAdd to PinterestAdd to Email

Actions (login required)

View Item View Item