Providing a Catalogue of Language Resources for Commercial Users

Bente Maegaard, Lina Henriksen, Claus Povlsen, Sussi Olsen, Andrew Joscelyne, Vesna Lusicky, Margaretha Mazura, Philippe Wacker

1 Citationer (Scopus)
41 Downloads (Pure)

Abstract

Language resources (LR) are indispensable for the development of tools for machine translation (MT) or various kinds of computer-assisted translation (CAT). In particular language corpora, both parallel and monolingual are considered most important for instance for MT, not only SMT but also hybrid MT. The Language Technology Observatory will provide easy access to information about LRs deemed to be useful for MT and other translation tools through its LR Catalogue. In order to determine what aspects of an LR are useful for MT practitioners, a user study was made, providing a guide to the most relevant metadata and the most relevant quality criteria. We have seen that many resources exist which are useful for MT and similar work, but the majority are for (academic) research or educational use only, and as such not available for commercial use. Our work has revealed a list of gaps: coverage gap, awareness gap, quality gap, quantity gap. The paper ends with recommendations for a forward-looking strategy.

OriginalsprogEngelsk
TitelProceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)
Antal sider8
ForlagEuropean Language Resources Association
Publikationsdato2016
Sider449-456
ISBN (Trykt)978-2-9517408-9-1
StatusUdgivet - 2016
BegivenhedLREC 2016 -
Varighed: 23 maj 201628 maj 2016

Konference

KonferenceLREC 2016
Periode23/05/201628/05/2016

Fingeraftryk

Dyk ned i forskningsemnerne om 'Providing a Catalogue of Language Resources for Commercial Users'. Sammen danner de et unikt fingeraftryk.

Citationsformater