KorAP Architecture – Diving in the Deep Sea of Corpus Data

Nils Diewald, Michael Hanl, Eliza Margaretha, Joachim Bingel, Marc Kupietz, Piotr Banski, Andreas Witt

12 Citationer (Scopus)

Abstract

KorAP is a corpus search and analysis platform, developed at the Institute for the German Language (IDS). It supports very large corpora with multiple annotation layers, multiple query languages, and complex licensing scenarios. KorAP's design aims to be scalable, flexible, and sustainable to serve the German Reference Corpus DeReKo for at least the next decade. To meet these requirements, we have adopted a highly modular microservice-based architecture. This paper outlines our approach: An architecture consisting of small components that are easy to extend, replace, and maintain. The components include a search backend, a user and corpus license management system, and a web-based user frontend. We also describe a general corpus query protocol used by all microservices for internal communications. KorAP is open source, licensed under BSD-2, and available on GitHub.

OriginalsprogEngelsk
TitelProceedings of the 10th conference of the Language Resources and Evaluation Conference
Antal sider6
ForlagEuropean Language Resources Association
Publikationsdato2016
Sider3586-3591
ISBN (Trykt)9782951740891
StatusUdgivet - 2016
BegivenhedLREC 2016 -
Varighed: 23 maj 201628 maj 2016

Konference

KonferenceLREC 2016
Periode23/05/201628/05/2016

Fingeraftryk

Dyk ned i forskningsemnerne om 'KorAP Architecture – Diving in the Deep Sea of Corpus Data'. Sammen danner de et unikt fingeraftryk.

Citationsformater