El'Manuscript-14
Conference
Organizing Committee
Programme Committee
Conference topics
Workshops & Tutorials
Important Dates
Registration & Submission
Registration fee
List of participants
Conference program
Conference Abstracts
Workshops program
Workshops Abstracts
Contact
Leisure & Tourism
Photos





Lost Password?
No account yet? Register
We have 9 guests online
RSS-feed
rss20.gif

The project is supported by the Russian Foundation for Basic Research, project #07-04-12140в

(c) "Information Technologies and Textual Heritage", 2008-2020

Bulgarian National Corpus PDF Print E-mail
Written by: Светла Коева   
Четверг, 07 Август 2014
Lecture

We will discuss several key concepts related to the development of corpora and reconsider them in light of recent developments in Natural Language Processing. We propose a data-driven approach to corpus design, which integrates the best practices of traditional corpus linguistics with the potential of the latest technologies allowing fast collection, automatic metadata description and annotation of large amounts of data.

We will illustrate this concept with a description of the compilation, structuring, documentation, and annotation (morphosyntactic tagging, lemmatisation, word-sense annotation, annotation of noun phrases and named entities) of the Bulgarian National Corpus (http://ibl.bas.bg/en/BGNC_access_en.htm; http://ibl.bas.bg/en/BGNC_en.htm; http://search.dcl.bas.bg/). We will conclude with a brief evaluation of the quality of the corpus and an outline of its applications in Natural Language Processing and linguistic research.


 
< Prev   Next >