Separating Named Entities

Authors	ULIPOVÁ Barbora GRÁC Marek
Year of publication	2014
Type	Article in Proceedings
Conference	Eighth Workshop on Recent Advances in Slavonic Natural Language Processing
MU Faculty or unit	Faculty of Arts
Citation
web	https://nlp.fi.muni.cz/raslan/2014/15.pdf
Field	Linguistics
Keywords	text corpus; mutual information; named entities
Description	In this paper, we analyze the situation of long sequences of mostly capitalized words which look like a named entity but in fact they consist of several named entities. An example of such phenomena is hokejista (hockey player) New York Rangers Jaromír Jágr. Without splitting the sequence correctly, we will wrongly assume that the whole capitalized sequence is a name of the hockey player. To find out how the sequence should be split into the correct named entities, we tested several methods. These methods are based on the frequencies of the words they consist of and their n-grams. The method DIFF-2 proposed in this article obtained much better results than MI-score or logDice.
Related projects:	Zastoupení ČR v European Research Consortium for Informatics and Mathematics Čeština v jednotě synchronie a diachronie - 2014