Naslov (eng)

Creating a Stop Word Dictionary in Serbian

Autor

Avdić, Aldina R.
Ljajić, Adela B.
Marovac, Ulfeta A.

Opis (eng)

Abstract: By using natural language processing techniques, it is possible to get a lot of information from the extraction of document topics through mapping of document key words or content-based classification of documents, etc. To get this information, an important step is to separate words that carries informative value in a sentence from those words that do not affect its meaning. By using dictionaries of stop words specific to each natural language, the marking of words that do not carry meaning in the sentence is achieved. This paper presents creating a stop word dictionary in Serbian. The influence of stop words to text processing is presented on three different data set. It is shown that by using proposed dictionary of Serbian stop words the data set dimension is reduced from 15% to 39%, while the quality of the obtained n-gram language models is improved.

Jezik

engleski

Datum

2021

Licenca

Creative Commons licenca
Ovo delo je licencirano pod uslovima licence
Creative Commons CC BY-SA 4.0 - Creative Commons Autorstvo - Deliti pod istim uslovima 4.0 International License.

http://creativecommons.org/licenses/by-sa/4.0/legalcode

Predmet

Keywords: stop words, Serbian, text mining, natural language processing, normalization.

Deo kolekcije (1)

o:28516 Radovi nastavnika i saradnika Državnog univerziteta u Novom Pazaru