Customisable Stop-Words in 110 Languages

Functions to generate stop-word lists in 110 languages, in a way consistent across all the languages supported. The generated lists are based on the morphological tagset from the Universal Dependencies.


tidystopwords: R package for multilingual stopwords

Authors: Silvie Cinková*, Maciej Eder
License: GPL-3

An R package containing customizable lists of stopwords in multiple languages; it attempts to follow tidy data principles.

The idea behind this package is to provide stopwords for less-resourced languages as well as give the user control over the stopword selection with respect to parts of speech. For the purposes of this package, stopwords are defined as forms of function words from closed parts of speech (e.g. prepositions, conjunctions, auxiliary verbs, and pronouns). The core generate_stoplist() function relies on multilingual_stopwords(), a large data frame derived from the current release of the Universal Dependencies Treebanks. We have included all languages. The data comes encoded in UTF-8. The vocabulary coverage for each language depends on the size, textual diversity, and annotation quality of the available treebanks. No manual post-editing was performed.

Installation

Install the package directly from the GitHub repository:

library(devtools)
install_github("computationalstylistics/stopwoRds", build_vignettes = TRUE)

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("tidystopwords")

0.9.1 by Maciej Eder, 5 years ago


Browse source code at https://github.com/cran/tidystopwords


Authors: Silvie Cinkova [aut] , Maciej Eder [aut, cre]


Documentation:   PDF Manual  


GPL (>= 3) license


Imports dplyr

Suggests knitr, rmarkdown


See at CRAN