Reproducible Text Classification Workflows

Dependency-light tools and tutorials for teaching reproducible text classification. The package covers HTML text extraction, sentence segmentation, text preprocessing, document-term matrices, TF-IDF, keyword extraction, cosine similarity, stratified cross-validation, classification metrics, and a multinomial Naive Bayes classifier. It modernizes the code accompanying Kobayashi, V. B., Berkers, H. A., Mol, S. T. Kismihok, G., and Den Hartog, D. N. (2017) The package replaces the original scripts in the paper.


textclassificationtutorial

textclassificationtutorial is an installable R package for learning and teaching reproducible text classification. It modernizes the code accompanying:

Kobayashi, V. B., Mol, S. T., Berkers, H. A., Kismihók, G., & Den Hartog, D. N. (2018). Text classification for organizational researchers: A tutorial. Organizational Research Methods, 21(3), 766–799.

The package turns the original sequence-dependent scripts into documented, testable functions.

Installation

Install the released version from CRAN, or obtain the development source from the GitHub repository. No package function, example, test, or vignette installs dependencies.

Quick start

library(textclassificationtutorial)

documents <- c(
  "Analyze customer data and build statistical models.",
  "Create dashboards and communicate analytical findings.",
  "Provide nursing care and support patients.",
  "Coordinate treatment with physicians and nurses."
)

labels <- c("data", "data", "care", "care")

clean <- preprocess_text(
  documents,
  stopwords = c("and", "with"),
  min_token_length = 2
)

dtm <- document_term_matrix(clean)
tf_idf(dtm)
extract_keywords(dtm, n = 2)

model <- fit_naive_bayes(dtm, labels)
predictions <- predict(model, dtm)
classification_metrics(labels, predictions, positive = "data")

What the package covers

  • HTML extraction from a file, string, or directory
  • Lightweight sentence segmentation
  • Transparent text normalization and stopword removal
  • Document-term and TF-IDF matrices
  • Keyword extraction and cosine similarity
  • Stratified repeated cross-validation folds
  • Binary classification evaluation
  • Multinomial Naive Bayes without a heavy modelling dependency

Read the tutorials with:

vignette("getting-started", package = "textclassificationtutorial")
vignette("model-evaluation", package = "textclassificationtutorial")

Design principles

Package functions never install dependencies, change global warning settings, open browsers, delete output directories, or depend on objects produced by a previous script. File-writing decisions remain with the user.

The lightweight sentence splitter and classifier are deliberately transparent for teaching. For production NLP, use language-specific tokenizers, carefully validated preprocessing, and models appropriate to the research design.

Development

Development dependencies must be installed explicitly by the developer before running documentation, tests, or package checks. The package never installs missing packages automatically.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("textclassificationtutorial")

0.1.2 by Vladimer Kobayashi, a month ago


https://github.com/vkobayashi/textclassificationtutorial


Report a bug at https://github.com/vkobayashi/textclassificationtutorial/issues


Browse source code at https://github.com/cran/textclassificationtutorial


Authors: Vladimer Kobayashi [aut, cre] , Stefan Mol [aut] , Gabor Kismihok [aut]


Documentation:   PDF Manual  


Apache License (>= 2) license


Suggests knitr, rmarkdown, testthat, xml2


See at CRAN