Text Alignment with Smith-Waterman

Find similarities between texts using the Smith-Waterman algorithm. The algorithm performs local sequence alignment and determines similar regions between two strings. The Smith-Waterman algorithm is explained in the paper: "Identification of common molecular subsequences" by T.F.Smith and M.S.Waterman (1981), available at . This package implements the same logic for sequences of words and letters instead of molecular sequences.


text.alignment

This repository contains an R package for aligning texts using the Smith-Waterman algorithm. This is especially usefull if you want to

  • Find names in documents even if they are not correctly spelled
  • Match 2 texts
  • Find relevant sequences of texts in other texts

Installation

  • For regular users, install the package from your local CRAN mirror install.packages("text.alignment")
  • For installing the development version of this package: remotes::install_github("DIGI-VUB/text.alignment", build_vignettes = TRUE)

Look to the vignette and the documentation of the functions

vignette("textalignment", package = "text.alignment")
help(package = "text.alignment")

DIGI

By DIGI: Brussels Platform for Digital Humanities: https://digi.research.vub.be

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("text.alignment")

0.1.5 by Jan Wijffels, 10 months ago


https://github.com/DIGI-VUB/text.alignment


Browse source code at https://github.com/cran/text.alignment


Authors: Jan Wijffels [aut, cre, cph] (Rewrite of functionalities from the textreuse R package) , Vrije Universiteit Brussel - DIGI: Brussels Platform for Digital Humanities [cph] , Lincoln Mullen [ctb, cph]


Documentation:   PDF Manual  


MIT + file LICENSE license


Imports Rcpp

Suggests knitr, markdown, rmarkdown

Linking to Rcpp


See at CRAN