A suite of functions for rapid and flexible analysis of codon usage bias. It provides in-depth analysis at the codon level, including relative synonymous codon usage (RSCU), tRNA weight calculations, machine learning predictions for optimal or preferred codons, and visualization of codon-anticodon pairing. Additionally, it can calculate various gene- specific codon indices such as codon adaptation index (CAI), effective number of codons (ENC), fraction of optimal codons (Fop), tRNA adaptation index (tAI), mean codon stabilization coefficients (CSCg), and GC contents (GC/GC3s/GC4d). It also supports both standard and non-standard genetic code tables found in NCBI, as well as custom genetic code tables.
Comprehensive Codon Usage Bias Analysis in R
Codon usage bias refers to the non-uniform usage of synonymous codons (codons that encode the same amino acid) across different organisms, genes, and functional categories. cubar is a comprehensive R package for analyzing codon usage bias in coding sequences. It provides a unified framework for calculating established codon usage metrics, conducting sliding-window analyses or differential usage analyses, and optimizing sequences for heterologous expression.
Biostrings and data.table backendsInstall the latest stable version from CRAN:
install.packages("cubar")
Install the latest development version from GitHub:
# Install devtools if not already installed
if (!requireNamespace("devtools", quietly = TRUE)) {
install.packages("devtools")
}
# Install cubar from GitHub
devtools::install_github("mt1022/cubar", dependencies = TRUE)
System Requirements:
Required Packages:
Biostrings (≥ 2.60.0) - Bioconductor package for sequence manipulationIRanges (≥ 2.34.0) - Bioconductor infrastructure for range operationsdata.table (≥ 1.14.0) - High-performance data manipulationggplot2 (≥ 3.3.5) - Data visualizationrlang (≥ 0.4.11) - Language toolsNote: Bioconductor packages will be installed automatically, but you may need to update your R installation if you encounter compatibility issues.
📖 Complete documentation is available within R (?function_name) and on our package website.
Here's a typical analysis workflow demonstrating key functionality:
library(cubar)
library(ggplot2)
# 1. Load and quality-check sequences
data(yeast_cds)
clean_cds <- check_cds(yeast_cds)
# 2. Calculate codon frequencies
codon_freq <- count_codons(clean_cds)
# 3. Calculate multiple metrics
enc <- get_enc(codon_freq) # Effective number of codons
gc3s <- get_gc3s(codon_freq) # GC content at 3rd positions
# 4. Analyze highly expressed genes
data(yeast_exp)
yeast_exp <- yeast_exp[yeast_exp$gene_id %in% rownames(codon_freq), ]
high_expr <- head(yeast_exp[order(-yeast_exp$fpkm), ], 500)
rscu_high <- est_rscu(codon_freq[high_expr$gene_id, ])
cai <- get_cai(codon_freq, rscu_high)
# 5. Visualize results
df <- data.frame(ENC = enc, CAI = cai, GC3s = gc3s)
ggplot(df, aes(color = GC3s, x = ENC, y = CAI)) +
geom_point(alpha = 0.6) +
scale_color_viridis_c() +
labs(title = "Codon Usage Bias Relationships",
x = "Effective Number of Codons", y = "Codon Adaptation Index")
?function_name) and online docsFor complementary analysis, consider these R packages:
This project is licensed under the MIT License - see the LICENSE file for details.