Probabilistic string similarity sketching software development for metagenomics and RNA-seq
INSTITUTION
Carnegie Mellon University, PA
PRINCIPAL INVESTIGATOR
Yun W Yu
FUNDING
$300K
YEAR
2025
MOONBASE SCORE
Still being scored
LOADING MOONBASE SCORE
Abstract
One of the major approaches modern biologists use for understanding living things is to sequence their “genomes,” which involves reading their genomes and comparing them to known biological databases. A lot of software for analyzing genomes relies on a bag of tricks that scientists have learned to make work using trial and error, but without mathematical proofs that they work. This research will provide rigorous mathematical proofs for when those analysis tricks are guaranteed to work, and also extend those methods to additional biological analysis problems. Specifically, this research will analyze analysis tricks from genome “alignment,” which measure how many differences there are between two genomes, and then apply those tricks to the problems of discovering new variants of proteins and measuring RNA levels in a cell. The broader impact of this work is that researchers will then be able to build faster genomic analysis software, improving our understanding of when living cells produce different forms of proteins. Researchers will perform an average-case analysis of the seed-chain-extend string alignment algorithm to prove bounds on speed and accuracy for sequence alignment and read mapping software. The researchers previously performed such an analysis in the substitution-only error model, but here are extending their analysis to a more biologically-plausible error model including indels, duplications, and subsampling. Some of the probabilistic subsampling techniques will be used to improve RNA-seq quantification and novel isoform discovery. The aim is to improve the speed of RNA-seq quantification by not mapping every individual read and to improve the accuracy of novel isoform discovery by filtering the reads to plausible novel isoform candidates using a subsampling pre-filter prior to mapping. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
Are you the primary organization running this research?
The two tools below are built for the principal investigator & host institution behind this project.