Moonbase
← Back to Awards
Directorate for Biological SciencesNSF · NSFNSF

Probabilistic string similarity sketching software development for metagenomics and RNA-seq

Yun W Yu·Carnegie Mellon University, PA·2025–2027·ACTIVE
Donate

INSTITUTION

Carnegie Mellon University, PA

PRINCIPAL INVESTIGATOR

Yun W Yu

FUNDING

$300K

YEAR

2025

MOONBASE SCORE

Still being scored

LOADING MOONBASE SCORE

Abstract

One of the major approaches modern biologists use for understanding living things is to sequence their “genomes,” which involves reading their genomes and comparing them to known biological databases. A lot of software for analyzing genomes relies on a bag of tricks that scientists have learned to make work using trial and error, but without mathematical proofs that they work. This research will provide rigorous mathematical proofs for when those analysis tricks are guaranteed to work, and also extend those methods to additional biological analysis problems. Specifically, this research will analyze analysis tricks from genome “alignment,” which measure how many differences there are between two genomes, and then apply those tricks to the problems of discovering new variants of proteins and measuring RNA levels in a cell. The broader impact of this work is that researchers will then be able to build faster genomic analysis software, improving our understanding of when living cells produce different forms of proteins. Researchers will perform an average-case analysis of the seed-chain-extend string alignment algorithm to prove bounds on speed and accuracy for sequence alignment and read mapping software. The researchers previously performed such an analysis in the substitution-only error model, but here are extending their analysis to a more biologically-plausible error model including indels, duplications, and subsampling. Some of the probabilistic subsampling techniques will be used to improve RNA-seq quantification and novel isoform discovery. The aim is to improve the speed of RNA-seq quantification by not mapping every individual read and to improve the accuracy of novel isoform discovery by filtering the reads to plausible novel isoform candidates using a subsampling pre-filter prior to mapping. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

Directorate for Biological SciencesInnovation: BioinformaticsADVANCES IN BIO INFORMATICSthroughunderstandingscientistsprobabilisticboundschainfilterdiscoveringworthyreflectsfilteringmathematicalisoformadditionalmeritpreviouslyeverywithoutthingssubsampling

Are you the primary organization running this research?

The two tools below are built for the principal investigator & host institution behind this project.