Publication:

Advancing virome analysis: New computational methods for viral profiling and characterization with applications to inflammatory bowel disease

Loading...
Thumbnail Image

Date

2026-05-09

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Jensen, Jordan S. L.. 2026. Advancing virome analysis: New computational methods for viral profiling and characterization with applications to inflammatory bowel disease. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Viruses are crucial components of microbial communities, including bacteriophages that shape bacterial populations and eukaryotic viruses that impact host health and ecosystem function. However, they are poorly captured by most current experimental and computational approaches, due to a combination of factors including the small size of viral genomes and resulting predominance of bacterial genomic content in sequencing data, diverse nucleic acid chemistries, lack of universal marker genes, and limited well-characterized viral reference data.1,2 Current approaches for detecting and classifying viruses in sequencing data remain limited, each constrained by trade-offs such as the exclusion of low-abundance viruses, lack of standardization, reliance on incomplete reference databases, and reduced sensitivity to highly divergent viruses. Consequently, the majority of observed viral sequences remain unclassified and underexplored. These challenges highlight the need for a standardized, comprehensive framework for viral profiling from multi-omic data, as well as broader methodological advances that enable improved investigations into how viruses influence human health.

To address these challenges, we developed the BAQLaVa algorithm for high-resolution profiling of >120,000 viral species (viral genome bins, VGBs) directly from metagenome (MGX) or metatranscriptome (MTX) sequencing data (Chapter 1). This method integrates two complementary modules: one based on VGB-specific nucleotide markers optimized for specificity, and another based on viral proteome sets optimized for sensitivity. The proteome-based module further enables assignment of previously unseen or near-neighbor viruses to their closest VGB, improving classification in the presence of incomplete reference databases. In comprehensive benchmarking, BAQLaVa substantially outperformed alternative profiling approaches, consistently achieving species-level recall and precision over 90%. We applied BAQLaVa to MGX and MTX samples from the HMP2 IBDMDB cohort to identify previously undescribed viral perturbations in inflammatory bowel diseases (IBD), (Chapter 2). Most notably, virome diversity was reduced in tandem with bacterial diversity during inflammation, in contrast to previous findings based on a narrower range of viral detection. A subset of viruses was enriched during IBD, and associated with carriage of abortive infection anti-defense systems such as AbiL and PD-𝜆-2, as well as genes involved in the regulation of lysogeny.

We further investigated improvements to viral analysis methodologies beyond direct viral quantification from MGX and MTX data, focusing on approaches enabled by BAQLaVa or that build upon its approach. We first examined phage-host prediction. Leveraging recent advances in ensemble host-prediction tools and the underutilized resource of virus-bacteria abundance patterns across large paired datasets, we developed an improved phage-host prediction framework and evaluated its performance in the HMP2 dataset (Chapter 3). Specifically, we evaluated the utility of virus-bacteria abundance covariation and co-occurrence and found that integrating these signals within an existing ensemble approach improves performance over all individual methods. We next explored deep learning approaches for viral taxonomic classification, addressing the need to classify novel or highly divergent sequences while supporting long-read sequencing data. We developed two neural network models: a BERT-based binary classifier to distinguish viral from non-viral sequences and a dense convolutional neural network to predict viral taxonomy at the genus level (Chapter 4). Both models demonstrated strong performance (binary classifier AUC = 0.88-0.98 across evaluation datasets; taxonomic classifier balanced accuracy = 80%), highlighting the potential of neural network-based approaches for viral classification independent of reference databases.

Together, these results represent important contributions to computational viral genomics, particularly in highly novel and health-associated environments. BAQLaVa enables sensitive and specific viral profiling from metagenomic and metatranscriptomic data, providing a scalable framework for virome epidemiology. Our phage-host prediction methods facilitate more systematic analysis of virus-host interactions across large datasets, while our exploration of neural network approaches demonstrates their effectiveness in capturing underlying viral sequence and taxonomic features, an insight that can be leveraged to further advance viral discovery. Finally, applying these approaches in the context of inflammatory bowel disease revealed disease-associated viral species, traits, and bacterial host interactions, offering a foundation for further investigation.

Description

Other Available Sources

Research Data

Keywords

Computational platforms, Databases, Inflammatory bowel disease, Phage-host prediction, Viral genomics, Bioinformatics, Virology, Microbiology

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories