Unlocking the Data Inside Molecular Maps
BU researchers receive a $3.37 million NIH grant to develop computational tools that could reveal new insights into cancer and other diseases
Scientists can now pinpoint individual RNA molecules inside tissue, creating extraordinarily detailed maps of where genes are active and how that activity relates to neighboring cells. Emerging technologies known as spatial subcellular transcriptomics are making these molecular maps possible—but the enormous datasets they generate are increasingly difficult to analyze.
A Boston University team is tackling that challenge through a $3.37 million National Institutes of Health (NIH) funded project led by Ruben Dries, assistant professor at the Chobanian & Avedisian School of Medicine, researcher at Boston Medical Center, and Hariri Institute faculty affiliate.The team will develop open computational infrastructure and new analytical methods to help researchers work with these massive datasets and uncover biological insights without relying on specialized computing resources.
Dries is joined by co-investigators Mark Crovella, professor of computer science and computing & data sciences, and Shariq Mohammed, assistant professor of biostatistics, both fellow Hariri Institute Faculty Affiliates.
“Many labs are now generating these datasets,” Dries says. But while increasingly turnkey technologies have made them easier to produce, obtaining usable insights takes many times longer.”

Making every molecule count
The research team is focusing on three related problems: managing the sheer volume of data, extracting more of the molecular information it contains, and finding patterns that hold up across multiple tissue samples.
Some of the groundwork grew out of challenges Eddie Ruiz (Genetics & Genomics PhD ’26) encountered firsthand during his graduate research with Dries, as spatial-omics datasets began pushing the limits of conventional analysis. With support from a Hariri Institute Graduate Student Fellowship, Ruiz explored emerging database technologies capable of handling much larger datasets while preserving familiar research workflows. That early research helped lay the foundation for part of the new NIH project, which he is now advancing as a postdoctoral associate in the Dries Lab.
The NIH project will leverage online analytical databases (OLAPs) not simply to store and query data, but to analyze and visualize it at scale. A key goal is to enable researchers to work with these datasets on their own computers rather than relying on costly GPUs or high-performance computing systems.

That approach is already taking shape in dbverse, an open-source software framework developed by Ruiz and colleagues that allows researchers to analyze datasets larger than a computer’s available memory without fundamentally changing their workflows. In a preprint released August 18, the team reports that its database-backed approach can perform some core operations up to 1,000 times faster than established methods and scale spatial-omics analyses to millions of cells on a computer with 16 GB of memory.
The work provides an early demonstration of the database-backed approach that the new NIH project will expand.
“We want to provide solutions that don’t require costly infrastructure,” Dries says. Making the analysis accessible to more researchers, he says, could ultimately broaden the technology’s impact.
Scale, however, is only one part of the problem. Current analyses don’t always take full advantage of what makes these datasets distinctive: the precise location of individual RNA molecules. Most approaches first assign molecules to cells and then analyze gene activity cell by cell. In doing so, they can lose some of the spatial information captured by the technology.
The researchers, including graduate student Wonyl Choi from the department of computer science, are developing methods that instead learn from the locations of the molecules themselves, drawing in part on concepts from physics about how interacting particles organize. The goal is to reveal features such as cell boundaries, tissue structure and biological processes that conventional analyses may miss, including in three dimensions.
A different challenge arises when researchers move from one tissue sample to many.
“An unusual arrangement of cells and molecules in a single tumor may be intriguing, but researchers need to know whether the same pattern appears repeatedly—and whether it is connected to something clinically meaningful,” says Dries.

The team is developing statistical methods to identify shared spatial patterns across samples and relate them to clinical or experimental outcomes. In cancer, for example, that could help researchers investigate whether recurring patterns of cells and gene activity are associated with how tumors respond to treatments such as immunotherapy.
“We want to dramatically accelerate what researchers can learn from these massive spatial datasets,” Dries says. “That means making them easier to analyze, using information that is often overlooked, and finding important patterns across multiple samples.”
Making tools accessible
The project is designed to make its methods available well beyond the Dries Lab. The resulting tools will be released independently and integrated into Giotto Suite, the team’s open-source platform for spatial biology, whose latest generation was published in Nature Methods in 2025.
For Dries, open software is only useful if researchers can actually put it to work. That means making tools straightforward to install, compatible with common analyses, and supported by documentation and examples—resources that can help both scientists entering the field and emerging AI-based research systems.
“Building the right tools isn’t enough—they also need to be easy to access, integrate and use,” Dries says. “That’s essential both for researchers entering the field and for developing AI systems that can help accelerate discovery.”
By pairing increasingly powerful experimental technologies with computational methods capable of keeping pace, the team hopes to narrow the gap between mapping individual molecules and understanding what their organization reveals about health and disease.
The source code for Giotto Suite may be found on our GitHub repository: https://github.com/giotto-suite.
Learn more about this work in this YouTube video on
