Making Room for Bigger Scientific Analyses

By Andrei Volic

When Eddie Ruiz, a postdoctoral researcher in the Dries Lab at the Chobanian & Avedisian School of Medicine, began analyzing the mouse organogenesis and spatiotemporal dataset (MOSTA), he encountered a problem confounding many researchers in computational biomedicine. MOSTA offers a detailed view of how cells and organs develop in mouse embryos, but its enormous scale exceeded the capabilities of available analysis tools.

MOSTA contains millions of cells and individual RNA molecules. Analysis required expensive GPUs and high-performance computing systems. Enabled by his Hariri Institute Graduate Student Fellowship (GSF), Ruiz built a tool that could efficiently analyze enormous datasets like MOSTA using a standard lab laptop.

The tool is dbverse, an open-source software framework which uses database-backed analytical methods to process and query data efficiently. It allows researchers to work with datasets much larger than their computer’s available memory while also keeping familiar analysis workflows.

In a preprint released in August 2026, lead author Ruiz and other Hariri Institute researchers report a new approach for large-scale scientific data analysis using embedded analytical databases. Across benchmarks, core scientific operations were shown to run up to 1,000 times faster than existing methods. The framework also enabled biological analysis that would otherwise exceed the memory available on conventional computers.

The development of dbverse helped establish a computational foundation that the team is now building on through a $3.4 million NIH award. The NIH-funded project expands the underlying infrastructure while developing new methods to preserve spatial information, identify patterns across tissue samples, and make large-scale molecular analysis faster and more accessible.

A conceptual logo for dbverse depicting a planetary database cylinder. Image courtesy of Edward C. Ruiz.

From Bottlenecks to Breakthroughs

Ruiz began developing dbverse to overcome memory bottlenecks he encountered early in his doctoral studies while analyzing high-resolution spatial maps of biological tissues. “I initially thought there was a problem unique to my dataset, but I eventually realized I was facing a bigger challenge affecting researchers across my field,” Ruiz said. 

Early prototypes solved some of the memory bottlenecks hindering his work for a particular dataset, but devising a solution that could be generalized to new kinds of datasets and analyses remained a central challenge.  

“The implementation of dbverse took several iterations to land on. Unlike a lot of biological software, dbverse is modeled at the level of core data structures used across scientific fields. It is agnostic to any dataset or biological representation,” Ruiz said. 

By designing dbverse around data structures shared across scientific fields, Ruiz showed it was possible to extend the approach beyond the biological datasets that first motivated the project.

Reaching Beyond the Lab

Funding from the Hariri Institute Graduate Student Fellowship (GSF) in 2023 supported early development of the project and helped Ruiz present his work at conferences throughout his doctoral training.

At DuckCon #5 in August 2024, Ruiz debuted dbverse and met the founders and developers of the database technology underlying the project. Feedback from industry and academic leaders helped shape the project’s further development. 

With the help of a travel award from the Sustainable Horizons Institute, Ruiz presented a second talk at the 2025 US Research Software Engineer Conference, highlighting ongoing development of dbverse and its potential to support collaboration among researchers using different programming languages.

Empowering Research Through Collaboration

Collaboration within the research team and contributions from the wider open-source software community played a crucial role in advancing dbverse, Ruiz said.

“dbverse is built on many open-source software packages, all of which rely on contributions from the open-source software community. Collaborations are an essential part of the research software lifecycle, and opportunities like the GSF enable trainees like me to pursue collaborations that may otherwise not have been possible.” 

Ruiz notes that some of his favorite collaborations were those with undergraduate students in the lab. Through Ruiz’s work with undergraduate student Timur Rizvanov in the Dries lab, the team twice secured funding from the Undergraduate Research Opportunities Program (UROP)  program to develop new applications of dbverse for scalable interactive visualizations. They have since released open-source software packages that demonstrate these capabilities and plan to release a preprint describing the work. 

Looking to the Future

“Looking ahead, what’s exciting is that the concepts behind the dbverse framework, initially written in the R statistical programming language, have the potential to be applied to other scientific programming languages like Python and Julia,” Ruiz said. “This may help us standardize the way we perform common scientific operations through a shared, scalable language in databases. This could lead to a lot of benefits like better reproducibility, but also perhaps increased collaboration between research communities that have historically been isolated from each other due to programming language barriers.”

What began as an effort to solve a computational problem in Ruiz’s own research has grown into an open-source framework with the potential to make large-scale scientific analyses more accessible. As Ruiz continues expanding dbverse to new applications, the Hariri Institute Graduate Student Fellowship’s impact extends beyond the project itself. It demonstrates the role the Hariri Institute can play in empowering junior researchers to pursue bold ideas that might otherwise remain out of reach.