Advancing Responsible AI: Naomi Saphra Joins CDS as Assistant Professor

Headshot of Naomi Saphra, Boston University Faculty of Computing & Data Sciences
Naomi Saphra

Boston University’s Faculty of Computing & Data Sciences is pleased to announce Naomi Saphra as an incoming Assistant Professor, starting in Fall 2026. Saphra joined BU following her previous role as Kempner Research Fellow at Harvard University, and brings a bold research agenda focused on deeply understanding language model training—spanning linguistics, interpretability, deep learning optimization, and behavioral evaluation. Her work seeks to decode how neural language models generalize, discover scientific insight, and behave in real-world settings.

“Naomi’s interdisciplinary vision—melding interpretability, language modeling, and AI for scientific understanding—adds a powerful new dimension to our faculty,” said Azer Bestavros, Associate Provost for Computing & Data Sciences. “She will help BU advance both the rigor and relevance of our AI research, especially at the edges where models meet real‐world complexity.”

Saphra earned her PhD from the University of Edinburgh, and has held positions at NYU, Google, and Facebook. Her appointment marks a strategic expansion of CDS’s core strengths—bringing in expertise that will complement ongoing work in AI explainability, scientific discovery, and training dynamics. Students and colleagues alike can look forward to engaging with her research across multiple domains, many of which intersect natural and social sciences, helping to shape BU’s role as a leader in responsible, interpretable AI.

“CDS is expanding rapidly, and language model analysis is one of the core areas it’s been hiring in. I’m excited to be part of a growing “supergroup” of language model analysis labs, and we’ve already begun planning joint meetings and collaborations,” Saphra said. “CDS also has a broad community encouraging collaboration between AI researchers and natural scientists or even philosophers, so I can’t wait to expand into more interdisciplinary work.”

In this Q&A, Saphra discusses her teaching philosophy, her research focus, her passion for roller derby, and more.


Q&A

How would you describe your teaching philosophy, and how does it reflect your work in AI and NLP?

I believe that when a student is confused or surprised, the resulting friction is a perfect learning opportunity. In AI, we know that the model learns the most from examples that it answers incorrectly. Humans are the same: The only way for us to learn is to constantly push up against the boundaries of our understanding. This is one problem with leaning too heavily on AI: without failure, we learn nothing.

Your research digs into language model training dynamics—why is understanding how models learn just as important as what they learn?

I look at it similarly to biology. The biologist Theodosius Dobzhansky famously said that “nothing in biology makes sense except in the light of evolution.” If you don’t understand how biological traits come about, you can’t know what the observed traits are for or how they relate to each other. Similarly, without considering the training process, there is very little we can really explain about its end product in a language model.

You’ve woven together linguistics, optimization, interpretability, and more. How do you see these fields intersecting at CDS?

CDS is full of people from all of these areas, who deal with their intersection. Just in the last couple of years, CDS has hired a number of people in language model analysis. Some come from cognitive interpretability, some from linguistics, and some from optimization. It’s always a pleasure to find yourself with the opportunity to work with some people who know more than I do about linguistics, some who know more about interpretability, and some who know more about optimization—and to help be the glue binding them together.

In addition to language modeling, you’ve started collaborating with natural and social scientists. Looking ahead to your time at CDS, what kinds of real-world scientific questions are you most excited to explore in collaboration with colleagues at BU?

Language has been a rich testbed for developing ideas about interpretability because it’s so learnable for models and because it’s easy for most humans to recognize patterns in language data. Areas like chemistry, animal communication, or astronomy are so much harder, but because we have methods connecting data structure with model structure in language, we have a starting point.

Bias and interpretability are big themes in your research. How do you think we can make AI systems more transparent—and fair—for everyone?

My top priority right now is to emphasize the scientific method in interpretability research: can we understand a model well enough to anticipate what its failure modes or edge cases might be. This is important for general model evaluation, but also for finding how it might misbehave dangerously or harmfully for humans who come to rely on it. Biases in a model might be subtly observed over only a few examples, but in aggregate and by inspecting the underlying reasons for its behavior, we can find how it might hurt individuals and society.

Beyond the Bio: You’re active in roller derby and also perform stand-up comedy—two high-energy, creative worlds. What do you enjoy most about each?

I mentioned above that friction and confusion are important opportunities for deep learning. That’s actually something I’ve learned from my hobbies! For example, in roller derby, a lot of leagues have a tradition of clapping when someone falls down to acknowledge that they’ve stretched against their limits and taken a risk.

In comedy, a lot of the transfer to teaching comes from workshops and classes I’ve attended on clown, not stand-up. Clown traditionally has a pedagogical approach that’s very emotionally involved. In clown, exercises often construct situations where mistakes are inevitable, and while we try as hard as we can to avoid failure—it’s not nearly as funny to fail on purpose—we can take advantage and lean into our mistakes, treating each error or moment of friction as an opportunity to deeply explore something true and funny.

Learning is a lot like comedy: as long as everything goes smoothly, you can memorize, or you can add little tools to your toolbox. But really deep understanding comes from overstretching and then exploring the resulting confusion.

By Maureen McCarthy