- Starts: 9:45 am on Tuesday, August 11, 2026
- Ends: 11:45 am on Tuesday, August 11, 2026
ECE PhD Thesis Defense: Christopher Liao
Title: Applications of Metric Learning to Image Retrieval and Multimodal Retrieval
Presenter: Christopher Liao
Advisor: Professor Brian Kulis
Chair: TBD
Committee: Professor Brian Kulis, Professor Venkatesh Saligrama, Professor Wei-Lun Chao, Professor Kayhan Batmanghelich
Google Scholar Link: https://scholar.google.com/citations?hl=en&user=iockppwAAAAJ
Advisor: Deep metric learning algorithms aim to learn a feature space where metric distances reflect the semantic relevance of inputs. In this thesis, we first study robust techniques to adapt pretrained encoders for image retrieval on domain-specific datasets. We then explore a parameter efficient way to adapt pretrained image-language encoders for text-to-image retrieval. Finally, we examine practical issues arising from deploying a dual encoder approach for crossmodal search, and mitigating strategies.
Firstly, we introduce a novel contextual loss function for image retrieval. Many prior loss functions focus on learning a correct ranking of training samples, but strongly overfit semantically inconsistent labels and require a large amount of data. To address these shortcomings, our loss implicitly enforces semantic consistency among neighbors while converging to the correct ranking.
Secondly, we focus on extremely parameter efficient methods (requiring only tens of parameters) for out-of-distribution few-shot adaptation of text-image encoders. Driven by the surging interest in contrastive image-language pretrained models (CLIP), a large body of research has emerged around zero-shot evaluation using GPT descriptors. Inspired by these prior works, we present two more flexible methods, named descriptor and word soups, which do not require an LLM at test time and can leverage training data to increase out-of-domain target accuracy.
Finally, we explore the practical issues of large-scale approximate nearest neighbor search (ANNS) arising from the modality gap phenomenon of CLIP models. Many follow-up studies of CLIP theoretically showed that dual encoder models suffer from the modality gap, where embeddings from the different encoders reside in disjoint subspaces. In contrast, industry standard ANNS systems using hashing or other quantization schemes assume that the nearest neighbor(s) to a query reside with high probability within the query's neighborhood. We show theoretically that crossmodal ANNS suffers from low recall due to the large distance between text queries and the image centroids used for coarse quantization. Accordingly, we propose paired k-means, a simple clustering algorithm that improves nearest neighbor recall by storing centroids in query space instead of image space, and extend this analysis to locality sensitive hashing (LSH).
- Location:
- PHO 339
