SEA: October - Multimodal Retrieval

Name: SEA: October - Multimodal Retrieval
Start: 2025-10-31T16:00:00+01:00
End: 2025-10-31T17:00:00+01:00
Location: Lab 42

Hosted by Maarten de R. and Yubao T.

Meet the group

SEA: Search Engines Amsterdam

No reviews yet

Details

📍 Location: Room L3.36, Lab42, Science Park Amsterdam
💻 Zoom link: https://uva-live.zoom.us/j/67649187004

We’re excited to announce our upcoming SEA Talk on Multimodal Retrieval, featuring two speakers:

Kishan Parshotam – Flume AI

Title: The Unindexed Trillion Dollars Industry – Why search engines overlook this information, and how Flume is tackling it
Abstract: To be announced
Bio: Kishan Parshotam is the CTO and Co-Founder of Flume, where he is helping build next-generation search for unindexed industries. Previously, he led AI at AltScore, helping the fintech secure $3.5M in funding, and co-founded Just A.I., which developed scalable machine-learning tools for smallholder farmers. He also worked on computer vision research at Prosus Group, contributing to image search optimization and publishing at CVPR. Kishan holds a Master’s in Artificial Intelligence from the University of Amsterdam (2020).

⸻

Jingfen Qiao – University of Amsterdam
Title: Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction
Abstract: Visual Document Retrieval (VDR) is an emerging research area that focuses on encoding and retrieving document images directly, bypassing the dependence on Optical Character Recognition (OCR) for document search. A recent advance in VDR was introduced by ColPali, which significantly improved retrieval effectiveness through a late interaction mechanism. ColPali’s approach demonstrated substantial performance gains over existing baselines on an established benchmark.

In this study, we investigate the reproducibility and replicability of VDR methods with and without late interaction mechanisms by systematically evaluating their performance across multiple pre-trained vision–language models. Our findings confirm that late interaction yields considerable improvements in retrieval effectiveness; however, it also introduces computational inefficiencies during inference. We further examine the adaptability of VDR models to textual inputs and assess their robustness across text-intensive datasets when scaling the indexing mechanism. Finally, we explore how query-patch matching contributes to VDR performance, finding that query tokens tend to match visually similar or contextually related patches rather than exact counterparts.

Bio: Jingfen is a fourth-year PhD student at the IRLab, University of Amsterdam. Her research focuses on developing language models for dense and sparse retrieval, with a particular interest in multimodal document retrieval.

Counter: SEA Talks #291 and #292.

Events in Amsterdam, NL

Information Architecture

Science

Technology

Information Retrieval

SEA: October - Multimodal Retrieval

SEA: Search Engines Amsterdam

Details

Members are also interested in