Skip to content

Details

This meetup session will focus on ACL and SIGIR 2026 and feature three speakers from University of Amsterdam: Yibin Lei, Zahra Abbasiantaeb and Dylan Jia-Huei Ju. They will present their papers accepted at ACL and SIGIR 2026.

Location: Science Park 904, Room C3.161
Date: Friday, June 26
Time: 16:00-17:00
Zoom link: https://uva-live.zoom.us/j/65011610507

Details below:
Speaker #1: Yibin Lei

Title: Making Large Language Models Efficient Dense Retrievers
Abstract: Recent work has shown that directly fine-tuning large language models (LLMs) for dense retrieval yields strong performance, but their substantial parameter counts make them computationally inefficient. While prior studies have revealed significant layer redundancy in LLMs for generative tasks, it remains unclear whether similar redundancy exists when these models are adapted for retrieval tasks, which require encoding entire sequences into fixed representations rather than generating tokens iteratively. To this end, we conduct a comprehensive analysis of layer redundancy in LLM-based dense retrievers. We find that, in contrast to generative settings, MLP layers are substantially more prunable, while attention layers remain critical for semantic aggregation. Building on this insight, we propose EffiR, a framework for developing efficient retrievers that performs large-scale MLP compression through a coarse-to-fine strategy (coarse-grained depth reduction followed by fine-grained width reduction), combined with retrieval-specific fine-tuning. Across diverse BEIR datasets and LLM backbones, EffiR achieves substantial reductions in model size and inference cost while preserving the performance of full-size models.
Bio: Yibin Lei is a PhD student at the University of Amsterdam, where he works on information retrieval in multilingual and low-resource scenarios.He has co-authored multiple papers in leading conferences like ACL, EMNLP, and ICLR. Additionally, he has served as a committee member/reviewer for SIGIR, ACL ARR, NeurIPS and ICLR.

Speaker #2: Zahra Abbasiantaeb

Title: Conversational Gold: Evaluating Personalized Conversational Search System Using Gold Nuggets
Abstract: As personalized conversational search systems increasingly rely on Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG), assessing the quality, relevance, and factual alignment of long-form responses remains a critical challenge. Traditional evaluation metrics often fall short when judging complex, multi-turn conversational outputs. In this presentation, we introduce CONE-RAG, a novel nugget-based evaluation paradigm designed to bridge the gap between retrieval effectiveness and generation quality.
Moving beyond high-level surface metrics, CONE-RAG establishes a structured framework for long-form answer generation evaluation by centering on "gold nuggets"—concise, atomic units of essential information extracted directly from relevant source passages. We outline how the framework operationalizes automatic response evaluation through automated nugget extraction and semantic matching, explicitly linking generation quality back to retrieval performance. Attendees will gain insights into how this paradigm enables a deeper, more granular diagnostic assessment of RAG systems, establishing a rigorous new standard for evaluating personalized, context-aware conversational search.
Bio: Zahra Abbasiantaeb is a Ph.D. student at the University of Amsterdam's Information Retrieval Lab (IRLab). Her research centers on improving conversational information seeking systems through advanced evaluation, query understanding, and personalized search framework design. Over the past three years, she has co-organized the TREC iKAT track, aiming to foster standard test collections and robust benchmarks for conversational systems.

Speaker #3: Dylan Jia-Huei Ju

Title: Search for Coverage: Learning Coverage-Aware Retrieval with Augmented Sub-Question Answerability

Abstract: Long-form RAG introduces a new challenge of coverage-based ranking, where retrieval must ensure the inclusion of comprehensive and diverse relevant nuggets (i.e., facts) that can be synthesized into a thorough output. In this talk, I will present CoveR, a retrieval method optimized for coverage-aware retrieval. CoveR is trained with coverage contrastive and distillation objectives, designed to produce query representations that capture broader aspects of an information need. To support model training, we construct SCOPE, a dataset of 90K training pairs built with synthetic coverage signals derived from sub-question answerability judgments. Our experiments demonstrate that CoveR effectively balances relevance and coverage. We hope this study can facilitate further research on coverage-based ranking and more dynamic retrieval architectures for RAG.

Bio: Dylan Jia-Huei Ju is a PhD student at the IRLab, University of Amsterdam, specializing in Information Retrieval and NLP with a focus on long-form Retrieval-Augmented Generation. His work targets the interaction between retrieval and generation, including retrieval evaluation and reranking for nugget coverage.

Counter: SEA Talks #308, #309 and #310.

Related topics

Events in Amsterdam, NL
Artificial Intelligence
Natural Language Processing
Information Retrieval
Search, Information Retrieval

You may also like