Über uns
Regensburg is full of data science expertise, both in industry and in academia. Our aim is to bring together people who share an interest in this area and offer an environment for networking in an informal setting. We continue to have speakers with a range of backgrounds offering insights into a wide spectrum of data science ranging from enterprise search to music recommendation, from automatic fact-checking to avoiding harms and biases, from generative approaches to automatic question-answering. And that is not even everything. Other topics include large language models, industry use cases of natural language processing and the list goes on and on ... We have speakers from industry (e.g. Bloomberg, Netflix, Amazon, Spotify, Deloitte ...) and universities (CMU, Queen Mary, Essex, Regensburg ...). Want to present? Drop us a message. We also travel around having stopped in Berlin, Leipzig and Tampere so far. Want us to stop at your location? Let's discuss ...
Kommende Veranstaltungen
1

Similarity, Reasoning & Fairness in Legal NLP (a Canada Special)
University of Regensburg, Universitätsstraße 31, Regensburg, BY, DEDear all,
Summer is over, back to business! Ready? Here we go ...
We are very honoured to have two special guests all the way from Canada who will pop by in Regensburg just to talk to YOU! What's not to like?
Both Adam and Yiran have substantial insights and experience when it comes to processing natural language texts in the legal domain. This is based on academic research as well as downstream industry applications.
If you want to hear about where LLMs are heading, what the pitfalls and challenges are in practical applications such as legal NLP, what the differences are between academic papers and the real world, or if you just want to meet some nice people and chat about AI, data science and NLP, then why not come and join us Friday next week? Sign up and we will be waiting for you.
Looking forward to seeing you all,
Udo, David, Gregor, Samy and Markus (plus the rest of the team)Speaker 1:
Adam Roegiest (VP, Research & Technology at Zuva)Title:
Aboutness Is Not Equivalence: The Limits of Embedding SimilarityAbstract:
The cosine similarity of text embeddings is often framed as measuring the "semantic" similarity of the text fragments, reflecting their origins in distributional semantics. This shorthand is often reasonable among technical practitioners who understand how embeddings are calculated and the limitations that entails. When terms like "semantic similarity" are used in front of non-experts, this can encourage misleading intuition about what the measure captures. For example, a legal practitioner might reasonably expect two "semantically similar" legal clauses to impose the same rights, requirements, and obligations, or very nearly so, but is this what cosine similarity actually measures?In this talk, we show that such an assumption does not hold, at least for contractual language, and that care is warranted when calculating domain-specific measures of similarity. In a first study, we compare a reference clause with two variants: a minimally edited version that negates/reverses its legal meaning and a substantially reworded version that preserves its intent. Across embedding models, we see that the negated clauses are consistently more "semantically" similar than intent-preserving rewordings. In a second study, we explore whether explicit reasoning over clause obligations can better approximate lawyer's assessment of clause similarity than do embeddings. While embeddings remain competitive in this more naturalistic setting, prompt-based methods outperform them in most settings and provide interpretable explanations, as opposed to the opaque scores of embeddings. These results provide a broader lesson extending beyond the legal domain. Embeddings are effective at identifying pieces of text that are about the same thing, but domain-specific similarity will likely rely on which differences matter to the end user.
---
Speaker 2:
Yiran Hu (PhD student, University of Waterloo)Title:
Evaluation of LLMs for legal applicationsAbstract:
When large language models (LLMs) are deployed in domain-specific tasks, it is essential to evaluate their trustworthiness. In the legal domain, robustness and fairness are two particularly critical aspects.In this talk, I will first discuss about some existing ways for integrating legal knowledge into LLMs, then I will present two of my recent papers: “J&H: Evaluating the Robustness of Large Language Models under Knowledge-Injection Attacks in the Legal Domain” and “LLMs on Trial: Evaluating Judicial Fairness for Large Language Models.” The first paper investigates whether LLMs rely on genuine legal domain knowledge when performing legal reasoning, particularly under knowledge-injection attacks. The second paper adopts a counterfactual evaluation framework grounded in legal theories of judicial fairness to assess fairness-related behaviors of LLMs in judicial decision-making scenarios.
21 Teilnehmer
Vergangene Veranstaltungen
38


