Beyond the Guardrail: Auditing Hidden Modes of Failure in Generative AI
Details
Large language models and emerging multimodal systems increasingly incorporate safeguards against harmful, biased, and factually unreliable outputs. But how robust are these safeguards beyond conventional, static evaluations? In this talk, I present three connected lines of research that expose vulnerabilities through interaction, self-consistency, and multimodal behavior.
First, I introduce the toxicity rabbit hole, an iterative framework that repeatedly asks a model to make its own previous generation more toxic. Across multiple LLMs and more than a thousand identity groups, this stress test reveals antisemitism, racism, misogyny, and other forms of identity-directed harm that guardrails can fail to stop. Second, I present HAUNT, a dynamic, contamination-resistant framework for evaluating factuality hallucinations under conversational nudges by asking models to generate, verify, and then confront their own falsehoods. Finally, I examine multimodal systems and uncover a striking disconnect between what models refuse to infer and what they are willing to construct.
These three studies suggest that A.I. safety evaluation is a moving target. As models become more capable, interactive, and multimodal, failure modes can emerge in ways that static benchmarks and surface-level refusals fail to capture.
Ashique KhudaBukhsh
https://www.rit.edu/directory/axkvse-ashique-khudabukhsh
Assistant Professor
Department of Software Engineering
Rochester Institute of Technology
About the Speaker:
Ashique KhudaBukhsh is an Assistant Professor in the Software Engineering Department and is affiliated with the ESL Global Cybersecurity Institute, where he directs the Social Insight Lab. Dr. KhudaBukhsh's current research lies at the intersection of natural language processing (NLP) and AI for Social Impact, as applied to: (i) globally important events arising in linguistically diverse regions, requiring methods to tackle practical challenges involving multilingual, noisy, social media texts; (ii) polarization in the context of the current US political crisis; and (iii) auditing generative AI systems and platforms for unintended harms. In addition to having his research accepted at top artificial intelligence conferences and journals, his work has also received widespread international media attention, including coverage from The New York Times, BBC, Wired, Times of India, The Indian Express, The Daily Mail, VentureBeat, and Digital Trends. A detailed resume can be found here.
With three published collections of poems to his credit, an experience of directing music at a New York theater play, occasional dabbling at journalism and column-writing, and a recent success at swimming 50 meters underwater (a navy seal requirement), Ashique enjoys his multiple distractions that keep him away from work.
* * * * * * * * * * * * * * * * * * * * * * * * * * * *
This is a talk with audience Q&A presented by the University of Toronto's Centre for Ethics that is free to attend and open to the public. Free refreshments will be provided at the event. The talk will also be streamed online with live chat here.
About the Centre for Ethics (http://ethics.utoronto.ca):
The Centre for Ethics is an interdisciplinary centre aimed at advancing research and teaching in the field of ethics, broadly defined. The Centre seeks to bring together the theoretical and practical knowledge of diverse scholars, students, public servants and social leaders in order to increase understanding of the ethical dimensions of individual, social, and political life.
In pursuit of its interdisciplinary mission, the Centre fosters lines of inquiry such as (1) foundations of ethics, which encompasses the history of ethics and core concepts in the philosophical study of ethics; (2) ethics in action, which relates theory to practice in key domains of social life, including bioethics, business ethics, and ethics in the public sphere; and (3) ethics in translation, which draws upon the rich multiculturalism of the City of Toronto and addresses the ethics of multicultural societies, ethical discourse across religious and cultural boundaries, and the ethics of international society.
The Ethics of A.I. Lab at the Centre For Ethics recently appeared on a list of 10 organizations leading the way in ethical A.I.: https://ocean.sagepub.com/blog/10-organizations-leading-the-way-in-ethical-ai
