Skip to content

Details

Your Training Data Has a Lawyer Problem (And Your Lawyer Has a Data Science Problem)

Abstract:

Every data scientist has assembled a dataset by scraping, licensing, purchasing, or quietly reusing something that was already sitting in the warehouse. Most of the time nobody asks where it came from — until a customer, a regulator, or opposing counsel does.
This talk is a product lawyer’s field guide to data provenance, delivered for people who actually build models. We’ll walk through the questions Legal asks and why — what makes a dataset licensable vs. merely available, why “it’s public” is the most expensive sentence in AI, how scraped data, user-generated content, and vendor-supplied corpora carry fundamentally different risk profiles, and what a data scientist can do at collection time that saves a year of pain at deployment time. We’ll cover real patterns from AI supplier reviews and contract negotiations — the retention clauses, the “we may use your inputs to improve our services” language, and the questions nobody asks the vendor until it’s too late.
No legal background required. Bring your messiest dataset question.

Sponsored by LBMC

As always free food and drinks. Come at 6PM to socialize and 6:30pm is the presentations.

Related topics

Events in Nashville, TN
Artificial Intelligence
Data Analytics
Data Science
Predictive Analytics
Data Management

You may also like