

About us
Charlottesville Data Science is a community for data scientists, AI engineers, machine learning practitioners, and all professionals, students, researchers, and enthusiasts working with data in Charlottesville and Central Virginia. Charlottesville is a growing data and technology hub, with the University of Virginia, including the UVA School of Data Science, established companies like S&P Global, Elder Research, and GA-Intelligence, and a dynamic ecosystem of homegrown startups. Let's connect these dots to share ideas, learn from each other, and grow the local tech community.
Our members include researchers and tech professionals with decades of experience, novices who have yet to write their first line of code, and everyone in between. If you're interested in learning more about cutting-edge work happening with data science, AI, machine learning, and related technologies in Charlottesville, you're in the right place, and you'll find a welcoming, supportive community of like-minded folks.
In addition to our Meetup group, please consider joining our Substack newsletter. It's the primary place we share community announcements and event updates with our members:
Have an idea for a future Charlottesville Data Science event? Fill out our Call for Proposals form and a member of our organizing team will get back to you!
Need to get in touch with the Charlottesville Data Science organizing team? You can reach us at organizers@cvilleds.org.
Upcoming events
1

Accuracy Isn’t Everything: Good Models are Lurking Everywhere
School of Data Science, 1919 Ivy Rd, Charlottesville, VA 22903, USA, Charlottesville, VA, USPlease join Charlottesville Data Science for the talk Accuracy Isn’t Everything: Good Models are Lurking Everywhere from Evzenie Coupkova, a mathematics PhD from Purdue. In this talk, instead of focusing on one “best” model, Evzenie will explore the broad landscape of possible models, consider what it means for a machine learning model to be “good” for a given problem, and examine the trade-offs between predictive accuracy and other desirable characteristics, like generalization and interpretability.
We'll be gathering in person at the UVA School of Data Science on the evening of Thursday, October 8. We look forward to seeing you there!
→ We're moving our events calendar to Luma! Please head over there to RSVP, and consider subscribing to our Luma calendar! ←
About the talk
Data scientists often focus on a single "best" model — the one with the highest accuracy. But what if we looked at all possible models? Would many others achieve similar accuracy and how does this depend on the dataset?
This talk sits at the intersection of Cynthia Rudin's work on interpretability and Sanjeev Arora's research on optimization and generalization of neural networks.
We start with simple, intuitive cases: linear models applied to Gaussian mixtures. We explore where the high- and low-accuracy models are located in parameter space, and how the proportion of "good" models change as the distance between class means grows.
We then turn to randomly initialized neural networks, asking the same question: what fraction of them are "good" classifiers? Using Sanjeev Arora's matrix-based approach to measuring dataset complexity, we examine how that complexity affects the fraction of "good" networks. The results of experiments on real-life datasets with varying sample sizes and dimensionality will be presented.
Underlying both parts of the talk is a general principle that holds for any model-data pair: if you know "good" models are abundant, you can improve generalization without giving up much accuracy. You can also choose a model that better fits your needs, one that is easily interpretable and suitable for high-stakes decision environments (such as healthcare or the judicial system), one that's sparse, or one that satisfies constraints imposed by a human expert.
If you want to shift your perspective and explore the full landscape of models, this talk is for you.
About the speaker
Evzenie Coupkova holds a PhD in Mathematics from Purdue University and specializes in statistical learning theory, machine learning, and high-dimensional data analysis. Her research explores the dimensionality reduction, proportion of high-accuracy models and generalizability in connection with randomness and dataset labels. She has collaborated with industry partners and presented her work at conferences including IPAM-UCLA and SIAM MDS, blending mathematical rigor with practical data science applications.
5 attendees
Past events
44


