Skip to content

Details

Managing a growing technical community means answering hundreds of repeating questions across different channels every single day. In this session, we'll unpack the practical architecture behind the DataTalks.Club AI FAQ assistant, a production pipeline built to automate community support and deliver accurate answers directly where students ask them.

​We will walk through the entire system lifecycle, moving from raw community knowledge sources to a real-time, deployed Slack assistant. You’ll see the exact data engineering steps used to gather unstructured community inputs and turn them into a searchable vector index, followed by a live demonstration of the bot running end-to-end inside Slack.

We'll Cover

  • ​Student question ingestion and FAQ dataset contributions
  • ​Multi-source data extraction across Slack threads and YouTube transcripts
  • ​Vector database indexing and search retrieval workflows
  • ​RAG system architecture and LLM response generation
  • ​Slack API integration and automated bot deployment
  • ​End-to-end live demonstration inside the active Slack community


By the end of this session, you’ll understand how to build and deploy a production-ready, multi-source RAG pipeline that turns unstructured Slack threads and YouTube transcripts into a responsive, automated Slack assistant.

​We've also documented the full technical setup and step-by-step implementation in this article.

## ​About the Speaker

Alexey Grigorev is the Founder of DataTalks.Club and creator of the Zoomcamp series.

​Alexey is a software and ML engineer with over 10 years in engineering and 6+ years in machine learning. He has deployed large-scale ML systems at companies like OLX Group and Simplaex, authored several technical books, including Machine Learning Bookcamp, and is a Kaggle Master with a 1st place finish in the NIPS'17 Criteo Challenge.

**Join our Slack: https://datatalks.club/slack.html**

You may also like