Skip to content

Details

60-minute walkthrough of how engineering teams build production data pipelines on AWS in a single day. We cover architecture, live demo, and what your team needs to run this independently.

Full lab + screenshots: https://becloudready.com/workshops

Learn how to build an AWS data lake from scratch, live, using Amazon S3, AWS Glue, and Amazon Athena. This is a free, hands-on data engineering workshop, not a slide deck: you'll get sandbox AWS credentials and build a real, working data lake pipeline yourself, step by step, in 60 minutes.

We'll take raw CSV data, catalog it with an AWS Glue Crawler, query it with Amazon Athena SQL, then build a Glue ETL job that transforms it into partitioned Parquet for faster, cheaper queries. You'll see the exact cost difference between querying CSV and Parquet on the same data, measured live.

What you'll learn:

  • How to build an AWS data lake pipeline from S3 to Athena, hands-on
  • What an AWS Glue Crawler does and how it catalogs data without moving it
  • How to write and run a Glue ETL job (PySpark) to convert CSV to Parquet
  • Why Parquet and partitioning dramatically cut Amazon Athena query costs
  • - The raw → catalog → transform → catalog → query pattern used in real-world AWS data platforms

Who should attend:
Data engineers, analytics engineers, cloud engineers, and engineering managers learning AWS data engineering, Glue, Athena, or data lake architecture. No prior Glue/Athena experience needed.

Hosted by BeCloudReady (Databricks Registered Partner) and TorontoAI (10,000+ member tech community).

Related topics

New Career
Data Analytics
Data Science using Python
Machine Learning with Python
Software Engineering

You may also like