Webinar: Stop Copying Data for Vector Search
21 attendees from 10 groups hosting
Details
To attend, register here.
The data lake is supposed to be where all your data lives. Yet vector search has traditionally required copying embeddings into a separate vector database—adding duplicate storage, synchronization pipelines, and another system to operate. This webinar explores how that architecture is changing.
Tune in for a technical walkthrough of how Apache Pinot brings vector similarity search directly to Apache Iceberg and Delta Lake. We'll cover the evolution from local Pinot tables to tiered storage on Amazon S3 and finally to lake-native vector search using External Tables, showing how approximate nearest neighbor (ANN) search can run over open table formats without moving your data.
In this technical discussion, you'll learn hot to:
- Run vector similarity search directly on Apache Iceberg and Delta Lake using Apache Pinot and External Tables.
- Understand how HNSW enables fast ANN search and what changes are required to make it work over object storage.
- Combine semantic search with SQL filters in a single query over the same data.
- Evaluate a lake-native architecture for AI retrieval that keeps one copy of your data while simplifying search infrastructure.
