Skip to content

About us

The Real-Time Analytics meetup covers a range of topics around building Real Time Analytics systems; including use cases, technical deep dives, and best practices. 
Interested in speaking, organizing, or volunteering? Contact community@startree.ai

This meetup is organized by the founders of StarTree and original creators of Apache Pinot:
Apache Pinot is a realtime distributed OLAP datastore,  used to deliver scalable real time analytics with low latency. It can ingest data from batch data sources (S3, HDFS, Azure Data Lake, Google Cloud Storage) as well as streaming sources (such as Kafka). Pinot is used extensively at LinkedIn and Uber to power many analytical applications such as Who Viewed My Profile, Ad Analytics, Talent Analytics, Uber Eats and many more serving 200k+ queries per second while ingesting 1Million+ events per second.

Resources
> • What is Apache Pinot? https://www.startree.ai/what-is-apache-pinot
> • Launching At LinkedIn: The Story of Apache Pinot: https://www.startree.ai/blog/launching-at-linkedin-the-story-of-apache-pinot
> • For more info on Apache Pinot go to dev.startree.ai
> •Our community is active on slack! To join our slack, go to stree.ai/slack

Upcoming events

1

See all
  • Network event
    Webinar: Optimizing Data Latency and Query Latency in Open Table Formats

    Webinar: Optimizing Data Latency and Query Latency in Open Table Formats

    ·
    Online
    Online
    23 attendees from 10 groups

    To attend, register here.

    Data lakes came with an implicit trade. You got cheap storage and open formats, and in exchange you accepted that the data would be hours old and the queries would take minutes. Low cost, high latency. That trade is no longer necessary. Streaming data continuously into open table formats and running indexed, page-level query execution over those same files turns the lake from a batch archive into a near-real-time serving layer. Low cost and low latency, on one copy of the data, in object storage.

    This session walks the path an event takes from producer to end user. Confluent Tableflow represents Kafka topics and their schemas directly as Apache Iceberg and Delta Lake tables, with schematization, schema evolution, and catalog publishing handled for you, so there is no separate ingestion stack to build and no batch window to wait on. StarTree, built on Apache Pinot, then queries those same open tables with page-level precision, using indexes and pruning to touch a small fraction of the Parquet files a scan-based engine would read. The result is a single copy of governed data that serves both exploratory analysis and the sub-second, high-concurrency queries that user-facing applications and AI agents require.

    We will cover the architecture, show it running on a live stream, and share the latency numbers behind each stage.

    What You Will Learn

    • How to go from Kafka topic to queryable Iceberg table without custom pipelines
    • What page-level indexing changes about the economics of querying open table formats
    • Which application and AI use cases this architecture enables

    Who Should Attend

    • Data platform leaders and senior data engineers responsible for Kafka, Iceberg, Delta Lake, or analytical serving
    • Architects evaluating how to serve low-latency queries without duplicating lakehouse data
    • Application and AI platform teams that need fresh governed data under high concurrency
    • Photo of the user
    1 attendee from this group

Group links

Organizers

Photo of the user StarTree
Badge for StarTree
StarTree

Super Organizer