Instacart is the leading grocery technology company in North America, partnering with over 1,800 retail banners across more than 100,000 stores. Behind the scenes is a modern data stack built on AWS, Kafka, Flink, Databricks, and Snowflake, designed to handle billions of events, power real-time decision making, and support analytics across the entire business.
Metrics
~83 million orders in Q3 2025, over 1.5 billion lifetime orders.
Over 2 trillion events processed per year across data pipelines.
Snowflake warehouse holds over 250,000 tables and views.
Nebula handles 30 billion events and manages ~40PB in the Lakehouse.
Mode dashboards launched from 0 to 8,000 in six months.
Content is based on multiple sources including Instacart Blog and other public articles etc. You will find references to dive deep as you read.
Platform
AWS
AWS Instacart runs its entire data infrastructure on AWS, with Flink workloads on Amazon EKS and Karpenter handling node autoscaling for complex resource-isolation requirements across hundreds of concurrent pipelines.
Messaging & Ingestion
Kafka
Kafka is the central nervous system for event streaming, connecting every major layer of the stack. Instacart handles over two trillion events a year through their data pipelines.
Debezium
Instacart continuously ingests data across thousands of tables from production systems to Snowflake using a system built on Debezium and Kafka. Debezium monitors database write-ahead logs, publishes row-level changes into Kafka topics.
Processing
Spark
The ad org runs an internal service called Nebula on Databricks, using Spark for both structured streaming and batch processing. Nebula handles 30 billion events and manages ~40PB of data stored in the Lakehouse.
Flink
Flink is Instacart’s core streaming computation engine, used for real-time event routing, decision-making, data augmentation, ML feature generation, and OLAP ingestion. Originally on AWS EMR, they migrated to Kubernetes using the Flink Operator, saving 50 weeks of development effort, cutting infra costs by 50%, and reducing critical alerts from 30 to 0.
📖 Related Reading: Building a Flink Self-Serve Platform on Kubernetes at Scale
Orchestrator
Airflow
Instacart previously ran at least a dozen separate Airflow instances across teams. They consolidated into a single setup, standardizing how pipelines are scheduled and reducing the operational overhead of maintaining fragmented orchestration systems.
📖 Read More: The Next Era of Data at Instacart
Warehouse & Transformation
Snowflake
Snowflake is the primary analytics warehouse, holding over 250,000 tables and views in a single database. Snowflake ingests data from multiple sources including real time events, production databases etc.
DBT
DBT is Instacart's standard transformation tool for Snowflake, adopted in 2022 to replace multiple fragmented tools. It integrates seamlessly with Airflow for scheduling.
Read More: Adopting dbt as the Data Transformation Tool at Instacart
Lakehouse
Delta Lake
Delta Lake is the storage format powering Nebula's Lakehouse. To allow querying outside of Databricks, Instacart leverages UniForm, which automatically generates Iceberg metadata on top of Delta tables without creating a second copy of the data.
Catalog & Governance
Amundsen
Instacart uses Amundsen as their data catalog, certifying datasets as bronze/silver/gold to help consumers select the right data. Instacart It indexes Snowflake assets; tables, lineage, ownership, etc.
Immuta
Instacart needed governance spanning Snowflake, Databricks, and AWS without managing native controls per platform. Their previous RBAC approach required users to adopt different “identities” across platforms. Immuta decouples policy authoring from enforcement, policies are written once in plain language and applied uniformly across all platforms, with a single audit view.
📖 Related Reading: How Instacart Streamlined Its Approach to Data Policy Authoring
Dashboard
Mode
Mode replaced Instacart's previous BI tools; within six months they launched over 8,000 dashboards, reduced failed dashboard runs by 7x, and cut Snowflake costs by 3x. Instacart Mode sits at the end of the Airflow → DBT → Snowflake pipeline, surfacing curated analytics to stakeholders across the company.
📖 Read More: The Next Era of Data at Instacart
Related Content:
💬 Overall, Instacart blends AWS with open-source and managed services to power real-time event streaming, petabyte-scale analytics, and self-serve data across its grocery technology platform.







