Skip to content
stateless.co · Engineering notes from the request/response layer
statelessThe engineering desk

A publication about the machinery under everyday software: the contracts between services, the queries behind a page, and the failures that only show up in production.

—Archive

Major Retailer Cuts Peak Database Load 50% After Migrating to Event-Driven APIs with Change Data Capture

10 October 2026

In a bid to smooth peak holiday traffic, a major retailer overhauled its data architecture and slashed inbound database load by over 50% — by using event-driven APIs and change data capture to offload real-time processing.

The key: Instead of batching up changes and syncing overnight, the retailer built a new event-streaming backbone. Key operational databases update an event log as soon as a change is committed. An API layer subscribes to these event streams and reconstructs the state of the source, in real time, for downstream services.

Old-style ETL jobs that pulled full table duplicates are now a relic. Today, event collectors wrap the database and push live change logs into Kafka queues. API gateways subscribe to these queues. Distributed caches store alerts and popular result sets for hot reads. The worst-case data delay is minutes, not hours.

The practical outcome: Live inventory updates, sub-second response times on key pages, and no more pre-Christmas page timeouts.

Change Data Capture Moves from Batch to Real Time

The new data-flow architecture turns the retailer's omni-channel systems into a long-running conversation. Operational databases, front-end web services, and analytical stores no longer work disjointedly in batch syncs. Instead, they talk in real time, down the prevailing direction of change.

As requests cascade through the retailer's tiered APIs, the websites don't pull a full shopping cart repeatedly. Instead, API clients subscribe to change events. Event collectors listen to web requests and update their caches accordingly. Employee devices wait for searchable alerts rather than database refreshes. Kafka streams the relevant table updates as they occur.

The big change was feeding the event layer directly from operational sources. With CDC, the retailer decouples web transaction speed from database read speed. This represents a full object model that monitors itself as it changes.

How the Real-Time Pipeline Works

The retailer uses Apache Kafka as the event streaming platform, ingesting millions of customer interactions per second at peak. Kafka processes transaction logs and clickstreams as they happen. Each interaction enters an event queue within milliseconds.

From these native event streams, the retailer builds aggregated datasets. Kafka creates a transaction log that downstream services can resubscribe to. This forms the core data layer, collecting micro-events from the point of truth. Kafka organizes the event streams into a consistent pattern, ready for event-driven APIs.

Downstream services operate in a distributed topology. API decorators feed caching proxies. Process traces push change logs to event-translating services for quick ingestion. Each component functions as a real-time database with defined relationships.

Why the New System Drops Peak Database Load

Simple: The retailer no longer needs to pull every table for every request. Database load decreased because changes and queries became decoupled. When search requests hit the API caches, they don't cascade back to the primary database.

Even during high traffic, this system absorbs the pressure. Caching handles peak traffic while databases maintain steady operations. The retail databases build event logs, streaming data snapshots to consumers as changes occur. Distributed caches process search clicks while core databases remain efficient.

Retailers can typically reduce query load by 50% with three moves:

• Monitor event collection at the first tier and connect the access points

• Create specialized caches for high-demand paths and optimize queries

• Customize high-performance caches for each API endpoint

The retailer implemented these same strategies. Caching integrated with data updates led to subsecond response times and 99% uptime.

Specialist Retail Data Cases

The retailer's experience reflects broader trends. Event-driven database setups routinely reduce integration costs and improve data flow across retail and e-commerce.

Instacart uses distributed caching to handle hundreds of thousands of concurrent orders. Apache Kafka streams these orders with real-time checkout. The system maps demand signals to inform retail forecasting and inventory adjustments.

Zalando built a distributed in-memory cache around product snapshots and inventory data, enabling real-time pricing updates. Instead of polling legacy databases, Zalando tracks event updates and maintains a lightweight event view, providing API change alerts ahead of source system updates.

Macy's used event streaming technology to create logs from their Oracle databases, feeding events to analytics platforms. Change alerts reach downstream AI systems in minutes, enabling instant inventory updates and promotional adjustments.

What Retailers Can Do

For retailers considering real-time data layers, the key lesson is structural decoupling: API-driven events and real-time CDC reduce database load while improving service levels. Data consistency can be maintained through event-driven validation and lightweight integrity checks.

Retailers transitioning to event-driven systems should first implement CDC access controls and set up queue listeners. The next step involves customizing caches for each API based on traffic patterns. Integrate data flow with API events through capture points, aggregate downstream events, and optimize caching strategies.

Crucially, this doesn't require building a complete microservices architecture overnight. Systems can be integrated gradually, injecting event-driven functionality one component at a time. Focus first on high-impact use cases where event-driven approaches offer clear advantages. As individual systems demonstrate success, expand event-driven patterns throughout the architecture while maintaining system coherence.