Confluent, Redpanda and Three Rivals Just Named What Comes After the Lakehouse
Data Engineering

Confluent, Redpanda and Three Rivals Just Named What Comes After the Lakehouse

Five streaming vendors that normally compete head to head formed a working group to define streamhouse architecture, a bet that live data infrastructure needs its own category before hyperscalers define it for them.

PublishedSeptember 24, 2026
Read time6 min read
Share

Five competitors, one working group

On September 15, Redpanda, Aiven, Confluent, StreamNative, and Ververica announced the Streamhouse Working Group, a joint effort to publish a vendor neutral definition of what the founders call real time AI data architecture. These five companies compete directly for the same streaming workloads day to day, which makes a joint announcement worth pausing on. Vendors rarely agree on shared terminology unless they see a bigger threat than each other on the horizon, and naming that threat plainly is the fastest way to understand why five rivals just stood on the same stage.

That bigger threat is almost certainly the hyperscalers and the largest lakehouse platforms, who have the market power to define categories unilaterally if independent streaming vendors do not coordinate first. By publishing a shared definition now, this group is trying to establish streamhouse as a distinct architectural category before Databricks, Snowflake, or a cloud provider absorbs the term into their own existing product marketing and sets the terms every smaller vendor then has to compete against.

The lakehouse versus streamhouse distinction

The working group's framing, laid out by Redpanda's Alexander Gallego, draws a specific line: lakehouses were built for understanding data, streamhouse architecture is built for acting on it. A lakehouse answers questions about what happened last quarter by querying a large, mostly historical dataset. A streamhouse is meant to continuously capture, transport, transform, govern, and serve the current state of a business so that production applications and AI agents can act on it immediately, not after a batch job completes.

That distinction matters more than it might sound. Most AI agent use cases enterprises are piloting today, fraud detection, inventory rebalancing, customer support triage, dynamic pricing, need decisions made on data that is minutes or seconds old, not data that was correct as of last night's ETL run. A lakehouse architecture, however well built, is structurally the wrong tool for that job, and the working group is betting enterprises will realize that gap faster than vendors can retrofit lakehouses to close it.

The defensive logic behind the alliance

The commercial logic is worth naming plainly. Confluent and Redpanda compete aggressively on Kafka compatible streaming, and both have spent the last two years pitching their own version of unifying streaming with lakehouse storage formats. A shared working group leaves that underlying rivalry fully intact while letting them jointly own the vocabulary that customers will use when evaluating both platforms, keeping that framing out of Snowflake's or Databricks' marketing teams' hands entirely.

There is precedent for this pattern. Open table formats like Apache Iceberg gained enterprise trust partly because multiple vendors backed a shared standard rather than each pushing a proprietary format. If the Streamhouse Working Group follows that playbook and produces genuinely interoperable specifications rather than just a shared glossary, it could meaningfully lower switching costs between these five vendors, which would be good for buyers even if it was not the founders' primary motivation.

What a real definition needs to include

A working group announcement is a starting point, not a specification, and the immediate test is whether these five companies can agree on concrete technical commitments beyond a shared blog post. Real interoperability would mean compatible approaches to state management, consistent semantics for exactly once processing across vendor boundaries, and shared governance primitives so that data lineage and access policy travel with the data regardless of which vendor's streaming engine touched it last, a much deeper commitment than agreeing on a name for the category and one that will take actual engineering resources away from each vendor's own competitive roadmap to deliver on schedule, at a moment when every one of these companies is also racing to ship its own AI features under separate competitive pressure from customers who want answers now, not after a standards process concludes.

That is a much harder problem than agreeing on a name. Streaming engines differ substantially in how they handle backpressure, state store consistency, and schema evolution, and those differences are exactly where vendor lock in currently lives today, quietly, inside implementation details most buyers never examine until a migration forces the question. If the working group's next deliverable is a marketing glossary rather than a technical interoperability spec, enterprise architects should treat the announcement as positioning rather than a genuine architectural shift worth changing a procurement decision over.

What CTOs should watch for next

For enterprise data leaders currently building real time AI pipelines, the practical takeaway is to watch whether this group produces actual technical artifacts, reference architectures, conformance tests, or open specifications, within the next two or three quarters. A working group that stalls at the press release stage tells you the underlying vendors were not willing to concede real technical ground to each other, which is useful information for procurement conversations either way.

In the meantime, the framing itself is useful regardless of the working group's ultimate output. If your organization is planning AI agent deployments that need to act on live business state rather than yesterday's snapshot, streamhouse versus lakehouse is now a legitimate axis to evaluate vendors on, and asking each streaming or lakehouse vendor directly how they position against this new terminology is a fast way to separate genuine architecture from repackaged marketing.

How this compares to past standards fights

The streaming ecosystem has been through this dynamic before with table formats, and that history is instructive for reading this announcement correctly. Apache Iceberg gained enterprise trust largely because Netflix, Apple, and later a coalition of vendors backed a genuinely open specification rather than a single company's proprietary format, and that broad backing is precisely what let Iceberg become a default choice across otherwise competing platforms within a relatively short number of years, faster than any single vendor could have driven adoption alone.

Kafka's own protocol tells a similar story, an open specification that multiple vendors, Confluent included, built commercial businesses around without any single company fully controlling its evolution or extracting disproportionate rent from everyone else building on top of it. The Streamhouse Working Group is explicitly trying to replicate that pattern at an earlier stage, before any single vendor's proprietary streamhouse implementation becomes the de facto standard other companies are forced to be compatible with on someone else's terms, years into a product cycle when switching costs are already locked in.

Tagged#news#data#data-engineering#databases#analytics#lakehouse#streaming#Streamhouse#Confluent#Redpanda#Aiven#StreamNative#Ververica#streaming-architecture