Aisha Amaira has spent over a decade at the intersection of marketing and deep-tech infrastructure, helping global brands navigate the increasingly complex maze of CRM and customer data platforms. As a seasoned MarTech expert, she has witnessed the industry’s shift from rigid, all-in-one suites to the modern era of composable stacks that prioritize flexibility and data ownership. Today, she shares her insights on a pivotal shift in the landscape: the move toward truly open customer data infrastructure that bypasses proprietary warehouse limitations. Our conversation explores the evolution of data sovereignty, the technical mechanics of open-source compute, and why the next phase of enterprise data management relies on giving power back to the customer’s own lakehouse.
The discussion focuses on the transition from traditional composable architectures to open-source alternatives that utilize Apache Spark and Iceberg. We examine how European enterprises are navigating strict regulatory environments like DORA and the CLOUD Act, the financial impact of eliminating “compute taxes,” and the strategic importance of maintaining a decentralized data catalog.
Traditional composable CDPs were once hailed as the ultimate solution for flexibility, yet they often rely on proprietary cloud warehouses. How does this dependency create hidden bottlenecks for modern enterprises?
For years, the industry believed that simply moving away from “packaged” solutions was enough, but we realized that we had only traded one form of lock-in for another. When a CDP keeps data in a customer’s warehouse but still requires a specific proprietary cloud runtime to perform audience computation, it effectively traps the enterprise in a metered consumption model. You feel the weight of this every time you run a complex segment and see those “per-credit” compute taxes eat into your budget without mercy. This architecture forces brands to rent a foreign cloud’s brain to process their own information, which makes it nearly impossible to satisfy the hard data-residency requirements we see today. It creates a tension where marketing teams want to innovate with real-time data, but the CFO is constantly eyeing the rising costs of the underlying proprietary warehouse credits.
The concept of “Open CDP” introduces a stack that is open at every layer, from compute to storage. Could you explain the practical impact of using Apache Spark and Iceberg tables directly on a customer’s own lakehouse?
This is a fundamental shift in how we think about the “physical” location of data processing. By using open-source Apache Spark as the engine and open Apache Iceberg for storage, we are finally allowing the data to live as a first-class citizen in the customer’s own environment. Imagine your identity graphs, event streams, and segment memberships sitting as readable tables in your own object storage, accessible by any tool you choose, rather than being hidden behind a vendor’s proprietary API. This setup removes the need to copy data into a vendor’s cloud, which eliminates both the security risk and the latency of data movement. It provides a level of transparency where you can see exactly how identity resolution is happening on bare-metal compute, giving technical teams a sense of control that was previously lost in the “black box” of traditional SaaS.
Data sovereignty is no longer just a buzzword but a legal necessity for many. How does a “sovereign by design” architecture help organizations face challenges like DORA or the CLOUD Act?
For a long time, sovereignty was treated as a checkbox or an add-on, but in the current regulatory climate, that simply doesn’t hold up under the scrutiny of a regulator. When your data resides in your own storage and the compute is operated in a known, specific jurisdiction without traversing a foreign vendor’s cloud, you are building a defensive moat that is evidenced by the architecture itself. This is particularly vital for European organizations that must navigate the complexities of the CLOUD Act, where the risk of foreign government access is a constant concern. By keeping the catalog—whether it’s Databricks Unity Catalog or Google Cloud—under the customer’s governance, the enterprise ensures that they, and only they, hold the keys to the kingdom. It turns compliance from a headache into a tangible competitive advantage because you can prove to your customers that their PII never leaves your direct supervision.
High-volume B2C sectors like telecommunications and financial services often struggle with the sheer scale of audience computation. Why is the “Open CDP” model particularly effective for these cost-sensitive industries?
In industries like retail or telco, where you are dealing with millions of event signals every single hour, the traditional “metered” consumption model of cloud warehouses becomes a financial black hole. These organizations need continuous audience computation to stay relevant, but they can’t afford to be penalized for their own success as their data volume grows. By moving to an open compute model on the customer’s own lakehouse, these high-volume players can predict their total cost of ownership with much greater accuracy. They are no longer paying a premium to a third-party vendor just to “access” the processing power they already have at their disposal. We are seeing these firms regain the ability to run massive identity resolution tasks across fragmented touchpoints without the fear of a massive, unexpected bill at the end of the month.
Projjol Banerjea, the founder of Zeotap, has mentioned that “openness becomes the lock-in story, told in reverse.” How does the ability for a customer to leave a platform at any time actually create a stronger, more sustainable partnership?
It sounds counterintuitive, but the most secure relationship is one where the door is always unlocked. When a platform is built on open substrates—where the engine, the storage, and the catalog are all replaceable—the vendor is forced to provide constant, undeniable value every single day to earn their place in the stack. Customers stay because the CDP intelligence, the journey orchestration, and the consent governance are top-tier, not because their data is held hostage in a proprietary format. This transparency builds a deep level of trust; the enterprise knows that if their strategy shifts or a better engine emerges between 2026 and 2028, they aren’t stuck in a five-year legacy contract with no exit. It shifts the power dynamic entirely, making the vendor a partner in the brand’s growth rather than a gatekeeper to their own customer insights.
What is your forecast for the evolution of customer data infrastructure over the next few years?
I believe we are entering an era where the “warehouse-native” trend will evolve into a “lakehouse-integrated” standard, where the boundaries between the CDP and the enterprise data platform vanish entirely. Between now and 2028, we will see a massive exodus from proprietary data silos as companies realize that owning their “identity graph” is just as important as owning their brand. We will see the rise of autonomous brand engagement where AI agents sit directly on top of these open Iceberg tables, triggering journeys in real-time without ever needing an ETL process. The winners in this space will be the companies that stop trying to “own” the customer data and instead focus on providing the most sophisticated intelligence layer on top of the data the customer already controls. Flexibility will be the only currency that matters, and open-source compute is the gold standard that will back it.
