Key Takeaways
- An open table format and an open platform are different claims, and a vendor can hold the first honestly while the second stays closed.
- Lock-in extends beyond storage into the catalog, governance model, and performance tuning.
- The only reliable test of openness is what leaving costs, measured in engineering months and not egress bills.
- EU regulation is removing the fee-based part of exit cost, but the architectural part remains untouched.
- The clearest test of a vendor's openness is the quality of its answer when asked to walk through migrating off its own platform.
An open data lakehouse is supposed to mean freedom from vendor lock-in, and every lakehouse vendor now describes its architecture that way. Most of them are telling the truth about the part they are describing: data sits in Parquet, tables are managed by Iceberg or Delta, and any engine can read the files.
Then an enterprise decides to leave and discovers the files were never the hard part. The claim is about the whole stack, and the gap between that claim and the actual architecture only shows up on the day someone tries to exit.
What Does 'Open' Mean in an Open Data Lakehouse?
Open means two separable things: the table format anyone can read without a vendor's software, and the platform layers above it that determine whether your work is portable.
The first is settled technology. The second is where the commercial question lives.
Read that table from top to bottom, and the pattern is clear. Openness is strongest where standards bodies did the work and weakest where vendors compete on features. That pattern is the structural reason enterprises are walking away from vendor-locked architectures even when the storage layer underneath was open all along.
How Does Vendor Lock-In Happen Inside a Data Lakehouse Even When the Table Format Is Open?
Lock-in happens above the format, in the catalog, the governance model, the tuning, and the scheduler, where the work your team produces has nowhere else to run. Four mechanisms account for most of it.
Proprietary performance optimizations
Caching layers, statistics, indexing, and query acceleration apply inside one vendor's engine, so two years of tuning becomes two years of work that does not travel. Move the same tables to a different engine and the team is effectively starting the optimization work over, even though nothing about the underlying data changed.
Catalog gravity
The table registry decides which engines can find your tables, and a catalog reachable only through one vendor's API makes every other engine a second-class citizen regardless of file format. A second engine can often still read the files directly. Still, without catalog access, it loses schema evolution, snapshot isolation, and concurrent write safety, which makes it unusable for anything beyond ad hoc reads.
Governance metadata that will not export
Row filters, masking rules, classifications, and access policies are expressed in each platform's own model, and there is rarely a clean export, so you rebuild the policy set by hand. That rebuild is not just tedious; it is also where compliance gaps get introduced, since a rule that quietly did not carry over is easy to miss until an audit finds it.
Orchestration coupling
Jobs written against a platform's proprietary scheduler, notebook runtime, or SDK are rewrites, not migrations, because the job logic itself assumes APIs that only exist inside that one platform.
None of these requires bad faith from the vendor. They are the natural result of building useful features on top of a commodity storage layer, which is exactly why a vendor lock-in data platform assessment has to look at the layers where differentiation lives.
The practical counterweight is architectural discipline, and the format choices that build a data lake that doesn't lock you in are the foundation the rest of it rests on.
What Role Do Open Table Formats Like Iceberg, Delta, and ORC Play in Avoiding Lock-In?
Open table formats guarantee that your data and its metadata remain readable by other engines. They guarantee nothing about the layers where your team's work accumulates.
The catalog gap is the one the Iceberg community addressed directly, and the reasoning is instructive. Pluggable catalogs caused practical problems as the project grew, because each catalog had to be implemented in multiple languages and commercial offerings struggled to support many different catalogs and clients, so the community defined a single REST protocol any catalog can implement.
That standard now separates vendors who let any engine reach their catalog from those who do not, which makes support for it a sharper procurement question than which table format a platform uses.
Choosing between Iceberg, Delta, and ORC matters less than most evaluations assume. All three keep your files portable. For a working view of what the leading format does and where it fits, these Iceberg table features, benefits, and best practices cover the ground an architect needs before committing.
What Hidden Exit Costs Should Enterprises Watch for Before Committing to a Lakehouse Platform?
The costs that matter are engineering effort to rebuild what the platform held, not the bill for moving bytes.
Egress has always been the most visible exit cost and it is the one now disappearing. The European Commission confirms that the Data Act entirely removes switching charges, including charges for data egress, from 12 January 2027, with a transitional period running from January 2024 during which providers may still charge the costs they incur. For any enterprise with EU operations, the fee-based portion of exit cost has an expiry date.
What remains is everything the fee was distracting from:
- Governance rebuild: Every masking rule, row filter, classification, and access grant re-expressed in another platform's model, then re-tested against the same compliance requirements.
- Orchestration rewrite: Jobs authored against a proprietary scheduler or notebook runtime, ported and revalidated end to end.
- Performance re-tuning: The workload behaves differently on a different engine, so the optimization work starts again from scratch.
- Parallel running: Both platforms live while confidence is established, which means paying twice for a period nobody plans for accurately.
- Retraining: A team fluent in one platform's tooling needs weeks or months before it is equally productive elsewhere.
Each of these is measured in engineering months, which is why it is worth asking what a platform is built to avoid before you sign rather than after.
See how Acceldata's xLake OS Foundry keeps enterprises on open formats with no proprietary formats and no hidden exit costs.
How Can Enterprises Evaluate Whether a Lakehouse Platform is Genuinely Open?
The most reliable way to evaluate whether a lakehouse platform is genuinely open is to ask the vendor to walk through the exact steps and cost of migrating off it onto open-source tooling, then judge the answer they give.
A specific, unembarrassed answer means the exit path was designed. A vague one means it was not, and that is the finding. Six questions get you there:
- What format is our data returned in, and can any engine read it with no conversion step? A vendor that hesitates here is describing a format lock-in they have not had to solve.
- Which catalog do you use, and does it implement the Iceberg REST specification so other engines can reach our tables? This is the question that tests catalog gravity directly.
- Can we export governance policies in a form another platform could import, or do they get rebuilt by hand? A "rebuilt by hand" answer is the governance cost showing up before you have even signed.
- Do our job definitions run outside your platform without a rewrite? This is the orchestration coupling question in practical form.
- Which performance optimizations are proprietary, and what happens to that tuning when we move? This surfaces how much of the platform's speed is genuinely portable.
- Who has done this migration before, and can we speak to them?
The last one separates confidence from marketing. A vendor whose customers have successfully left usually knows the number and will give it to you, since the exit path being real is the argument.
The architectural side of this is worth understanding before the procurement conversation, and Acceldata's open, multi-cloud data platform architecture sets out what a multi-cloud data architecture looks like when portability is a constraint on the design.
Judging Openness by What It Costs to Leave
The real test of an open lakehouse comes when an enterprise decides to leave, not from the language used on a product page.
Before signing, get these in writing:
- The data export process, the format it returns, and the time it takes should all be documented up front.
- The catalog should implement an open specification that other engines can reach.
- Governance policies should be exportable in a form another platform could import.
- Every proprietary component should be named, along with what replacing it would involve.
- A named reference customer should confirm they have actually migrated off.
A platform that makes exit expensive was never fully open, whichever table format it stored data in. That is also why the question belongs in evaluation, since by renewal the leverage has moved.
Acceldata's xLake platform runs on the same open formats across cloud, on-premises, and hybrid environments from one control plane, so the compute and governance layers stay portable alongside the storage layer that was always open.
Find out what leaving would cost before you are committed to staying. Book a demo with Acceldata and walk through what an honest exit from your own stack would actually take.
FAQs: Open Data Lakehouse and Vendor Lock-In
Does using an open table format like Iceberg guarantee a lakehouse is free of vendor lock-in?
No. An open table format such as Iceberg reduces lock-in at the storage and table layer, but dependencies on catalogs, governance, compute engines, proprietary features, and operational tooling can still create switching costs.
How much does migrating off a proprietary lakehouse platform typically cost?
Migration costs vary widely based on data volume, workload complexity, proprietary dependencies, and the amount of code that needs to be rewritten. Beyond data transfer, enterprises may also incur costs for engineering effort, parallel operations, testing, and temporary infrastructure during the transition.
Can an enterprise mix open source and proprietary components in the same lakehouse?
Yes, enterprises can combine open source and proprietary components in the same lakehouse, using open technologies for core data layers while relying on managed services for specific capabilities. The main consideration is how proprietary dependencies affect portability, interoperability, and future migration options.
What questions should procurement ask a lakehouse vendor about exit costs before signing?
Procurement should ask about data export fees, proprietary formats or APIs, contract termination terms, migration support, and the effort required to move workloads and metadata. They should also clarify which components remain portable and what costs could arise from running the old and new platforms in parallel during migration.
Does open source lakehouse tooling cost more to operate than a fully managed proprietary platform?
Open source lakehouse tooling can cost more to operate when organizations must manage infrastructure, upgrades, security, and specialized engineering resources themselves. A fully managed proprietary platform typically reduces operational overhead but may introduce higher licensing, usage, and vendor lock-in costs.







