Data platforms that outlive vendors
We design lakehouses on open table formats and leave behind systems your own team can operate.
data and keys, EU only
nothing is held by us
no lock-in, exit ready
every choice, with what we rejected
your team runs it
nothing billed outside it
What we do
We build and repair data platforms: lakehouses a team can operate without a vendor on call, warehouses moved onto open table formats, and access control that can be shown to an auditor.
Lakehouse architecture
Apache Iceberg on object storage, with a REST catalogue, snapshot isolation and time travel. Engines chosen for the workload rather than for the vendor.
Platform engineering
Declarative pipelines, orchestration, schema evolution, table maintenance. The unglamorous work that decides whether a platform still runs in three years.
Sovereign infrastructure
Migration off managed services onto European providers or your own hardware, with the owner, licence and governing jurisdiction of every component written down.
Governance and compliance
Access control, classification, lineage and retention, mapped to what the GDPR requires and to the dates the AI Act sets.
How we work
These four rules decide how an engagement runs, whether the work lasts three weeks or three years.
Decisions in writing
Decisions are written down, with the reasons and with the alternatives that were rejected. A decision you cannot reconstruct in two years is a decision you will have to take again.
Your repositories, your accounts
What we build is code and configuration in your repositories, running on your accounts.
Every component accounted for
No component enters an architecture before its owner, its licence and its governing jurisdiction are stated.
Measured, not estimated
We say what we do not know. An estimate given without the measurement behind it is worth nothing to either of us.
Three ways to work with us
Every engagement is scoped and quoted for your situation: what needs to be built or moved, where it runs, and what constraints you are under. No list prices, no standard package.
Platform Assessment
Know where you stand before you commit a budget.
A few weeks · fixed scope · quoted after a first call
What's included
- Architecture and data-flow review of your current platform, batch and streaming
- Sovereignty and GDPR exposure map: where data lives, under which jurisdiction, who can reach it
- Cost and lock-in analysis of your current vendors
- Written decision records for every recommendation, with the alternatives we rejected
- A prioritised roadmap you can execute with or without us
Platform Build
Design and build a data platform you own end to end.
Several months · milestones or time and materials · quoted on scope
What's included
- Lakehouse architecture on open table formats and open engines
- Batch and streaming pipelines, orchestration, schema evolution, data quality
- Deployment on infrastructure you control: your cloud account, your data centre, or a European provider
- Governance built in: access control, classification, lineage, retention, audit trail
- Foundations ready for AI and analytics workloads
- Runbooks, documentation and hand-over so your own team can operate it
- Everything delivered into your repositories, from day one
Operate & Advise
Keep the platform healthy without hiring a full team.
Monthly retainer · sized to your platform · quoted on scope
What's included
- Monthly DataOps: upgrades, performance, cost review, incident follow-up
- Governance reviews as regulations and your data change
- Architecture advice for new use cases before they become projects
- Agreed response window for incidents
- Quarterly written review of the platform against your roadmap
Every engagement is quoted individually after a first conversation: the scope, the location, the constraints and the timeline all change the price, so we do not publish one. Contracted with DATASINK SARL, RCS Versailles 830 255 766, registered in France since 2017. Work is performed under EU jurisdiction.
What a written decision looks like
ADR-015 · Backup and disaster recovery
Status
Accepted · 2026-09-06
Context
One provider outage must not cost a customer their platform.
Decision
Back up to a second provider from the first byte of data:
object-locked bucket, cold storage, one-way replication.
Recovery is a full rebuild at the other provider, from
infrastructure as code and table snapshots. Tested every
quarter, dated report, measured RTO and observed RPO.
Rejected — active-active across two providers
Workable for compute, impractical for state: no native object
replication between providers, egress paid both ways, a single
commit point in the table catalogue. More incidents, each less
severe. For a team our size, a bad trade.
Excerpt from our own architecture log, shortened for this page.
Every architecture decision we make is recorded the same way: the context, the decision, and the alternative we rejected with the reasons it lost. That record is what lets someone reopen the question in two years without starting from nothing. You get the same thing for your platform, in your repository, dated and in plain language.
See how we workTell us what you are trying to build.
If the work is outside what we do, we will say so and point you elsewhere.
Track record
We work with the people who carry the consequences of a data architecture: the CIO who signs for it, the data lead who runs it, the data protection officer who answers for it.
DataSink has been registered in France since 2017. The work has been carried out for large French organisations in banking, rail transport, public investment, retail and telecommunications — the kind of environment where a data platform has to coexist with systems older than the team running them.
We do not publish client names without their agreement, and we do not publish their logos at all. What we can discuss in detail, under the confidentiality your own procurement requires, is what the work consisted of.
Frequently asked questions
The questions we get asked before a first engagement, answered plainly.