Skip to main content

I've been in data and engineering for a long time now, long enough to remember when "big data" meant standing up a Hadoop cluster and hoping your ops team could keep it alive. Long enough, too, to have watched more than one wave of "we're finally solving data silos" move through the industry, with a "we're finally democratising data" pitch right behind it.

None of that is new, and neither is the progress. Companies have been chipping away at silos, and at genuinely governed self-service access, for the better part of two decades. DATAVERSITY's 2026 research still puts silos as the top concern for 68% of organisations, not because nothing worked, but because the target kept moving under everyone's feet.

Here's what's actually changed, and why it matters more urgently now than it did even two years ago. The pipelines and platforms that were good enough for dashboards, monthly reporting, and a human analyst quietly sanity-checking a number before it went into a slide are not automatically good enough for agentic AI and the new wave of use cases sitting on top of them. An agent making an unattended decision needs something those older pipelines were never built to give it: data that's classified, lineage-tracked, and semantically legible without a person filling in the gaps from memory. That's the real story in this piece. Not that silos or democratisation are new problems, but that the old way of solving them has met a use case it wasn't designed for, and it needs to change.

That's not a reason for doom and gloom. It's a reason to look honestly at what's actually changed in the architecture, and what hasn't.

We've actually moved the architecture forward

Go back to the Hadoop era and the whole conversation was about storage: how do we get petabytes of data into one place without falling over. That was the entire ambition. Fast forward to now and the architecture has matured past recognition. Cloud-native lakehouses, data mesh as an operating model rather than just a buzzword, active metadata and semantic layers that actually let you find and trust data across domains without physically moving it. We stopped trying to force every dataset into one giant central warehouse and started building systems that respect where data naturally lives while still making it discoverable and governed. That shift, from "centralise everything" to "connect and govern everything," is the real engineering achievement of the last decade, and I don't think it gets enough credit. Gartner went as far as predicting that by 2030, universal semantic layers will be treated as critical infrastructure, sitting alongside data platforms and cybersecurity on the CIO's non-negotiable list. That's a strong signal of how far the conversation has moved from "where do we store it" to "how do we make it universally understandable."

And I want to give credit where it's due here, because a lot of organisations have put serious capital and serious years into this. Financial services firms rebuilding core data platforms around governed self-service. Retailers investing in data products so a merchandising team doesn't need to file a ticket to answer a basic question. Public sector bodies in the UK working through genuinely difficult data-sharing constraints to get closer to a single, trusted view of a citizen or a patient. These are multi-year, well-funded modernisation programmes, and they're why the intensity of the silo problem has come down even where the underlying fragmentation hasn't fully disappeared. The tools got sharper. The cost of living with some fragmentation went down. That's real progress.

Democratisation was never about removing the fences

I think this term gets misread a lot, especially by boards worried about AI adoption outpacing control. Data democratisation was never "give everyone access to everything." It's closer to the opposite: build the governance and the semantic layer well enough that access can be broad because it's safe. A sales lead should be able to ask a plain-language question and get a trustworthy answer without waiting on a data team, and that only works if the underlying platform has already done the hard work of classification, lineage, and access control. For anyone operating under GDPR or getting ready for the EU AI Act's data governance obligations, that isn't a nice-to-have, it's the only version of democratisation that survives an audit. Gartner's projection that non-technical users will originate the majority of new data integration flows this year only makes sense in a world where the guardrails are doing quiet, constant work in the background.

Data still needs to be democratised and shared, but what a human needs from that data and what an AI agent needs from it are not the same thing, and treating them as the same is where most existing self-service programmes fall short.

The part that still needs honest attention

Where I'll admit we're not finished: literacy. Roughly six in ten executives say their teams still don't have the confidence to self-serve, even when the access is sitting right in front of them. That's not a platform problem, and no amount of additional tooling fixes it on its own. It's a change management and enablement problem, and it deserves the same investment we've been putting into infrastructure. The organisations pulling ahead right now aren't the ones with the fanciest data stack. They're the ones treating data fluency as seriously as they treat the platform underneath it.

The hardest part was never the pipeline

If I'm honest, the technology was rarely the thing that slowed a programme down. What slows things down is people agreeing on who owns a dataset, who's accountable when it's wrong, and who gets to say yes when another team asks to use it. We've had the chance to work alongside data and engineering teams inside some very large, very established organisations, the kind with decades of history, multiple business lines, and data estates that grew through acquisition as much as design. Watching how those organisations navigate a question as fundamental as data ownership taught me something that no architecture diagram ever will: it is never easy, no matter how good the platform underneath it is. Ownership sits at the intersection of org chart, incentive structure, and old turf lines that predate anyone currently in the room.

The organisations that get this right don't treat it as a side conversation to the platform build, they treat it as the main event. They put clear accountability against data domains the same way they'd put accountability against a P&L. They give data stewardship real authority, not just a title. And they invest in the unglamorous work of getting people across functions to actually trust each other's numbers, which is a cultural achievement as much as a technical one. Gartner's own prediction backs this up: by 2030, six in ten organisations that successfully differentiate with AI will be led by executives who've prioritised mastery of human relational skills over pure technical depth. The platform gets you the plumbing. The culture gets you the outcome.

Agentic AI is compressing the technical side of this faster than most people expect

There's a newer piece of this story worth adding, because it changes the excuses available to an organisation, not the fundamentals. Agentic AI is now doing real work across the parts of the pipeline that used to make silo remediation a multi-quarter slog. At ingestion, agents can look at a legacy source, infer its schema, and propose the mapping into a shared model, work that used to mean weeks of an engineer manually reverse-engineering someone else's undocumented database. On governance, agents are increasingly enforcing classification, lineage, and access policy as code at the point a new source is connected, rather than governance arriving months later as a bolt-on audit exercise, which is exactly the sequencing that used to let shadow silos re-form faster than anyone could formally close them. In deployment, agent-assisted pipeline generation and self-healing on schema drift mean the distance between "we've been granted access to this dataset" and "it's live, monitored, and safe to query" has genuinely shrunk, from months in the old model to weeks or less. And on quality, continuous agent-driven anomaly detection and quarantine means a newly unified dataset earns trust faster, which matters more than people give it credit for, because a freshly opened-up silo that nobody trusts yet is functionally still a silo.

I'd be careful not to overclaim here. None of this touches the ownership and culture questions I've just spent several paragraphs on, and I don't think it ever will, because those are human decisions, not engineering ones. What it does mean is that organisations have fewer good excuses left. The technical lag that used to justify a slow, cautious rollout of governed access is compressing month by month. The remaining bottleneck is even more visibly the human one than it was two years ago.

Where this leaves us

I don't think data silos are a problem you solve once and close out. They're a condition you manage well or badly, and for the first time, managing it well is genuinely achievable at scale.

There's also a sharper reason to get this right now than there was even two years ago. The organisations doing this well aren't just serving dashboards and quarterly reports anymore, they're serving agents. An AI agent making a decision, or an LLM answering a plain-language question, doesn't tolerate ambiguity the way a human analyst quietly did for years, working around a stale field or a half-documented table because they'd learned to. Agents need data that's classified, lineage-tracked, and semantically described well enough for a machine to trust it unattended, not just a human who's built up years of institutional folklore about which numbers to believe. That's a materially different bar. It means the old way of handling data, tolerable when a person was always in the loop to sanity-check the output, quietly stops being good enough the moment you put an agent or a model in front of it. Getting data into a shape that a new-age AI platform can actually rely on isn't a separate initiative from fixing your silos, it's the same work, just with a less forgiving audience on the other end.

The work left isn't mostly technical anymore. It's cultural, and it's about building the confidence across an organisation to actually use what's been built. That's a much better problem to have than the one we started with.


Happy to chat more on this shantanu@clair-x.ai