Open records laws have pushed cities, counties, and states to publish more data than at any point in their history — payroll, permits, property assessments, contracts, inspections, and more, often refreshed on a regular cadence. By the letter of most transparency laws, the disclosure requirement is met the moment a file is posted.
By any practical measure of understanding, it usually isn't.
A dataset published as a wide, inconsistently-formatted CSV, with department-specific codes and no documentation, is technically public and functionally closed to almost everyone except a specialist willing to spend hours reconciling it. The gap between "the file exists" and "a person can use it" is where most of the value in open data quietly gets lost.
The compliance bar for transparency and the practical bar for understanding are not the same bar — and most public data efforts are only measured against the first one.
Two things have shifted at the same time. First, the volume and cadence of public data releases has kept increasing, driven by broader open-records mandates and more digitized government operations. Second, AI-based reconciliation — matching inconsistent schemas, extracting structure from unstructured records, and connecting entities across datasets — has gone from research-lab territory to something that can run in production, city by city.
That combination is what makes an Information Intelligence layer a realistic infrastructure category now, in a way it wasn't a decade ago.
Transparency should be judged the same way usability is judged anywhere else: not by whether a file was posted, but by whether an ordinary person — a resident, a journalist, a renter, a small business owner — can get a real answer to a real question in a reasonable amount of time, with the reasoning behind that answer visible and checkable against the public source.
That's a much higher bar than most current open data efforts are held to. It's also, we think, the correct one.
The opportunity isn't to publish more data — governments are already doing that at scale. It's to build the layer that sits on top of what's already public and makes it legible: connected across departments, explainable in its reasoning, and fast enough to answer the question someone actually has, not just the one the dataset was designed to answer.