Getting AI Right: How GovTech Is Making Sure Government Data Means What It Says
2 October 2026
Category
Building Tech at GovTech
Discover how GovTech's context layer helps AI not just read government data, but also understand it, so officers get answers that are accurate, traceable and grounded in policy.
.png)
Ask a colleague from another agency how many cases were resolved last quarter, and you'll likely hear: "It depends on what you mean by resolved." Ask an AI, and you will get a number, without hesitation.
With GovTech continuing to build and strengthen the data foundations that the public sector runs on, from cloud data platforms to automated pipelines and AI tools, moving data quickly and reliably between systems is now largely a solved problem.
But, as AI becomes more central to how the government works, a harder question has emerged: With good data in place, how do we make sure AI genuinely understands what that data means?
AI taking data at face value
Picture an officer putting that same question to a chatbot that gives conversational access to data, where the AI retrieves the relevant data to answer queries in plain language.
The AI scans the database, counts every row where the status is marked "RESOLVED", and returns 142, instantly and with complete confidence. Nothing about the number looks wrong.
The correct answer is 105. Thirty-seven of those cases were closed automatically after a period of inactivity. No officer handled them, and no resolution was ever verified.
The AI wasn't broken. Its query ran correctly against the data it was given. It simply didn't know what "resolved" meant in the agency's context. It could see that a case was marked resolved, but not how it got there. That is the limit of a database schema. It describes how data is structured, such as a status field and the values it can hold, but not what those values mean in policy and practice. An experienced officer would know to pause and ask a clarifying question. An AI tool, without guidance, has no way to make that distinction.
Closing the gap between what data records, and what the data actually means is the challenge GovTech's Data Practice is now taking on.
Facing the cost of inconsistent definitions
Imagine an orchestra of skilled musicians with no shared score. Each section plays well on its own, but together they produce noise.
Government data works the same way. Agencies built their systems at different times, for different purposes, and often use the same words to mean different things. "Resolved", "case" or "active" can carry different business rules from one system to the next. Human officers learn these differences through experience, and they know to ask before they answer. AI tools don't ask. They fill the gaps with assumptions.
The cost shows up in two ways.
Human friction: Officers spend time in alignment meetings reconciling spreadsheets that don't match, and chasing institutional knowledge that sits with a few colleagues.
Confident errors: AI tools guess how datasets relate to each other. They invent table joins or columns that do not exist, and produce reports that look credible but are wrong. These flawed reports can then reach decision-makers and shape the decisions they make.
These problems are a natural growing pain for a maturing data ecosystem. They tend to surface once the basics are in place, when data moves freely between systems and AI tools can reach it.
GovTech's Data Practice describes that maturity in three phases.
Data availability: Moving raw data reliably between systems through cloud warehouses and automated ingestion pipelines.
Data usability: Cleaning and structuring data, and enforcing quality standards, so systems can run efficiently.
Data understanding: Ensuring that people and AI models interpret the same metrics consistently across agency boundaries.
With the first two phases (opens in new tab) largely in place, the focus now shifts to data understanding, where the Data Practice team at GovTech is working at that frontier. Rather than reacting to errors once AI tools have already spread, it is building shared understanding into government data now, so that AI across the public sector grows on foundations that can be trusted.
Building a data handbook for Government AI
GovTech's answer is a context layer, a shared rulebook that sits between raw data and the AI tools that read it.
Think of it as a company handbook. Imagine hiring a brilliant expert who has minimal understanding of how your organisation operates on day one. They are capable, but they don't yet know your terminology, your history, or why things are done a certain way. A context layer is that handbook in digital form. It brings the AI up to speed instantly on the terms, policies, processes and past events that shape how the data should be read, so that the data it retrieves is relevant to the query.
The approach is deliberately narrow. Today's large language models (LLMs) already understand general concepts such as dates, currencies, arithmetic and everyday administrative language. By the team's estimate, that covers about 80% of what an AI tool needs to know. The context layer focuses on the remaining 20%, the knowledge unique to the public sector: official metric definitions, agency-specific rules, internal naming conventions and edge cases.
How that knowledge is stored is also of utmost importance. Traditional data governance tried to solve this problem by documenting every database field in static wikis and PDFs. It was a massive effort, and the documentation was often out of date almost as soon as it was published.
The context layer takes a different route. Instead of living in documents that go stale, definitions are written as code, an approach known as semantics-as-code. Like any other code, they are version-controlled, so every change is tracked and the latest definition is always the one in use. Alongside the definitions sit pre-tested queries that are known to return correct results. When a question comes in, the system passes only the relevant definitions to the AI before it retrieves any data.
What this looks like in practice
With the context layer in place, the officer's question is handled in three steps.
1. Interpret the request
The system identifies which policy domain the question touches and checks that the officer has permission to access that data.
2. Apply the rulebook
Instead of generating a query from scratch, the system retrieves the agency's official definition of "resolved", under which a case counts only when an officer has verified the outcome. It then uses a pre-tested query built on that definition, which excludes cases closed automatically after inactivity.
3. Verify before answering
An independent AI reviewer, using an approach known as LLM-as-a-judge, checks the result against policy before it reaches the officer.
This time, the officer receives 105, with a note explaining that 37 auto-closed cases were excluded and which definition was applied. The answer is transparent, traceable to its source, and grounded in policy. The officer can not only see the number but also understand how it was reached.
Scaling shared meaning across government
GovTech's Data Practice is preparing a proof of concept in selected public sector domains to test three outcomes:
Higher accuracy on complex questions, with a target of [85%+], by binding each query to official definitions.
Fewer invented database errors, such as non-existent columns, invalid joins or malformed SQL, by relying on pre-tested query templates.
Lower compute costs, with a target of [60–70%], by sending the model only the rules relevant to each question, which reduces the number of tokens processed.
Essentially, early results are expected to demonstrate meaningful improvements in accuracy and a significant reduction in the kind of subtle errors that are hardest to catch i.e. the ones that look right but aren't quite.
The long-term goal is consistency across whole of government, where core terms such as Citizen, Business and Case carry the same meaning in every system. As AI becomes a co-worker to officers across every agency, that shared understanding ensures it reads the numbers the way they do. It is what allows the public sector to modernise with AI at pace, without trading off accuracy or accountability, and to better serve citizens and businesses’ needs.
Starting with a single definition
Existing data systems do not need to be rebuilt to get started. A context layer grows one definition at a time, so any organisation can begin with what it already has and build from there. Here’s how you can start:
Pick one contested metric
Find the number that always comes back with "it depends", the one that sparks the most debate across your teams, and settle its definition first.
Put the definition in code
Write it, along with a tested query that uses it, as version-controlled code so every change is tracked.
Give it to AI at the point of asking
Make sure your AI tools receive the relevant definitions with each question, before they touch any data.
Let experts check the answers
Have the people who know the data review what the AI returns, so you can measure accuracy and build trust over time.
With every definition settled and every answer checked, people and AI come a step closer to reading data the same way. In time, the orchestra that once produced noise begins to play as one.
