Skip to main content

Testing AI in the Real World: How KJR’s VDML Methodology Builds Trust and Reduces Risk

KJR’s VDML methodology helps organisations make sure their AI can be trusted by continually testing its reliability, fairness and risks from development through to everyday use.

Author avatar
Amy Lepinay 10 August 2026 · 4 min read
Testing AI in the Real World: How KJR’s VDML Methodology Builds Trust and Reduces Risk

VDML logo in white and blue

Testing AI in the real world: what government should be asking before deployment

Artificial intelligence is already finding practical uses across government services. It can help classify documents, identify patterns in large datasets, support contact centres, analyse images and assist staff to find information more quickly.

For government organisations, however, a successful pilot is only part of the story. The more important question is what happens when an AI system encounters real people, real data and conditions its designers did not anticipate.

An AI tool might perform well during a demonstration and still produce inconsistent results when the data changes, when it encounters an unusual case, or when it is used by a different group of people. Unlike conventional software, many AI systems do not simply follow a fixed set of instructions. Their results depend heavily on the information used to develop them and the circumstances in which they operate. Performance can also change over time.

That creates a particular challenge for government, where decisions can affect access to services, privacy, safety and community trust.

Start with the consequences of getting it wrong

Before asking how accurate an AI system is, it is worth asking what an incorrect result would mean.

Consider an AI system helping prioritise enquiries. An error might mean somebody waits longer for assistance. In another setting, an incorrect result could expose personal information or influence a decision affecting an individual.

Those consequences should determine how the system is evaluated.

This is one of the principles behind KJR's Validation-Driven Machine Learning (VDML) methodology. Developed from our experience testing AI and complex digital systems, VDML starts by defining what the system is expected to achieve and identifying the risks surrounding its use. Testing continues through development, integration and operation rather than being confined to a final checkpoint before release.

For government leaders, the useful idea here is straightforward: decide what evidence you need to trust the system before you decide whether it is ready.

Accuracy alone doesn't tell you enough

A headline accuracy figure can conceal important information.

Imagine an AI system achieves 95 per cent accuracy overall. Public servants still need to know what happened in the remaining five per cent. Were errors evenly distributed? Did particular circumstances cause more failures? Were some groups more likely to receive an incorrect result? What happens when the system encounters information that differs substantially from the material used to develop it?

These questions become particularly important when AI is used across diverse communities, locations and service environments.

Testing should therefore include realistic scenarios and unusual cases, along with the everyday situations expected to account for most use. Teams should examine the quality and representativeness of the underlying data, the consequences of incorrect outputs, privacy implications and whether staff can understand enough about a result to act appropriately.

This provides decision-makers with something more useful than a single performance score: evidence about where the system works, where its limitations sit and what controls may be required.

Government experience shows why context matters

KJR has seen this principle play out in work involving government health data.

In a project with Datarwe and Queensland Health, the challenge involved removing personally identifiable information from intensive-care patient data so that the information could be used safely for research and other purposes.

The work involved defining the problem, assessing privacy risks and validating the AI against its intended operating environment. The resulting process achieved more than 99 per cent accuracy in detecting personally identifiable information.

The percentage is significant, but the broader lesson for government is more valuable.

AI assurance needs to reflect the consequences and circumstances of the particular service. A method appropriate for analysing public information may be inadequate when dealing with patient records, vulnerable people or decisions carrying significant consequences.

Deployment is another stage of testing

Approval to go live should not be the end of scrutiny.

The information an AI system encounters can change. Community behaviour changes. Organisational processes change. New circumstances appear that were absent from the original testing.

An AI system that met expectations when introduced therefore cannot automatically be assumed to maintain the same performance indefinitely.

Government organisations need practical ways to detect deterioration, investigate unexpected results and determine when a system should be reviewed. KJR's VDML approach incorporates this ongoing monitoring because evidence collected before deployment tells you how a system performed at that point in time; operational evidence tells you whether those assumptions continue to hold.

This also gives governance bodies better information. Instead of relying primarily on assurances that an AI system has been “tested”, executives and oversight groups can ask more concrete questions: What was tested? Under which conditions? What limitations were identified? What happens when the system produces an unexpected result? How will we know if its performance changes?

Trust needs evidence

Government agencies have good reasons to explore AI. Used appropriately, it has the potential to help staff work with large amounts of information, improve access to services and automate tasks that consume valuable time.

Public trust will depend partly on how carefully those systems are introduced.

That means understanding an AI system's limitations as well as its capabilities. It means testing under conditions that resemble the environment in which public servants and communities will actually use it. And it means continuing to gather evidence after deployment.

The objective should be relatively simple: when an AI-supported government service makes or informs a decision, the organisation should be able to explain why it has confidence in that system, where that confidence has limits, and what evidence supports the judgement.

That is a much more useful foundation for responsible AI adoption than assuming a successful pilot will translate automatically into a trustworthy public service.

Published by

Amy Lepinay Marketing Coordinator, KJR

About our partner

KJR

KJR provides independent quality engineering that gives organisations the confidence to deploy complex, high-risk technology and AI systems. We focus on decisions, not just defects, helping government move from ambition to outcomes that are practical, responsible and built to last. We partner closely with public servants to deliver complex initiatives in highly regulated environments. Our strength lies in understanding how government really operates, from policy intent and procurement through to security, privacy, accessibility and ethics - and translating strategy into delivery with confidence. Unlike large consultancies that prioritise scale, or vendors that lead with tools, KJR is commercially independent and vendor-agnostic. We specialise in real-world implementation, supporting agencies to design and deliver transparent and explainable AI and digital solutions aligned to whole-of-government frameworks. KJR brings: * Deep experience delivering technology programs within government constraints* A strong commitment to responsible and human-centred AI* End-to-end capability across strategy, delivery, assurance and change* A focus on capability uplift, leaving agencies stronger and more self-sufficient Founded in 1997, KJR is known for working shoulder-to-shoulder with government teams to de-risk innovation and deliver lasting impact. KJR helps government move faster - safely, transparently and with lasting impact.

Learn more