Skip to main content

Two AI questions every government technology programme should ask its testers.

Agencies are under pressure to adopt AI and to assure systems that already contain it. These are two different testing problems, mixing them up is where programmes can come unstuck.

Author avatar
Gary Brookes 24 August 2026 · 3 min read
Two AI questions every government technology programme should ask its testers.

Every government technology programme is having an AI conversation right now. 


Sometimes it starts with ambition: a minister's commitment, a digital strategy or a productivity target. Perhaps the new platform you are implementing turns out to have AI features switched on already, whether anyone asked for them or not.

In quality assurance, that conversation keeps collapsing two very different questions into one:


The first is: Should we use AI to do our testing? 
The second is: How do we test systems that contain AI? 


They sound similar but they are not, and programmes that mix them up tend to get the worst of both, with inflated expectations on the first and a blind spot on the second.

Question one: using AI to test

The honest answer is that AI is already extremely useful in the testing process, within limits that matter more in government than anywhere else.

Where it genuinely helps: 

  • Drafting test cases and scenarios from requirements, which testers then correct and extend. 
  • Maintaining large regression suites, where AI-assisted tooling can flag which tests a change is likely to break. 
  • Triaging defects, clustering duplicates and suggesting severity, so people spend their time on judgement rather than sorting.

Now the limits. A general-purpose AI tool is not a substitute for a specialist tester, any more than a spell-checker is a substitute for an editor. 

It doesn't know your legislation, your business rules, your integration landscape or your users. It will happily produce plausible test cases that miss the one vital scenario that matters. Plausibility is exactly what makes the gap dangerous. Every AI-generated artefact still needs a competent human to check it. The productivity gain is real, but it is smaller than the demonstrations suggest.

One limit that is relevant to the public sector: data. The moment a tester pastes production-derived data, citizen records or sensitive configuration detail into a consumer AI tool, the programme has a security incident, not a productivity gain. 

If your testers are using AI tools today (and statistically, they are), the first governance question is not which tool. It is 'Where is the data going?'.

Question two: testing systems that contain AI

This is the harder, uninvited problem. Modern platforms increasingly ship with AI built in: document summarisation, chat interfaces, recommendations and triage features. A department can acquire an AI risk without ever having captured their AI requirements.

Testing these systems is a different discipline from testing ordinary software, in three ways:

  1. The pass/fail line moves. A payroll calculation is either right or wrong. A model that summarises correspondence or prioritises casework is more or less accurate, so acceptance has to be expressed as thresholds. What error rate? Measured how? On whose data? If a programme's acceptance criteria for an AI feature look identical to its criteria for everything else, the feature has not really been tested.
  2. The data is part of the system. Bias and gaps in training or configuration data translate directly into behaviour in production. That makes validating datasets (ie their variety, their labelling and their fit to the people the system will serve) testing work rather than a data team's private concern. In government the population served is everyone, so "the data under-represents this group" is not a technical footnote. It is a service-delivery and fairness finding.
  3. The work doesn't end at go-live. Ordinary systems mostly stay tested once tested. Model-driven behaviour drifts as data, prompts and vendor updates change underneath it. So AI features need a monitoring and re-validation rhythm after release, closer to how agencies already treat security than how they treat UAT.

None of this relaxes the fundamentals, eg. accessibility obligations don't bend because a feature is AI-powered. A chat interface that fails WCAG is a compliance failure regardless of how clever the model is. Performance, integration and security testing all still apply. AI just adds surface area.

Five questions for programme leaders

If you carry accountability for a programme with AI anywhere in it, in the toolchain or in the product, five questions will surface most of the risk:

  1. Does our test strategy treat "AI as a tool" and "AI in the system" as separate problems with separate controls?
  2. What is our rule for what data may enter which AI tools? Would our testers recognise a breach of it?
  3. For each AI feature we are accepting, what is the measurable threshold? Who sets it?
  4. Who validates the data the AI depends on? Against what definition of the people it serves?
  5. What happens after go-live? Who watches for drift and how often?

If any of those has no owner, that is your gap. It is far cheaper to close it in test than to explain it afterwards.

Gary Brookes is the founder of the Luvo Group, a Sydney-based group providing independent software testing and quality assurance, technology advisory and specialist technology recruitment to public sector and enterprise organisations across Australia and New Zealand.

Published by

Gary Brookes Director, Luvo Pty Ltd

About our partner

Luvo Pty Ltd

Luvo Group is a Sydney-based technology services group helping public sector organisations deliver technology change that works. Three businesses, one accountable partner.Luvo Testing — independent software testing and quality assurance: test strategy and management, functional and regression testing, automation, performance, WCAG accessibility testing, independent QA and project assurance, and managed testing. Deep specialism in ERP and Microsoft Dynamics 365 Finance & Operations programmes. SIT, UAT and regression cycles, Azure DevOps test management uplift, defect management and programme rescue, including legacy AX-to-D365 migrations.Luvo Solutions — technology advisory, AI and process optimisation: AI readiness assessments, AI governance workshops, executive AI briefings, technology health checks, vendor selection reviews and implementation recovery. Independent and vendor-aware rather than vendor-led.Luvo Talent — specialist technology recruitment: D365 consultants, project managers, business analysts, developers, architects, test managers and automation engineers. Permanent and contract, drawing on a network built over two decades in the Australian technology market.We are independent. We do not build the systems we test and we do not sell the software we advise on, so our advice answers one question only: will this work when it matters? For government programmes carrying public accountability and legislative deadlines, that independence is the point.Our consultants have delivered for Australian public sector and enterprise organisations across transport, policing, emergency services, higher education, health and financial services. Sydney-based, we deliver Australia-wide and into New Zealand, mobilise fast, put senior people on the work, and are transparent about scope, assumptions and pricing.Book a discovery call to discuss your programme, your technology decisions or your critical roles

Learn more