Where QA Meets AI Governance

An AI system can pass its tests before release and still behave unexpectedly when people begin using it. New inputs, changing data, model updates, and different operating conditions can expose weaknesses that were not apparent during testing. For organizations relying on AI, this raises the question: what evidence tells us that the system remains acceptable for its intended use?

At STARWEST 2026, I explored this challenge in my presentation, “Testing AI Systems That Change Over Time.” The conversation focused on how testers and quality assurance professionals can evaluate AI-enabled features whose behavior may vary. It also connects directly to governance because deciding what counts as acceptable behavior requires agreement about risk, responsibility, and the boundaries of use.

Consider a GPS that recommends one route in the morning and another in the evening. A different answer does not automatically mean the system failed. Traffic, roadwork, or weather may have changed. We judge the recommendation by whether it remains appropriate under the circumstances.

AI testing requires a similar understanding of context. A customer service assistant might provide the same answer in several ways without compromising quality. However, if it invents a refund policy or promises something the organization cannot deliver, that variation crosses a meaningful boundary. Testing must distinguish between differences that are acceptable and differences that could cause harm.

Governance helps establish those boundaries. Quality Assurance (QA) provides evidence about whether the system stays within them.

Imagine an AI tool that summarizes customer complaints. Its wording and sentence order may vary, but it should preserve the customer’s concern, relevant facts, and requested resolution. It should not introduce details that were never provided. A technically fluent summary can still misrepresent the customer’s experience, especially if an employee uses it to make a decision without checking the original complaint.

Screenshot

Evaluating that tool means looking beyond whether it produces a readable summary. The organization needs to define what information must be preserved, what errors require escalation, and who has the authority to accept or reject the output. Those decisions give testers concrete criteria to evaluate.

The depth of testing should also reflect the consequences of being wrong. A poor product recommendation and an incorrect output used in a consequential decision do not warrant the same level of verification. As the potential impact increases, teams need stronger evidence, more demanding scenarios, and clearer review procedures. This is why realistic testing matters. Clean demonstrations show how a system performs under selected conditions. Real users bring incomplete information, ambiguous requests, unfamiliar language, and unexpected situations. Including these conditions in testing helps organizations understand where the system’s reliability begins to weaken and where human judgment becomes necessary.

That work must continue after launch. Monitoring can reveal patterns that earlier testing missed, but someone must be responsible for interpreting those signals and acting on them. An increase in unsupported claims, repeated corrections, or failures affecting users should prompt a defined response. Collecting information has limited value if no one knows when to investigate, restrict use, or pause the system.

My SAFER AI™ Protocol connects these questions at the moment of reliance, when someone moves from receiving an AI output to acting on it. Scope establishes the intended use. Authority identifies who can decide or escalate. Failure Awareness considers what could go wrong. Evidence supports the decision to rely, and the Record preserves what was checked and why the decision was made.

For QA professionals, this creates a direct connection between testing and organizational accountability. Test results inform decisions about whether, where, and under what conditions an AI system should be used. Governance makes the responsibility for those decisions explicit. As AI becomes part of everyday workflows, confidence needs to be supported by evidence that remains relevant as conditions change. Passing a test is valuable, but responsible reliance requires continued attention.

When your AI system behaves differently tomorrow, will your organization know whether the change is acceptable, and who is responsible for deciding?

 

Related Articles