Digital Maturity as the Foundation.
Before exploring how AI can be applied in regulated industries, it is important to consider digital maturity. AI is often seen as a tool that can solve any problem. While it can automate and accelerate almost any task, its success depends on access to the right information.
AI cannot fill knowledge gaps or read minds. To perform effectively, it needs access to documentation that is available and legible to AI, such as business processes, business rules, security roles, and data requirements. Digital maturity is not just about having the right tools and licenses; it is also about making knowledge accessible and structuring it so that AI can use it. Without this foundation, even the most advanced AI solutions will have limited value.
Defining Intended Use and Human Oversight
When introducing AI into your portfolio of tools, it is crucial to be specific about the task you want it to solve. This is essential in a GxP context because the task will determine the validation requirements for the AI tool.
At Cegeka Denmark, we have worked with Microsoft products such as Copilot, Power Automate, and Power Platform to develop tools that solve concrete tasks and support specific workflows. The goal was never to replace people, but to accelerate and support the completion of specific tasks by shifting the most time-consuming work from execution to review. This makes human-in-the-loop oversight an essential part of using AI. At the core of these initiatives were two questions: What task do we want the AI agent to solve, and how do we want it to solve it?
In my opinion, it only makes sense to define, design, and build AI agents with the same level of detail as any other software system. A thorough analysis must take place so that clear requirements, functionality, and success criteria are defined and derived from one another, creating bidirectional traceability. Human-in-the-loop quality gates must be incorporated to ensure the quality of AI-generated outputs. Human judgement must remain decisive. Every AI workflow should include at least one human-in-the-loop quality gate to ensure that no work is performed blindly.
What Annex 22 Means for AI
The European Commission’s draft Annex 22 on Artificial Intelligence identifies three key areas that must be addressed when working with AI in a GxP context:
- Intended use
- Metrics and evidence
- Monitoring and human-in-the-loop (HITL) oversight
Most importantly, Annex 22 places particular emphasis on continuous monitoring because data sources change, prompts evolve, user behavior varies, and model updates can affect responses. This is not in opposed to practices of non-regulated businesses; rather, pharmaceutical companies must be able to demonstrate compliance, especially with the intended use of the AI tool.
This is where it becomes challenging: how can you maintain control when an AI system’s inputs, configuration, and outputs may change over time?
Maintaining Control of AI
By combining analytics (what users do) with evaluation (how the agent performs), we gain insight into agent quality and can improve it systematically over time.
In regulated industries, intended use is essential. Systems must be validated to demonstrate that they perform as intended. But how can this be achieved when an AI model or its operating context may change over time?
AI tools must be tested throughout the development phase, just like any other software system. However, control of AI agents is established through continuous monitoring after the tools are published. This is similar to how software systems are maintained and subjected to regression and smoke testing. AI tools can also be tested at multiple levels, including testing the models used and evaluating prompt responses.
For solutions developed using Microsoft products, testing capabilities are available to assess whether an agent behaves as expected and continues to serve its intended use. It can also be useful to define acceptable quality thresholds and tolerated variation in responses and behaviors so that minor changes do not automatically indicate failure.
Defining “Correct” in a Non-deterministic World
Because AI responses can vary, testing cannot rely solely on exact matches.
Instead, evaluations focus on how well the response meets the intent of the question and the agent.
An agent can be tested in multiple ways and at different levels. One obvious approach is to validate responses using:
- Accuracy thresholds
- Hallucination levels
- Grounding compliance
- Safety compliance
- Consistency
- Intended use → What scenarios are tested?
- Monitoring (HITL) → How do we ensure ongoing control?
- Metrics and evidence → What defines success?
Unlike traditional testing, the expected result may relate to the intended outcome rather than to a single fixed answer. Human review therefore remains essential, especially where context and judgement determine whether an answer is acceptable.
Back to GxP and Validation
From a GxP perspective, agent testing is about control, traceability, and evidence:
By enforcing bilateral traceability between requirements, design documentation, test cases and the product, tracking in metrics can be done to determine the test coverage of requirements and thereby if the tool serves the intended use. To validate the success of an AI, success must be determined as well as accuracy thresholds.
Continuous monitoring must be done through continuous testing as well as designing workflows for AI use with Human-in-the-Loop as a key activity.
Introducing AI agents into validated systems and regulated environments requires the same discipline applied to other software systems, combined with monitoring and control that addresses AI’s variable behavior. It is crucial that accountability remains through deliberate Human-in-the-Loop quality gates. However, what determines the level of automation and use of AI is the digital maturity of the organization.