How to Validate AI in Clinical Research

Thoughts from the Digital Tools & AI: Practical Applications for Sites panel at SCRS ANZ in Melbourne, July 24, 2026.

By Ellery Nadarajah, Chief of Staff at Delfa

Introduction

Sites are being handed AI whether they asked for it or not. An eQMS ships with an assistant, a CRM adds one. Plenty of teams leave those features switched off. At SCRS ANZ 2026, a great question came up: how do you validate something that keeps regenerating, and where do you take the first step?

The obstacle is not a blanket prohibition on AI. Teams need a clear strategy and a defensible way to validate the technology. Moving from a demonstration to a live workflow depends on defining the job, the controls, and who remains responsible.

A feature can appear impressive during a vendor demonstration before a site has determined whether it can be used safely and consistently in practice. The site still needs to understand the tool’s purpose, limitations, data flows, controls, and ownership.

Deterministic software can often be tested against fixed expected outputs. Generative systems introduce an additional challenge because the wording may vary between interactions. Validation therefore needs to examine whether the system continues to meet defined requirements across representative inputs and operating conditions.

The short answer. A generative AI system should be validated against its intended use, not against the expectation that it will produce identical wording every time. The site needs to demonstrate that the complete workflow, including the people who use and oversee it, performs a defined job within approved limits. The controls should reflect what the system does and the consequences if it fails.

A six-step framework for sites

The wording may vary, but the permitted task, boundaries and handover rules do not. We would use six steps to turn an open-ended question about validating AI into a review of a defined workflow.

1. Start by defining the job the AI is allowed to perform

Validation becomes more practical once the system’s job is defined precisely. Before testing a clinical research AI system, define its task in specific terms.

The site should document:

  • What task the system performs
  • What information it may access
  • What outputs or actions it may produce
  • What it must not say or do
  • Which decisions remain with site staff
  • When it must stop or escalate
  • Which outputs require human review
  • What records will be retained
  • How incidents and changes will be managed
  • Who is responsible for approval and ongoing oversight

2. Test realistic and difficult scenarios

Delfa tests agents before release using synthetic patient profiles and cases written by people. Some test conversations are routine. Others involve a participant who is confused about a medication, mentions a possible adverse event, appears distressed, or is speaking from a noisy environment.

We check whether the agent stays within the approved conversation, avoids prohibited actions, and hands the interaction to site staff when required. Post-release, we continue to monitor calls and outcomes, and interactions remain subject to QA.

3. Human oversight should reflect the risk

AI risk depends partly on what happens after the system produces an output.

For higher-risk steps, a person should review the output before it affects a participant or a study decision. The audit trail should show what the system received, what it produced, whether it escalated, what happened next, and where a person took responsibility.

Delfa’s agents do not determine clinical trial eligibility. An agent may mark someone as potentially eligible based on an approved pre-screening flow, but a person makes every eligibility decision. The agents do not provide medical advice. A clinical question, mention of an adverse event, or sign of distress triggers a handover to site staff.

Participants are told at the start that the call is automated, and there is always a route to a person. Every conversation is logged.

4. Ethics committees are reviewing a new method rather than a new activity

Clinical research sites have used recruitment calls, appointment reminders, and pre-screening questions for decades. Ethics committees already review these activities when a person follows an approved script. The HREC/IRB or relevant review body needs enough information to assess what has changed.

Delfa’s agents use constrained conversation flows rather than unrestricted generation. The scripts submitted for review describe what the agent may say, what it may ask, and when it must hand the conversation to site staff.

Every site running a Delfa agent has taken it through its ethics committee with the relevant scripts, and every submission so far has been approved. Some committees asked questions before approval, which is the review process working as intended.

Depending on the study and jurisdiction, Delfa prepares the submission pack with the site. It can include participant-facing scripts, proposed amendment language, AI disclosure wording for the participant information and consent form (PICF), a description of relevant data flows, escalation procedures, and a validation summary. We give the committee enough information to examine the system and its controls.

5. Reliable data and workflows come before AI

AI cannot repair a broken process. It scales the process and data it receives. If participant information is scattered across spreadsheets, inboxes, and paper logs, the output can sound polished without becoming dependable.

At one site where we implemented Delfa, there was no clinical trial management system (CTMS). AI performed only a small share of the useful work. Most of the value came from replacing spreadsheets with a structured, searchable participant database.

That outcome was less pronounced than a demonstration of AI, but it gave the site a reliable foundation for recruitment work.

6. Start with one measurable recruitment bottleneck

Sites do not need to digitize everything before improving one part of the participant recruitment process.

A site moving away from paper can begin with practical changes such as SMS visit reminders or a digital pre-screening log. These steps create structured data and clearer workflows that can support later automation.

The same principle applies when a site is assessing several AI products. Start with one measurable bottleneck, such as:

  • Missed appointments;
  • Slow follow-up with new leads;
  • Coordinator hours spent on initial screening;
  • Incomplete pre-screening records; or
  • Potential participants lost between referral and booking.

Pilot one tool against that problem and measure the result. Assess whether staff can see an improvement in the workflow.

Questions to include in every clinical AI vendor review, and how we answer them at Delfa

1. Where is participant data stored?

Participant data is always stored in the region of the deployment. For Australian sites, that means Australian resident storage.

2. Where is it processed?

Storage location is not the whole answer: a vendor can store data locally while routing it offshore for processing.

For Australian deployments, AI-related processing involves encrypted transmission to US-based infrastructure, with no data retained outside Australia; all data is returned for in-country storage.

3. Who else can access the data?

Data-sharing agreements, and where required, business associate agreements, are in place with all subprocessors.

4. Is the data used to train an AI model?

No. Participant data is never used to train any AI model.

5. Does the data come back to us?

Our CTMS integrations automatically return relevant information to the site’s system of record.

6. How many sites have submitted the tool for IRB/HREC review, and what happened?

Ask what materials were submitted, whether any submission was rejected or required revision, and what controls the site retains after approval.

Delfa has supported IRB/HREC submissions across its live studies. We provide templated documentation covering four components:

  • A description of how the service works;
  • The pre-screening script;
  • The conversation flow the agent follows; and
  • The knowledge base of grounded responses it can draw from.

No submission has been rejected or required substantive revision. Sites retain control over approved scripts, guardrails, and escalation rules, and once IRB-approved, scripts are not modified without returning to review. We work alongside sites throughout the submission process and share template materials for review before any deployment begins.

7. What certifications can you show us?

Ask for audit reports; they should be available on request.

Delfa is SOC 2 Type II certified, and HIPAA and GDPR compliant.

8. Where exactly is the human in the loop?

Ask which outputs are automated, which get reviewed, and whether there is an audit trail.

Our agents handle outreach, pre-screening, and scheduling. Coordinators manage warm transfers and escalations, can intervene at any point, and have a full audit trail of every interaction.

9. What happens when agents get something wrong?

Ask what evals they run, how accuracy is measured, and who is accountable. Be sceptical of anyone promising zero hallucinations.

Out-of-scope queries escalate to a human rather than guessing. Sites set their own escalation thresholds. We run continuous evals on agent outputs, and there is a named point of contact for any issue.

Ellery Nadarajah, Chief of Staff at Delfa · 05.08.2026

← Back to News