Clinical Trials and AI
Artificial intelligence can support clinical trials from protocol development through outcome adjudication. The evidence is strongest for bounded tasks such as retrieval, extraction, prescreening, and measurement support. A faster intermediate task is not evidence of better enrollment, more representative participation, a valid endpoint, or a successful therapy.
After completing this chapter, readers should be able to:
- distinguish trial retrieval, prescreening, eligibility adjudication, consent, enrollment, retention, and clinical outcomes
- evaluate evidence for TrialGPT and other patient-to-trial matching systems
- assess AI-supported protocol, eligibility, site-feasibility, endpoint, and safety-monitoring workflows
- interpret clinical trial outcome prediction studies, including the CTO and TOP benchmarks
- design trials of AI interventions with appropriate comparators, outcomes, monitoring, and version control
- select SPIRIT-AI, CONSORT-AI, DECIDE-AI, and related reporting guidance for the correct study stage
- define physician and investigator accountability when AI affects participant-level decisions
Introduction
Clinical trials generate multiple opportunities for computational support: translating a protocol into executable criteria, finding potential participants, forecasting site workload, measuring endpoints, detecting candidate safety events, and assessing whether development should continue. Large language models add flexible retrieval and text reasoning, but they also introduce nondeterminism, prompt sensitivity, version drift, and plausible unsupported output.
No single percentage describes whether AI “works” in trials. A model may retrieve the correct trial yet misclassify one exclusion criterion. A system may clear charts rapidly without increasing consent or enrollment. An endpoint model may be reproducible without being clinically meaningful. A trial-outcome predictor may classify historical labels without forecasting future results. Each claim must retain its task, population, design, comparator, endpoint, denominator, uncertainty, and evidence strength.
This chapter focuses on physician-investigator and trial-site decisions. Upstream target discovery, molecular design, dose selection, sponsor portfolio decisions, and regulatory strategy are covered in the Life Sciences AI Handbook clinical-trials chapter.
Map the Trial Decision Before Selecting the Model
The phrase “AI for clinical trials” can refer to at least six distinct decisions. A useful evaluation begins by assigning the system to one row.
| Trial function | Model output | Immediate endpoint | Downstream endpoint that still requires measurement |
|---|---|---|---|
| Protocol and eligibility design | candidate wording, criterion structure, simulated cohort | extraction accuracy, logical validity, investigator acceptance | feasibility, representativeness, safety, estimand validity |
| Trial retrieval and prescreening | ranked trials or patient list | recall, precision, criterion accuracy, review time | referral, consent, enrollment, retention, equitable access |
| Site feasibility and operations | predicted accrual, dropout, workload, or duration | calibration and error in external sites and future periods | on-time delivery without compromised consent or data quality |
| Endpoint measurement | score, segmentation, classification, or event candidate | repeatability, agreement, sensitivity, specificity | unbiased treatment-effect estimation and clinical meaning |
| Trial-outcome prediction | probability of success, completion, transition, or approval | temporally external discrimination and calibration | better decisions, fewer harmful errors, improved development value |
| Trial of an AI intervention | recommendation or automated action | model and workflow performance | patient benefit, safety, equity, resource use, implementation durability |
The downstream endpoint is not optional. A prescreening tool should not be described as improving recruitment when the study measured only record classification. An outcome benchmark should not be described as improving drug development when no decision-impact study was performed.
Protocol Development and Eligibility Design
Protocols contain repeated, structured elements that make them attractive targets for language models: eligibility criteria, schedules of events, endpoint definitions, prohibited therapies, dose-modification rules, and reporting requirements. AI can propose normalized concepts, identify internal inconsistencies, compare versions, and translate text into candidate database queries. The output remains a draft until clinical, statistical, operational, and regulatory review is complete.
Eligibility criteria are executable clinical logic
An eligibility criterion is not merely a sentence. It is a rule with a time window, measurement source, allowable uncertainty, exception structure, and relationship to other criteria. “No recent myocardial infarction,” for example, is incomplete without defining the event, the look-back period, the authoritative source, and how suspected but unconfirmed events are handled.
An AI-assisted protocol workflow should therefore test:
- Semantic fidelity: Does the structured rule preserve the protocol’s meaning?
- Temporal logic: Are index dates, look-back windows, and sequence requirements explicit?
- Data availability: Can the required variable be obtained from the intended source at the time of screening?
- Missingness: Does absent information trigger exclusion, manual review, or a request for clarification?
- Logical execution: Do generated queries return the intended patients without syntax or join errors?
- Clinical justification: Does each restriction protect safety or interpretability, or does it unnecessarily narrow access?
Criteria2Query 3.0 illustrates the gap between extracting a concept and executing a safe cohort query. In an evaluation using 518 concepts from 20 trials, GPT-4-based concept extraction achieved an F1 score of 0.891. Evaluation of generated SQL for five trials still found 29 errors, including 10 logic errors (Park et al., 2024). Language models can draft executable criteria, but investigators and data specialists must validate the logic and resulting cohort before the query affects recruitment.
Trial Pathfinder retrospectively emulated completed non-small-cell lung cancer trials using records from 61,094 patients. Broadening selected criteria more than doubled the modeled eligible population on average, while modeled treatment hazard ratios changed by an average of 0.05. This supports empirical review of inherited exclusions, not automatic relaxation or prospective safety equivalence (Liu et al., 2021).
Adaptive design is not synonymous with AI
Adaptive trials use prespecified rules to modify aspects of an ongoing study in response to accumulating information. Bayesian models, simulations, or machine learning may support planning, but the design remains a statistical and regulatory construct. Error control, estimands, stopping rules, adaptation timing, operational bias, and independent oversight must be specified whether or not AI is used. The FDA’s final guidance treats adaptive design as a planned design method, not an AI label (FDA, 2019).
AI may be useful for scenario generation, candidate subgroup discovery, or operational forecasting before the protocol is locked. Data-driven discoveries made after outcomes are observed should be labeled exploratory unless the design and analysis plan preserve confirmatory inference.
Patient Matching, Prescreening, and Recruitment
Patient-to-trial matching usually contains three separate tasks:
- Retrieval: Find trials that may be relevant to a patient.
- Criterion review: Compare patient evidence with each inclusion and exclusion criterion.
- Workflow action: Present candidates for investigator review, contact, consent, and enrollment.
Performance at one stage does not determine performance at the next. Retrieval recall can be high while criterion precision is poor. Accurate prescreening can still fail to increase enrollment because the trial is unavailable locally, the patient cannot be contacted, the investigator disagrees, the patient declines, or capacity is limited.
TrialGPT: important evidence with a bounded claim
TrialGPT uses separate retrieval, criterion-level matching, and ranking modules. It was evaluated on 183 synthetic patient summaries with more than 75,000 patient-trial annotations. Retrieval recalled more than 90% of relevant trials after reducing the collection to less than 6% of its original size. Manual review of 1,015 patient-criterion pairs found 87.3% criterion-level accuracy, and a small two-physician user study found a 42.6% reduction in screening time (Jin et al., 2024).
These results support further evaluation of LLM-assisted retrieval and criterion review. They do not establish performance on longitudinal EHRs, increased enrollment, equitable access, retention, or patient benefit. The more than 75,000 annotations were generated trial-relevance labels across synthetic cohorts, not 75,000 independently adjudicated expert criterion decisions.
What later studies add
| Study | Design and denominator | Supported finding | Boundary |
|---|---|---|---|
| TrialGPT (Jin et al., 2024) | 183 synthetic summaries; 1,015 manually reviewed patient-criterion pairs | strong retrieval and criterion-level performance; shorter screening in a small user study | synthetic patients; no enrollment or outcome endpoint |
| Prospective molecular tumor board evaluation (Gueguen et al., 2025) | 157 sequential patients; 2,164 patient-trial pairs; four public tools | prospective use was feasible, but mean precision and recall were approximately one third | oncology setting; tool outputs still required expert review |
| Randomized prescreening trial (Unlu et al., 2025) | 4,476 EHR-identified patients randomized to AI-assisted or manual prescreening for one heart-failure trial | more than 99% of completed AI screening occurred within 15 days versus 50 days manually | eligible proportions were similar, 20.8% and 21.1%; clinical benefit was not tested |
| Multimodal matching pipeline (Callies et al., 2025) | 7,021 labeled patient-criterion pairs from 485 patients, 36 trials, and 30 sites | 87% criterion-level accuracy in retrospective validation | sponsor-developed system; concentrated indications and sites; no randomized enrollment endpoint |
| Real-world TrialGPT adaptation (Syed et al., 2026) | 149-case expert corpus and 55-record manual comparison | demonstrated that a research pipeline could be adapted to local records | small validation sets; no enrollment or patient outcome endpoint |
| AI-triggered oncology notifications (Mazor et al., 2025) | single-center randomized trial of 20,707 patients | therapeutic-trial enrollment was 2.20% with AI notification and 2.03% with usual practice | difference 0.18 percentage points, 95% confidence interval −0.25 to 0.58; notification alone did not increase enrollment |
The randomized prescreening study provides the clearest evidence of workflow acceleration because the comparator and allocation were explicit. It also illustrates why denominators matter: AI processed the queue faster, but the proportion identified as eligible did not increase (Unlu et al., 2025). The oncology notification trial extends the pathway to an enrollment endpoint and found no significant increase, showing that identifying or notifying clinicians about candidates does not remove the remaining barriers to participation (Mazor et al., 2025).
Minimum evaluation set for a matching system
A deployment report should include:
- the number of patients or records entering the system
- the number of trials searched and their recruitment status
- recall and precision at both trial and criterion level
- the percentage labeled “insufficient information”
- false-negative review, especially for safety-relevant exclusions
- time per record and total coordinator workload, including adjudication
- referral, contact, consent, enrollment, and retention denominators
- yield by disease, site, age, sex, race and ethnicity, language, insurance, geography, disability, and socioeconomic variables that can be evaluated lawfully and responsibly
- model, prompt, retrieval corpus, protocol version, and data-extraction version
- downtime, latency, cost, overrides, and reasons for investigator disagreement
A patient should not be excluded solely because the record is incomplete or the model infers likely nonadherence. Missing evidence should route to clarification or human review unless the protocol explicitly defines the missing value as disqualifying.
Site Feasibility, Operations, and Safety Surveillance
AI can estimate candidate counts, forecast accrual, prioritize sites, classify protocol deviations, extract adverse-event candidates, and predict dropout or trial duration. These tasks often use registry, EHR, claims, operational, and unstructured text data. Their usefulness depends less on internal discrimination than on calibration at future sites and times.
TrialBench illustrates the breadth of the research problem. It provides 23 AI-ready datasets across eight tasks: duration, dropout, serious adverse events, mortality, approval outcome, failure reason, eligibility design, and dose finding (Chen et al., 2025). It is a dataset and baseline suite, not evidence that these predictions improve operations.
Operational models require temporal validation because sites, standards of care, competing trials, registry practices, and sponsor behavior change. Evaluation should also separate events the model could know at the decision time from information documented after the outcome. Post-outcome notes, registry status changes, later publications, and phase-transition data create label leakage when used to simulate an earlier forecast.
For safety surveillance, high recall may be appropriate for generating a review queue, while final causality, seriousness, expectedness, and reportability remain accountable medical and regulatory judgments. False negatives require targeted review because aggregate accuracy can conceal rare but consequential missed events.
Endpoint Measurement and Outcome Adjudication
AI can support radiology, pathology, physiologic signals, digital measures, and clinical-event extraction. A valid endpoint role requires more than agreement with one reader. The model must be fit for the endpoint definition, specimen or device workflow, population, intervention, trial phase, and intended analysis.
In completed metabolic dysfunction-associated steatohepatitis trials, AI-assisted pathology has been evaluated for repeatability, reproducibility, and agreement with expert reads. A multisite validation using 1,481 biopsy cases found that AI-assisted reads improved accuracy for selected components while maintaining noninferiority for others (Pulaski et al., 2025). This supports a bounded pathology-assistance context. It does not establish that the same model is a validated surrogate endpoint, transfers to different specimen preparation, or improves patient outcomes.
LLMs can also support event preadjudication. Fu-LLM was evaluated in a secondary analysis using 1,046 telephone follow-up vignettes from three centers in a randomized trial and classified candidate deaths, hospitalizations, and medication use (Shi et al., 2025). The study supports preadjudication research, not replacement of an independent clinical events committee. Its secondary design, task-specific vignettes, human comparator, and model-development context should travel with the performance claim.
Endpoint validation questions
- Is the model output the endpoint, a component of the endpoint, or a review aid?
- Was the reference standard defined before model evaluation?
- Were readers blinded to treatment assignment and model output where appropriate?
- Are repeatability, reproducibility, missingness, and technical failure reported?
- Was performance evaluated across sites, devices, specimen preparation, demographic groups, and disease severity?
- Could treatment assignment change data quality or model behavior and create differential measurement error?
- Is the analysis based on a locked model, and how are software updates handled?
- Does the endpoint measure how a patient feels, functions, or survives, or is its relationship to patient benefit independently justified?
If AI changes endpoint ascertainment after randomization, investigators should assess whether misclassification differs by study arm. Nondifferential error can reduce power; differential error can bias the treatment effect in either direction.
Trial Outcome Prediction and the CTO Benchmark
“Clinical trial outcome prediction” can describe different targets: trial completion, primary-endpoint success, phase transition, regulatory approval, adverse events, dropout, duration, or the direction of a specified treatment effect. These labels are not interchangeable.
CTO, TOP, CTOP, and CTRP are different terms
| Term | Meaning | Correct interpretation |
|---|---|---|
| CTO | Clinical Trial Outcome benchmark | a large, dynamically constructed outcome-label dataset, not a prediction model |
| TOP | Trial Outcome Prediction benchmark used with HINT | an older benchmark for retrospective model development |
| CTOP | clinical trial outcome prediction | a generic abbreviation used in some papers and products |
| CTRP | Clinical Trial Result Prediction | a separate NLP task predicting higher, lower, or similar endpoint direction from background and PICO information |
The peer-reviewed CTO resource contains 125,840 automatically aggregated trial labels and an 11,012-trial manually curated subset. Its labeling pipeline combines trial records, publications, phase-wise tracking, news, stock movement, trial metrics, and LLM-derived interpretations. CTO labels achieved an overall F1 score of approximately 0.91 against prior expert annotations (Gao et al., 2026).
CTO addresses a real research bottleneck: historical trial outcomes are difficult to assemble at scale. It also makes label validity the central question. A statistically significant endpoint, phase progression, commercial news, sponsor stock movement, trial completion, and regulatory approval represent different constructs. A model trained on a combined label predicts that construct, not necessarily biological efficacy or patient benefit.
The dataset and code are public. The authors state that updates are planned at least annually but depend partly on OpenAI and SerpAPI credits, so reproducibility requires recording the dataset snapshot and pipeline version. Two authors disclosed cofounder or consulting relationships with a company developing AI tools for clinical research (Gao et al., 2026).
HINT is a historical model built on the TOP benchmark. It combines drug, disease, eligibility, pharmacokinetic, and historical-trial information and reported retrospective test-set F1 scores of 0.665, 0.620, and 0.847 for phases I, II, and III (Fu et al., 2022). These results do not establish prospective forecasting because the evaluation was retrospective and some representations may encode information unavailable at the historical decision time.
Minimum bar for a credible forecast
- Define the target: State whether success means endpoint attainment, completion, phase transition, approval, or another outcome.
- Freeze the prediction time: Use only information available when the forecast would have been made.
- Use temporal separation: Evaluate on future trials, not a random split that mixes eras or related development programs.
- Control relatedness: Prevent leakage across phases, indications, molecules, sponsors, publications, and duplicated registry records.
- Report calibration: A development decision needs reliable probabilities, not only ranking or F1 score.
- Compare with decision makers: Evaluate against simple baselines and the actual clinical-development process.
- Measure decision impact: Determine whether using the forecast improves decisions, costs, participant safety, or development value.
- Audit abstention and failure: Models should identify cases in which information is insufficient or outside the development domain.
Retrospective discrimination is demonstrated research performance. Reliable prospective trial forecasting remains unproven until timestamped external prediction and decision-impact evaluation are available.
Simulations, Digital Twins, and External Controls
“Digital twin” can describe mechanistic simulation, patient-level outcome forecasting, synthetic comparators, or repeated prediction from longitudinal data. These uses have different assumptions and regulatory implications. A model that reconstructs outcomes in a completed trial may be useful for calibration, but it has not necessarily demonstrated prospective prediction or a credible counterfactual treatment effect.
A defensible simulation claim should specify:
- the precise context of use, population, intervention, endpoint, and decision
- which inputs were available before the predicted outcome
- whether parameters were fitted to the trial being “predicted”
- a prespecified error tolerance and clinically meaningful failure threshold
- validation on independent trials, sites, and time periods
- performance of each important submodel rather than one aggregate match score
- uncertainty, missingness, model discrepancy, and sensitivity to assumptions
- who performed the validation and whether the result was independently replicated
FDA’s final guidance for credibility assessment of computational models in medical-device submissions and its draft framework for AI supporting drug and biologic regulatory decisions both emphasize risk and a defined context of use, although their scopes and legal status differ (FDA, 2023; FDA, 2025). An in silico trial credibility framework proposed by FDA scientists similarly evaluates the evidentiary support for important submodels rather than accepting one overall percentage (Pathmanathan et al., 2024).
A vendor statement that a simulation matched a published result within a few percentage points is insufficient without the trial, endpoint, population, prediction time, tolerance, and independent evaluator. Synthetic and external controls also require credible exchangeability, time-zero alignment, endpoint harmonization, and sensitivity analysis. Detailed sponsor-level treatment is provided in the Life Sciences AI Handbook.
Bias, Access, and Participant Protection
Trial matching can reproduce disparities even when demographic fields are excluded. Fragmented records, fewer specialist visits, limited molecular testing, inaccessible language, unstable housing, transportation barriers, and prior exclusion from research can alter what information is available and which trial appears feasible.
A 2026 controlled evaluation held clinical content constant while varying identity descriptors across 58 protocols and nine LLMs. Eligibility judgments were generally stable, but homelessness produced the largest adverse shift and stronger negative judgments about adherence and resources (Soffer et al., 2026). This was a simulated counterfactual evaluation, not a study of observed enrollment decisions. It nevertheless shows why models should apply explicit criteria rather than infer motivation, trustworthiness, or resources from social identity.
An equity audit should compare the full pathway:
| Stage | Equity question |
|---|---|
| Data availability | Who lacks complete, timely, interoperable records? |
| Model eligibility | Who is incorrectly excluded, sent to review, or ranked lower? |
| Investigator referral | Who is contacted after a model flag? |
| Access | Who can reach the site, communicate in a supported language, and complete required visits? |
| Consent | Who is offered understandable, voluntary participation? |
| Enrollment and retention | Who enrolls, remains, withdraws, or is lost to follow-up? |
| Outcomes | Are safety and efficacy measured with equivalent validity across groups? |
Automated matching should expand investigator attention, not automate assumptions about willingness or feasibility. Participant contact, consent, privacy, and protocol exceptions remain governed by the approved protocol, institutional review board, applicable law, and local policy.
Trials Evaluating AI Interventions
When AI itself is the intervention, model accuracy becomes one part of a complex intervention that includes data flow, interface design, clinician behavior, patient response, escalation pathways, and organizational capacity. The unit of randomization should match the contamination risk. Patient randomization may be reasonable for a discrete diagnostic output; clinician, unit, or facility randomization may be necessary when exposure changes practice patterns.
A pragmatic cluster-randomized trial in 16 Kenyan primary-care facilities tested an LLM-enabled clinical decision-support system. The trial enrolled 9,691 patients treated by 103 clinical officers. Fourteen-day treatment failure occurred in 2.2% of the intervention group and 2.0% of the control group, with no statistically significant difference in the primary outcome (adjusted odds ratio 0.77, 95% confidence interval 0.55–1.08; P = .13) (Agweyu et al., 2026). The trial is important because it measured patient outcomes in routine care. It did not show clinical superiority.
Design requirements for an AI intervention trial
- Intervention specification: model name and version, prompts, retrieval sources, thresholds, hardware, integration, and update policy
- Human-AI interaction: who sees the output, required training, discretion to override, and management of uncertainty
- Comparator: usual care as actually delivered, an existing decision tool, or an alternative workflow
- Clinical endpoint: patient-centered benefit and harm appropriate to the intended use
- Process endpoints: use, adherence, overrides, time, alert burden, technical failures, and workload transfer
- Safety: predefined AI-related events, escalation rules, independent monitoring, and analysis of missed or delayed care
- Equity: access, performance, treatment, and outcomes across relevant populations and sites
- Version control: a locked intervention during confirmatory evaluation or a prespecified update and analysis plan
- Implementation: contamination, learning effects, site readiness, cost, and durability after trial support is withdrawn
The Evaluating Clinical AI chapter provides the general claim-to-evidence ladder and local-evaluation framework. A randomized trial can still be biased if allocation, missing outcomes, adherence, contamination, or analysis are poorly handled.
Reporting Standards
Reporting guidance improves transparency but does not validate an intervention or remove risk of bias. The base trial guideline and the AI extension should be used together.
| Study stage or design | Primary guidance | Use |
|---|---|---|
| Protocol for an interventional trial | SPIRIT 2025 plus SPIRIT-AI | specify the intervention, workflow, inputs, outputs, errors, human role, and analysis before enrollment |
| Completed randomized trial | CONSORT 2025 plus CONSORT-AI | report allocation, intervention version, human-AI interaction, failures, outcomes, and subgroup analyses |
| Early clinical evaluation of AI decision support | DECIDE-AI | report feasibility, safety, human factors, workflow integration, and preliminary performance |
| Prediction-model development or validation | TRIPOD+AI | report data, model development, performance, validation, and applicability |
| Digital health intervention | CONSORT-EHEALTH when applicable | report platform access, engagement, delivery, and digital-intervention details |
SPIRIT-AI adds AI-specific protocol items to SPIRIT, including the intended use, input and output handling, user skills, workflow integration, and error management (Cruz Rivera et al., 2020). CONSORT-AI extends randomized-trial reporting for AI interventions, including system version, human-AI interaction, handling of unavailable or poor-quality inputs, and analysis of errors (Liu et al., 2020). DECIDE-AI addresses early, small-scale clinical studies in which safety, feasibility, and human factors must be established before confirmatory evaluation (Vasey et al., 2022).
The underlying trial standards were updated in 2025. Investigators should use CONSORT 2025 for randomized-trial reports (Hopewell et al., 2025) and SPIRIT 2025 for protocols (Chan et al., 2025), supplemented by the applicable AI extension.
Regulatory, Governance, and Data Controls
Regulatory evaluation depends on context of use and risk. FDA draft guidance for AI supporting regulatory decisions about drugs and biologics proposes a risk-based credibility framework for a specific question of interest and context of use (FDA, 2025). FDA and EMA subsequently published joint Good AI Practice principles for drug development, emphasizing human-centric design, risk-based use, data governance, performance assessment, lifecycle management, and clear communication (FDA and EMA, 2026).
For physician-investigators, the operational controls are concrete:
- define whether the system is research software, an operational aid, an endpoint tool, or part of the investigational intervention
- document data provenance, authorization, minimum necessary access, retention, and secondary use
- preserve protocol and model versions with timestamps
- log model inputs, outputs, evidence, edits, overrides, and final decisions
- prohibit silent exclusion when information is missing or the model abstains
- establish escalation for safety-critical disagreements and technical failures
- monitor drift in the patient population, protocol corpus, data pipeline, and model
- disclose sponsor, vendor, investigator, and dataset conflicts of interest
- distinguish a reporting or regulatory submission role from clinical validation or authorization
Local institutional, jurisdictional, and sponsor requirements may be stricter. FDA draft guidance is not final guidance, and neither FDA guidance nor Good AI Practice principles constitute authorization of a particular system.
Physician-Investigator Checklist
Before adopting AI in a trial, document the following:
Decision and evidence
- What exact decision or task will the system support?
- What evidence level supports that use: retrospective benchmark, external validation, prospective workflow study, randomized trial, or regulatory qualification?
- Does the study endpoint match the claim being made?
- Is the comparator credible and contemporaneous?
Data and model
- Were the data available at the time of the modeled decision?
- Are cohort construction, missingness, exclusions, and linkage errors reported?
- Could later trial phases, publications, registry changes, or outcomes leak into the inputs?
- Are the model, prompts, thresholds, retrieval corpus, and software versions fixed and recoverable?
Participant pathway
- What happens after an AI flag, nonflag, or abstention?
- Who reviews safety-critical criteria and protocol exceptions?
- Are referral, contact, consent, enrollment, retention, and outcomes measured separately?
- Are false negatives and disparities reviewed, not only aggregate accuracy?
Trial integrity
- Could AI use reveal allocation, alter outcome ascertainment, or create differential missingness?
- Are endpoint readers and adjudicators blinded where appropriate?
- Is the update policy compatible with the statistical analysis plan?
- Are adverse AI-related events and technical failures predefined and monitored?
Transparency
- Is the protocol registered and the analysis prespecified?
- Are applicable SPIRIT, CONSORT, AI-extension, and risk-of-bias tools used?
- Are code, model cards, prompts, data dictionaries, and validation artifacts available when lawful?
- Are commercial relationships and conflicts disclosed with the performance claim?
Clinical Bottom Line
AI can reduce work in trial retrieval, prescreening, measurement, and event extraction. The best current studies support selected intermediate tasks, including a randomized demonstration of faster prescreening. They do not justify autonomous eligibility decisions or a general claim that AI increases enrollment or predicts trial success.
CTO, TOP, HINT, and TrialBench make outcome-prediction research easier to reproduce and compare. Their value is methodological. Prospective timestamped forecasting, calibration, decision impact, and patient consequences remain the necessary next tests.
Continue Reading
- Clinical Research with AI for literature review, real-world evidence, causal reasoning, and general research reporting
- Evaluating Clinical AI for the claim-to-evidence ladder, prospective evaluation, and local validation
- Clinical Data and AI for provenance, missingness, leakage, and transportability
- AI in the Life Sciences: Clinical Trials for drug-development strategy, external controls, and sponsor-level decisions
- Vendor Evaluation Framework for commercial due diligence