Verify AI clinical trial summaries against publications, registry records, protocols, outcomes, participant flow, effect estimates, harms, and human review.

A Summary Is Only as Reliable as Its Evidence Chain

An AI assistant can turn a dense trial paper into fluent prose in seconds. That fluency can hide a serious failure: the draft may blend an early abstract with a later journal article, promote a secondary outcome to primary status, omit people who left the study, translate a wide confidence interval into certainty, or declare an intervention safe because the harms paragraph was short. Each sentence may sound reasonable while the chain beneath it is incomplete.

A clinical trial summary is not simply a shorter paper. It is a set of claims connected to an exact publication, trial registration, protocol, statistical analysis plan, outcome definition, analysis population, result, and publication-status check. Some links may be unavailable. Others may disagree for legitimate or concerning reasons. The reviewer’s job is to preserve those conditions, not let a model smooth them away.

This workflow is for public-facing editorial summaries of clinical research. It is not medical advice, a systematic review, peer review, regulatory guidance, or a substitute for a clinician, statistician, research-methods specialist, or governing policy. AI may organize a sanitized evidence packet and draft bounded language. It must not recommend treatment, assess an individual’s eligibility for a trial, decide that evidence is sufficient, or approve its own summary.

1. Freeze the Exact Study and Publication Version

Begin with a study identity card outside the model conversation. Record the full publication title, authors, journal or repository, DOI, PMID when available, trial-registration identifier, sponsor protocol number, publication type, version, publication date, correction date, and the date each source was checked. Add the condition, intervention, comparator, population, setting, recruitment period, planned follow-up, and sponsor exactly as the records describe them.

Do not match a trial by title or acronym alone. One registered trial can produce a protocol paper, primary-results paper, secondary analysis, subgroup analysis, safety follow-up, conference abstract, press release, and plain-language summary. Two unrelated studies can share an acronym. A preprint, accepted manuscript, and version of record can also differ. Attach every proposed claim to the exact artifact that supports it, and state when the publication is a secondary or follow-up analysis rather than the primary report.

2. Check Publication Status Before Reading for Conclusions

Start at the publisher’s current landing page, not a detached PDF or copied abstract. Follow the DOI and look for a correction, erratum, addendum, expression of concern, retraction, withdrawal, or new version. Crossref’s Crossmark service can expose publisher-supplied updates that affect interpretation or credit, including corrections and retractions. Crossref also cautions that Crossmark is a way to inspect update metadata, not a guarantee that the record is complete or correct.

For biomedical literature, inspect PubMed’s linked notices when the article is indexed there. The National Library of Medicine explains how PubMed links errata, retraction notices, retracted publications, expressions of concern, comments, and corrected republications based on information supplied by journals. Record the status, notice title, locator, and check time. Absence of a linked notice is not proof that no concern exists.

Put the draft on HOLD when a retraction makes the proposed claim unusable, an expression of concern materially affects it, a correction has not been incorporated, or the version cannot be identified. Do not quietly summarize the old wording and add “updated” from memory.

3. Name the Study Design and Respect What It Can Support

Identify the design from the methods and protocol: randomized or nonrandomized, controlled or uncontrolled, blinded or open label, parallel, crossover, cluster, single arm, superiority, noninferiority, equivalence, pilot, or another stated design. Record the comparator, allocation method, number of sites, planned and actual follow-up, and which groups were blinded. Do not infer these features from the word trial, a phase label, or an abstract’s conclusion.

The NIH guide to understanding clinical studies distinguishes well-designed randomized controlled trials, which can support causal inference under their conditions, from observational studies, which can identify associations but cannot by themselves establish cause and effect. Even randomization does not make every conclusion causal: departures from assignment, missing data, unblinding, post hoc analyses, and measurement choices still matter.

Calibrate the verbs to the design and result. “Was associated with,” “the groups differed,” and “reduced” are not interchangeable. Avoid “proved,” “works,” “prevents,” and “causes” unless the evidence and qualified review justify that exact strength.

4. Reconcile the Paper With the Registry, Protocol, and Analysis Plan

Locate the exact registry record from the identifier in the paper or protocol. On ClinicalTrials.gov, use the record history as well as the current view. The protocol data-element definitions describe fields such as study design, arms, enrollment, primary and secondary outcomes, outcome timeframes, protocol documents, and statistical analysis plans. Absence of a document is not evidence that it never existed.

Build a reconciliation table with one row per material item and columns for the earliest available plan, later amendment, registry results, publication report, source locator, date, and status. Compare population, sample size, allocation, interventions, comparators, primary and secondary outcomes, measurement method, analysis metric, timeframe, subgroup definitions, stopping rules, and planned analyses.

A difference is a question, not an automatic misconduct finding. Protocols can change for documented scientific, operational, or safety reasons. What matters for the summary is when the change occurred, whether it was declared, how it affects interpretation, and whether a qualified reviewer accepts the explanation. Keep unexplained material discrepancies visible and on HOLD.

ClinicalTrials.gov records are submitted by sponsors and investigators. Its official About page says the U.S. government does not review or approve the safety and science of every listed study. Its results quality-control guidance says NLM review looks for apparent errors, deficiencies, and inconsistencies; it does not assess the scientific design or ensure that information is truthful and non-misleading. Never describe registration or completed QC as scientific endorsement.

5. Lock Every Outcome to Its Role, Metric, and Timeframe

Create an outcome ledger before drafting conclusions. For each outcome, record its exact name, whether it was primary, secondary, other prespecified, or post hoc, the measurement variable, analysis metric, aggregation method, assessment timeframe, analysis population, and source locator. The ClinicalTrials.gov results definitions keep these elements separate and allow outcome data to be organized by arm or comparison group with measures of precision.

Do not collapse outcomes that share a friendly label. “Pain at four weeks,” “change in pain from baseline at twelve weeks,” and “proportion with a clinically defined response” answer different questions. Preserve how a composite outcome was defined. Distinguish a surrogate measure from a direct patient outcome. If the primary outcome was assessed at twelve weeks, a favorable four-week secondary result does not replace it.

Registration timing matters too. Record whether the outcome appeared before enrollment, in a later amendment, or only in the report. Do not call an analysis prespecified merely because it appears in the current registry view. Conversely, do not call a clearly labeled post hoc analysis improper solely because it was exploratory; summarize its status and limits accurately.

6. Rebuild Participant Flow, Denominators, and Missingness

Follow participants from eligibility assessment through assignment, treatment, follow-up, and analysis. Record, by arm where applicable, how many were screened, randomized or assigned, received the intended intervention, completed each relevant follow-up, entered each outcome analysis, and were excluded or lost, with the reported reasons. The denominator for one outcome may not be the denominator for another.

Preserve the analysis population exactly as reported: intention-to-treat, modified intention-to-treat, per-protocol, as-treated, safety population, complete cases, or another definition. Do not rename these groups or silently divide an event count by total enrollment. Record how missing observations were handled and whether sensitivity analyses changed the interpretation. If the publication does not explain a discrepancy between assigned and analyzed counts, mark it unresolved.

The official CONSORT 2025 guidance provides a minimum reporting set and participant-flow diagram for randomized trials, including who was randomized, received the intended intervention, and was analyzed. Use it to ask whether reporting is complete, not as a certification stamp. CONSORT does not make a trial well designed, and it should not be applied as though every nonrandomized design were a randomized trial.

7. Report Effect Estimates and Precision, Not a Significance Shortcut

For every result that enters the summary, capture the value in each group when available, the between-group contrast or effect estimate, units, direction, confidence interval or other reported precision measure, analysis population, timeframe, and prespecified analysis. Preserve adjusted and unadjusted estimates as separate results. Verify simple arithmetic independently when the inputs and method are clear, and label any derived calculation as the reviewer’s calculation rather than a reported result.

A p-value does not describe the size, practical importance, or certainty of an effect. “Statistically significant” does not mean large, clinically meaningful, replicated, or free from bias. “Not statistically significant” does not prove no effect, equivalence, or safety. A wide interval can remain compatible with materially different outcomes. Do not turn “the interval included no difference” into “the treatments were the same.”

Relative and absolute effects answer different reader questions. If both are reported or can be validly derived from matching denominators and timeframes, provide both with clear labels. Never borrow a baseline risk from another population to manufacture an absolute result. Keep rounding, multiplicity adjustments, subgroup status, and noninferiority or equivalence margins tied to the actual analysis.

8. Give Harms the Same Verification Discipline as Benefits

Build a separate harms ledger. Record how adverse events were defined and collected, the observation window, number exposed in each arm, events and participants with events, serious adverse events, withdrawals attributed to events, deaths when reported, and any source-stated severity or relationship assessment. Do not change “investigator judged related” into a proven cause, or “no significant difference” into “no risk.”

A small or short trial may be unable to detect rare, delayed, or subgroup-specific harms. Passive collection can produce a different picture from systematic questioning. If harms are absent from the abstract, inspect the full report, tables, supplements, and registry results. If they are not reported, write “not reported in the sources reviewed,” not “none occurred.”

Do not use a trial summary to tell an individual to start, stop, switch, or avoid treatment. Do not characterize an investigational intervention as approved because it was studied, or as safe because participants were monitored. The final post should direct medical decisions and trial-participation questions to qualified health professionals and the responsible study team.

9. Preserve Funding, Conflicts, and Access Conditions

Record funding sources, sponsor identity, author conflict disclosures, the sponsor’s stated role in design, data collection, analysis, writing, and publication, who had access to the data, and any reported publication restrictions. Also record registration timing, protocol and statistical-analysis-plan availability, data-sharing statement, and whether the report identifies an independent monitoring body when relevant.

A disclosed relationship is not proof that a result is false. Lack of a disclosed relationship is not proof of independence. Report the disclosure accurately and let the evidence reviewer assess how it changes confidence. Do not ask AI to infer undisclosed motives from affiliation names, writing style, or funding alone.

Separate transparency from validity. A fully disclosed trial can still have weak methods, and a sparse disclosure statement does not let an editor invent missing roles. Mark absent information as not found in the checked sources and place material uncertainty before the release owner.

10. Minimize Data Before AI Enters the Workflow

Prefer public, authoritative source documents. Do not paste participant-level data, unpublished protocols, confidential peer-review reports, internal safety discussions, private correspondence, or restricted datasets into a general AI system. Remove direct and indirect identifiers, access tokens, internal file paths, signatures, contact details that are not needed, and proprietary annotations. Preserve stable placeholders and source locators in the controlled evidence record.

Redaction is not authorization. A document may remain confidential, licensed, embargoed, or unsuitable for model processing after names are removed. Confirm the organization’s tool, privacy, copyright, security, and research-governance rules before use. Keep the reidentification key and full source packet outside the prompt.

The model should never be asked to identify trial participants, infer individual outcomes, recommend enrollment, diagnose a condition, or translate aggregate evidence into personal medical advice. If the public summary needs a plain-language explanation, draft it from verified aggregate claims and have a qualified human check both accuracy and accessibility.

11. Draft From a Claim Ledger, Not From a Document Dump

Give AI a bounded evidence ledger containing only approved fields: exact identifiers, source version, design, population, intervention, comparator, outcome role and timeframe, denominators, effect estimate and precision, harms, limitations, funding, publication status, and sentence-level locators. Mark each field as VERIFIED, INFERENCE, CONTEXT, or UNKNOWN. Tell the model to preserve those labels and use VERIFY rather than fill a blank.

After drafting, map every factual sentence back to one or more ledger entries. Check that citations actually support the full sentence, including scope, population, date, and degree of certainty. A correct citation attached to an overstated sentence does not make the sentence correct. Verify quotations against the source and avoid copying more text than the publication rules permit.

NIST’s Generative AI Profile describes confabulation, automation bias, and information-integrity risks, including confident false content and invented citations in consequential domains.

12. Require Qualified Review and a Named RELEASE or HOLD

Assign responsibilities before the draft is written. A research editor owns document identity and publication status. A methods or statistical reviewer checks design, outcomes, analyses, estimates, precision, denominators, and missingness. A clinician or subject specialist reviews health interpretation and audience safety when the summary makes clinical claims. A privacy or governance owner reviews nonpublic inputs. The publication owner decides RELEASE or HOLD for the exact version.

The release record should name the summary revision, evidence-packet revision, source check time, reviewers, unresolved limitations, intended audience, correction route, and re-review triggers. Return HOLD when identifiers do not match, the current publication status is unclear, a material registry discrepancy is unexplained, outcome role or timeframe is blurred, denominators cannot be reconstructed, benefit language exceeds the estimate, harms are omitted, restricted data entered an unapproved tool, or qualified review is missing.

Recheck after a correction, retraction, expression of concern, registry update, new protocol or analysis-plan release, additional follow-up, safety notice, regulator action, or material change to the article. A verified summary is dated evidence, not a permanent verdict.

A Worked Example: One Favorable Result Is Not the Whole Trial

Consider a fictional two-arm trial whose abstract says the intervention improved symptoms. An AI draft turns that into: “The treatment was proven safe and effective.” The identity card shows the paper is the primary-results publication and the registry identifier matches. The publication-status check finds no linked correction, but that only establishes what was found at the check time.

The reconciliation table then shows that the prespecified primary outcome was symptom change at twelve weeks. A four-week secondary outcome favored the intervention, while the twelve-week primary estimate had a wide confidence interval that included no difference and clinically important effects in more than one direction. Of 180 randomized participants, the paper’s highlighted analysis used a smaller per-protocol group. The harms table used everyone who received at least one dose and reported withdrawals in both arms. The sponsor funded the trial and had a disclosed role in analysis and writing.

The reviewers return HOLD on the AI sentence. A defensible revision says that, in this single trial, a prespecified short-term secondary outcome favored the intervention, while the primary twelve-week result remained uncertain; it states the relevant denominators and precision, summarizes reported harms without declaring safety, identifies the sponsor role, and explains that the finding should not guide an individual treatment decision. The point is not to make the result sound negative. It is to make every adjective answerable to the evidence chain.

Fifteen Questions Before Publishing

  • Do the DOI, PMID, registry ID, protocol ID, title, and authors identify the same study and exact publication?
  • Was the publisher page checked for the current version and linked notices?
  • Are correction, retraction, and expression-of-concern checks dated and recorded?
  • Is the study design named from the methods rather than inferred from promotional language?
  • Were registry history, protocol, and analysis plan compared with the report where available?
  • Is every outcome labeled primary, secondary, other prespecified, or post hoc?
  • Are measurement, metric, aggregation method, and timeframe preserved?
  • Can the reader follow participant counts from assignment to each analysis?
  • Are analysis populations, exclusions, and missing-data handling stated accurately?
  • Are effect size, units, and precision reported instead of relying on a p-value?
  • Are absolute and relative effects based on matching populations and timeframes?
  • Were harms, withdrawals, and observation windows checked with the same care as benefits?
  • Are funding, sponsor role, conflicts, and unavailable disclosures represented without speculation?
  • Were restricted data withheld from AI and every factual sentence traced to approved evidence?
  • Did qualified reviewers record RELEASE or HOLD for this exact summary?

If one answer is missing, do not ask AI to make the gap read smoothly. Keep the claim narrower, label the uncertainty, find the authoritative record, or hold publication.

Scope Notes and Authoritative Sources

These sources support evidence tracing and reporting questions; they do not produce a universal appraisal score or certify a trial, summary, treatment, or AI workflow. Registry obligations, disclosure rules, medical-writing standards, and required reviewers depend on the study, organization, audience, and jurisdiction.

Improve the Draft Without Breaking the Evidence Chain

Use AI to refine verified wording and structure while keeping source checks, clinical interpretation, privacy decisions, and publication approval with qualified humans.

Open AI Humanizer