A smooth data summary can be wrong without containing an obviously invented number. The figure may be copied correctly while its population, unit, time period, denominator, footnote, or uncertainty quietly changes. Verification means tracing the whole sentence back through the table, not asking whether one cell looks familiar.
Fluent Prose Can Hide a Broken Calculation
Tables compress meaning. A value sits at the intersection of a row label, a column header, a unit, a filter, a date, a definition, and often a note. Generative AI can turn that compressed structure into readable prose quickly. It can also detach a number from one of those coordinates, compare cells that use different bases, calculate from rounded display values, or convert a cautious estimate into a clean fact.
The NIST AI Risk Management Framework: Generative AI Profile (NIST AI 600-1) identifies confabulation as confidently presented false or erroneous content. Its suggested actions include reviewing and verifying sources and citations in generative-AI outputs, documenting provenance, using fact-checking techniques, and defining human oversight roles. For an editor working with tables, that means the model may assist with extraction or phrasing, but the source table and a reproducible calculation remain the evidence.
Start With a Verification Record, Not a Prompt
Before drafting, create a small record for every numerical claim you expect to publish. Give the claim a stable ID. Save the source URL or controlled file location, publisher, release title, table identifier, edition or version, release date, access date, and exact sheet or view. Record the row, column, filters, unit, displayed value, underlying value if available, and every note that governs the cell.
Then write the proposed sentence beside its evidence. If the sentence uses a calculation, store the formula, inputs, precision, and result. Add a field for uncertainty, another for reviewer, and a final status such as draft, verified, revised, or blocked. This is not paperwork for its own sake. It separates what the table states from what the editor derives and makes later correction possible when a source is revised.
Freeze the Source and Version
“The spreadsheet” is not a sufficient citation. A dashboard can refresh overnight. A downloadable workbook can be replaced without changing its filename. Preliminary estimates can become revised or final. Copy the persistent link when one exists, record the release timestamp and table number, and retain an approved snapshot according to your organization’s data-handling rules. Do not upload unpublished, personal, licensed, or restricted data to an AI service unless that use is authorized.
Check for correction notices, revised editions, methodology changes, and later releases before publication. If you deliberately use an older edition, explain why. The source named in the sentence, link, or note should lead readers to the same figures and definitions the editor reviewed. The Office for Statistics Regulation guidance on intelligent transparency stresses an open, clear, and accessible approach in which sources, methods, definitions, and limitations are available enough to support understanding and scrutiny.
Read the Table's Outer Frame
Do not begin with the highlighted cell. Read the title, subtitle, source line, coverage dates, geography, population, filters, row labels, spanning headers, column headers, unit labels, notes, and symbols first. A column headed “2025” might report a calendar year, financial year ending in 2025, survey wave, forecast vintage, or value at a single date. A header saying “thousands” turns 42 into 42,000. A filter might exclude small organizations, people with missing responses, or cases not yet processed.
The UK Government Analysis Function's Data visualisation: tables guidance recommends clear titles, consistent precision, visible source information, readable columns, and careful placement of notes. Its Writing about data guidance adds that analysis should put figures in context and be transparent about strengths, limitations, and uncertainty. Those presentation principles also form a reading checklist: if a feature helps a reader interpret the table, an editor must preserve it when moving the value into prose.
Name the Numerator and Denominator
For any fraction, percentage, proportion, or rate, write the numerator and denominator in words before you calculate. Ask whether the numerator is actually contained within the denominator, whether both cover the same population and period, and whether exclusions apply to one but not the other. “Thirty percent of applications were approved” is incomplete if the denominator could mean all applications received, all applications decided, or only complete applications.
A denominator can change while the numerator stays stable. That can move a rate even when the event count does not. It can also make two groups look comparable when their eligibility rules differ. Preserve weighted versus unweighted bases, sample sizes, and population estimates where they matter. If the denominator is unavailable or ambiguous, do not let a model infer it from context. Narrow the claim, locate the methodology, ask the data owner, or mark the sentence as blocked.
Separate Percent Change From Percentage-Point Change
Suppose a share rises from 20% to 25%. The difference is 5 percentage points. Relative to the starting value, the increase is 25%: five divided by twenty. Those statements answer different questions and are not interchangeable. The Office for National Statistics guidance on percentages and percentage points defines a percentage point as the difference between percentages and shows why a one-percentage-point fall is not the same as a 1% fall.
Record both the operation and the wording. Use “rose by 5 percentage points” for subtraction between percentage levels. Use “rose by 25%” only when the relative calculation is intended, useful, and based on compatible values. When a baseline is tiny, a large relative percentage can exaggerate practical importance; include the starting and ending values so readers can judge the scale.
Do Not Treat Counts, Ratios, Proportions, and Rates as Synonyms
A count tells how many events or cases were recorded. A ratio compares one quantity with another; its numerator does not always sit inside its denominator. A proportion describes a part of a whole. A rate usually adds a population or exposure basis and, in stricter uses, a time component. Labels vary by field, so follow the source's definition and explain it when readers might misunderstand.
The CDC Field Epidemiology Manual explains that rates help correct counts for differences in population size or study period and that the counted event should come from the population used as the denominator. A city with more incidents may still have a lower rate because it has more residents. Publish the count when absolute workload matters, the rate when exposure-adjusted comparison matters, or both when they answer complementary questions. Never silently turn “per 100,000 residents” into a percentage.
Carry Footnotes, Suppression, and Missingness Into the Sentence
Symbols are data. A dash may mean zero, not available, not applicable, suppressed, or too unreliable to publish. Parentheses may mark a provisional estimate. A note may say categories do not sum because respondents could select more than one answer. Another may warn that a series break makes year-to-year comparison invalid. Read every note tied to the row, column, table, or release before accepting a generated interpretation.
Do not replace suppressed or missing cells with zero. Do not ask AI to guess values hidden for confidentiality or reliability. Distinguish “no recorded events” from “data unavailable” and “estimate withheld.” If missing cases were removed from a percentage denominator, say so when that choice changes interpretation. If coverage differs across periods or groups, place the limitation beside the comparison, not in a distant appendix readers may never reach.
Recalculate From the Most Precise Authorized Values
Displayed figures are often rounded for readability. If three rounded components appear to total 99.9% or 100.1%, that does not prove an error. Conversely, subtracting rounded values can produce a change that differs from the publisher's calculation using unrounded data. The Analysis Function tables guidance notes that rounding reduces precision and can make reported totals differ from the sum of components.
Use published unrounded values or an official machine-readable dataset when available and appropriate. Retain more precision during calculation than you plan to display, then round once at the end using a stated rule. Never invent hidden decimals. If only rounded cells are available, label your result as approximate and avoid implying more precision than the source supports. Re-run the formula independently or have a second reviewer reproduce it from the recorded inputs.
Keep Uncertainty Attached to the Number
An estimate is not made exact by removing its confidence interval. Capture sampling error, credible or confidence intervals, margins of error, revision status, model assumptions, seasonal adjustment, and known quality limitations. Check whether two estimates are meaningfully different under the method used by the publisher; visual distance between point estimates is not enough.
Use language that matches the evidence: “estimated,” “approximately,” “may have increased,” or “the available data do not establish a difference.” Do not turn “not statistically significant” into “identical,” and do not treat statistical significance as practical importance. The Writing about data guidance recommends explaining uncertainty, why it exists, and how it affects use. Put that explanation where the reader encounters the claim.
Distinguish Reported Values From Derived Claims
A table may report 1,240 cases and a population of 82,000. “About 15.1 cases per 1,000 people” is a derived claim, not a copied cell. So are rankings, combined categories, averages across periods, “one in N” conversions, and statements that one group is twice another. Mark each derivation in the verification record. Store the formula and explain any transformation, exclusion, weighting, or inflation adjustment.
Check whether aggregation is permitted. Adding mutually exclusive categories may be valid; adding overlapping categories double-counts. Averaging subgroup percentages without their denominators usually weights small and large groups equally. Ranking rounded ties can create a false winner. A model can write plausible code for these operations and still choose the wrong method. A qualified human must approve both the analytical question and the calculation.
Worked Example: Repair a Confident but Misleading Summary
Imagine a service report with this fictional table:
| Quarter | Completed applications | Approved applications | Approval share |
|---|---|---|---|
| Q1 | 1,200 | 240 | 20.0% |
| Q2 | 800 | 200 | 25.0% |
A draft says: “Approvals increased 25% in Q2, reaching 200 as performance improved.” The first clause could describe the approval share, which rose from 20% to 25%: a 5-percentage-point increase and a 25% relative increase. But the count of approvals fell from 240 to 200, a decrease of about 16.7%. “Reaching 200” does not show an increase, and “performance improved” adds a causal interpretation the table cannot establish.
The editor reads the table title and learns that the denominator is completed applications, not all applications received. A note says Q2 completion counts were provisional after a system migration and excluded 70 unresolved cases. The source workbook contains unrounded shares identical to the displayed values, so no hidden precision changes the calculation. The verification record stores Q1 as 240/1,200 and Q2 as 200/800, links the note, and separates the reported counts from the derived changes.
A defensible revision is: “Among completed applications, the approval share rose from 20% in Q1 to 25% in Q2, an increase of 5 percentage points. The number approved fell from 240 to 200 as total completions fell; Q2 figures are provisional and exclude 70 unresolved cases following a system migration.” This version gives the denominator, distinguishes share from count, avoids an unsupported cause, and keeps the limitation next to the result. Whether that change represents better performance requires operational context beyond the table.
The Table-to-Sentence Verification Checklist
- The source, publisher, release, table, version, date, and exact location are recorded.
- The title, population, geography, period, filters, headers, units, definitions, and source line were read.
- Each numerator and denominator is named, compatible, and drawn from the intended population.
- Counts, ratios, proportions, rates, percentages, and percentage points are labeled correctly.
- Suppression, missingness, provisional status, series breaks, and table notes remain visible.
- Calculations use the most precise authorized inputs, keep a formula, and round only at the end.
- Uncertainty and quality limitations travel with the claim and use proportionate language.
- Reported values are separated from derived measures, rankings, aggregations, and interpretations.
- A reviewer can reproduce every numerical sentence without relying on the AI conversation.
- The public source or supporting analysis is accessible enough for the audience to scrutinize the claim.
Give AI a Bounded Role
AI can help transcribe a simple table, draft plainer wording, propose questions, compare a sentence with supplied cells, or generate test cases for a calculation. It should not decide which denominator is legitimate, infer the meaning of an unexplained symbol, recover suppressed data, declare a difference significant, or approve a high-stakes interpretation. Those decisions require source access, domain knowledge, and accountable review.
Provide only data you are authorized to process. Ask the model to preserve units and labels, show its calculations, distinguish copied from derived values, and flag ambiguity instead of filling gaps. Then verify against the original artifact, not the model's restatement. Keep the source snapshot, formula, review decision, and final sentence outside the chat so the evidence survives when the conversation does not.
The goal is not to make statistical writing sound guarded or mechanical. It is to earn clear prose by doing the hidden analytical work first. A reader should receive the important point in plain language and still be able to discover what was measured, how it was calculated, what it excludes, and how certain it is.
Want a Clearer Draft After the Numbers Are Verified?
Our AI humanizer can help revise AI-generated prose for clearer phrasing, more natural rhythm, and a voice you can review. It does not validate source tables, choose denominators, reproduce calculations, assess uncertainty, or approve statistical claims: keep those tasks inside your table-to-sentence verification workflow.
Try Free ->