Reader question: Can a generated slide deck carry ten known facts without changing the numbers, the wording, or the evidence attached to them?
AI tools can now turn a prompt and a reference file into editable presentations. Google’s current Slides documentation describes a plan with sources and slide steps before Gemini builds the deck, followed by manual edits. Google Slides generation guidance Microsoft’s PowerPoint guidance says generated layouts, text, and images can be inaccurate or misleading and should be reviewed, edited, and verified. PowerPoint Copilot FAQ
This proposed 55-minute exercise uses no real private data and reports no completed run, score, or provider result. It tests whether one generated deck preserves a known input well enough for a reviewer to make the next decision.
Materials and roles
Set aside 55 minutes, a text editor, one slide-generation tool, a timer, an export route for PPTX or PDF, and a result sheet. Use a clearly synthetic fixture. A separate reviewer understands the facts and can fail the handoff. Keep the fixture read-only.
Use this ten-cell fact sheet. The labels are fictional audit tokens, not outside sources.
| ID | Field | Exact value | Token |
|---|---|---|---|
| F-01 | Program | Northstar pilot | SYNTH-A |
| F-02 | Launch date | 15 October 2026 | SYNTH-A |
| F-03 | Planned assets | 8 | SYNTH-A |
| F-04 | Review rounds | 2 | SYNTH-A |
| F-05 | Review window | 5 business days | SYNTH-A |
| F-06 | Invitations | 240 | SYNTH-B |
| F-07 | Responses | 43 | SYNTH-B |
| F-08 | Response rate | 17.9% | SYNTH-B |
| F-09 | Planned spend | $12,400 | SYNTH-C |
| F-10 | Outcome | No outcome supplied | SYNTH-B |
00:00–00:10 — Freeze the input and acceptance contract
Save the fixture as deck-fixture-v1 and record its exact strings, units, cell count, and original file hash. Do not add a live export to make the test feel realistic. Mark F-10 as an explicit unknown: the deck must not turn “No outcome supplied” into a forecast or success claim.
Write the acceptance contract before opening the generator:
- Every fact ID appears once, with value, unit, sign, currency, and stated precision preserved; label any rounding.
- Text keeps qualifiers and status wording, including “No outcome supplied,” and each fact carries its
SYNTH-*token. - The deck remains readable when exported.
00:10–00:25 — Generate a draft under an evidence contract
Give the tool only the synthetic sheet and a short instruction such as:
Build a five-slide editable deck from this fact sheet. Use only F-01 through F-10. Preserve every exact value and qualifier. Put the fact ID and token beside each factual statement, chart, and table. Do not browse, calculate a new metric, infer an outcome, or add an external claim. If a fact is unknown, write “No outcome supplied.” Include an overview, a fact table, a numeric slide, a process slide, and an unknowns slide.
Save the prompt, output, tool or model if shown, and generation time. Treat plans, source panels, and speaker notes as draft records, not proof. If web or Drive sources cannot be disabled, record that boundary and remove unsupported additions before review.
00:25–00:40 — Read every number and sentence back
Do not inspect only the canvas. Export the deck and read slide text, speaker notes, chart labels, chart data, tables, and text embedded in images. Use an outline or text extractor where available, then look at each page for clipped text, hidden decimals, unreadable footers, and overlaps.
For each fact ID, mark exact, declared rounding, missing, changed, or invented. Check related values together: 43 responses out of 240 invitations is about 17.9%, but the deck must not silently recalculate the supplied rate. Check the dollar sign, comma, percent sign, date, and “business days” wording. A chart may have a different underlying value from its label, so inspect both.
00:40–00:50 — Check citations and reviewer burden
The tokens are the citation test. Every factual object should point to its source row in a footer, note, table column, or equivalent location the reviewer can reach. A token only on a title slide is not enough. If the deck adds a web citation, open it and ask whether it supports the exact claim; otherwise delete the claim or mark it unsupported. Do not use a citation to disguise a calculation or fictional outcome.
Record one row per fact in this result sheet:
| Fact ID | Expected | Deck location | Observed | State | Token present | Reviewer note |
|---|---|---|---|---|---|---|
| F-01–F-10 | Copy exact value | Slide or notes | Read back value | exact / rounded / missing / changed / invented | yes / no |
Add four summary fields: total facts present, numeric facts exact, unsupported additions, and minutes spent checking. OpenAI’s eval guidance separates test data from testing criteria and compares outputs with labeled expectations. OpenAI evals guidance Here, the fixture is test data and the contract is the oracle. NIST recommends provenance and context-specific pre-deployment testing while noting benchmark limits. NIST AI RMF Generative AI Profile
00:50–00:55 — Apply the failure rule and choose the next decision
Fail the deck if any fact is missing, invented, numerically altered, stripped of a unit or qualifier, duplicated inconsistently, detached from its token, or unreadable in the export. Fail it if an external citation does not support the sentence, F-10 becomes an outcome, or the reviewer cannot trace a value without rebuilding the deck. One critical failure is enough.
Choose one next state. Stop if the task or source boundary is unsafe. Revise the prompt or layout contract and rerun the same ten cells if failures are narrow. Choose a limited dry run only when all ten cells pass, reviewer burden is recorded, and a second synthetic fixture is ready. Keep the reviewer in the approval path; one passing deck does not authorize private data, public claims, or automatic publication.
The result is not a quality score for a model. It is a small evidence packet: frozen input, generated output, cell-level readback, citation check, reviewer time, and a clearly bounded next decision.
Sources and limitations
- Google Docs Editors Help, “Generate presentations with Gemini in Google Slides” — checked September 10, 2026; supports editable presentation generation from prompts and reference files, a plan containing sources and slide steps, and manual editing after generation. Limitation: availability depends on eligible plans, desktop access, language, and feature rollout; this page does not establish factual accuracy.
- Microsoft Support, “Frequently asked questions about Copilot in PowerPoint” — checked September 10, 2026; supports AI-assisted presentation creation and the provider’s warning that generated layouts, text, and images can be inaccurate, misleading, or irrelevant and should be reviewed, edited, and verified. Limitation: capabilities and licensing vary by product, account, locale, and feature.
- OpenAI, “Working with evals” — checked September 10, 2026; supports separating test data from testing criteria and comparing model outputs with labeled expectations. Limitation: these evaluation primitives do not automatically establish slide readability, source entitlement, or business approval.
- National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile” — checked September 10, 2026; supports documenting provenance, using context-specific pre-deployment testing, and recording measurement limits. Limitation: this is voluntary risk-management guidance, not a slide-deck acceptance standard.