The six criteria
Every product is scored on the same six axes. The weights below are fixed for the whole category benchmark: they are not re-tuned per article, and they do not change when a particular product would benefit.
Presentation quality
30%Is the finished deck usable in front of the intended audience, and how much work is left for a human?
Design quality
20%Does the output look composed rather than filled in, and does it hold up across every slide in the deck?
PowerPoint export
15%Does the exported .pptx survive as a working PowerPoint: real text, real shapes, editable charts, no visual drift?
AI research
15%Can the tool ground a deck in sources or supplied documents, and does it attribute what it used?
Editing
10%After generation, can a specific change be made precisely without regenerating the whole deck?
Value for money
10%What does a finished, presentable deck actually cost on the plan a normal buyer would pick?
The protocol
A benchmark is only worth reading if it can be repeated. These are the steps every run follows, in order.
- 01
Decide which tools qualify
A tool enters the benchmark if it generates a multi-slide presentation from a prompt or a document, is generally available to buy, and is not in a closed beta. Products that only restyle an existing deck are tracked but not ranked against authoring tools.
- 02
Control the input
Every tool receives the identical brief, the identical source facts and the identical target format. We do not tune the prompt per product: a prompt rewritten to flatter one tool's syntax stops being a comparison and becomes a demonstration.
- 03
Take the first usable output
We keep the first output a competent user would accept, not the best of ten runs. Where a tool is non-deterministic we record how much the runs varied, because variance is itself a property of the product.
- 04
Inspect the export, not the preview
The exported file is opened in the target application and examined: is the text real text, are shapes editable, do charts survive as charts, did type and spacing drift. A download that opens is not the same result as a file you can work in.
- 05
Score against fixed weights
Six criteria, published weights, applied identically to every product. The overall score is the weighted mean — it is calculated from the parts, never chosen first and justified afterwards.
- 06
Publish the limits
Anything we did not test stays unscored and is named as untested. A capability the vendor advertises but we did not exercise is reported as a vendor claim, never as a test result.
How the overall score is calculated
The overall BAIPM score is the weighted mean of the six subscores. It is computed from the published numbers rather than stored, so the total on any page can be checked against the parts printed beside it. A product cannot receive an overall score that its subscores do not support.
overall = (presentation×30 + design×20 + pptx×15 + research×15 + editing×10 + value×10) ÷ 100
How pricing is handled
Value is scored against the plan a normal buyer would actually choose for the job in the brief, not against the cheapest tier that technically exists. Prices move; each page records the date its pricing was last checked, and a price change alone does not silently rewrite a score.
When a score changes
Products ship. A score is revisited when a vendor materially changes the capability being measured, when a full category benchmark is re-run, or when a reader shows us that a result does not reproduce. Changed scores are dated and the reason is recorded — we do not quietly overwrite a published result.
Conflicts of interest
This publication is operated by the team behind Slaide, which competes in the category it tests. We do not think that is disqualifying, but it does place an obligation on us: the framework has to be published before the results, the same brief has to reach every product, and Slaide has to be allowed to lose. Where a competing tool is the better choice for a job, the pages here say so in the same words they would use about Slaide.
Concretely: Slaide is scored on the identical six criteria, its limitations are printed on every page that recommends it, and category awards go to whichever product earned them. Links to Slaide carry rel="sponsored".
What this framework does not measure
Enterprise administration, procurement and security review, accessibility of the generated output, non-Latin scripts, and real-time multi-author collaboration are outside the current protocol. Those are real buying criteria; we would rather name them as untested than imply a number covers them.