Prompt testing · without an eval stack

You are guessing which prompt actually works.

Three versions, three outputs, and a gut feeling about which one reads better. Prompt Experiment Scorecard turns that into a weighted score, a ranked winner and a decision you can show someone.

  • Score two to four variants against an editable weighted rubric.
  • Review blind: labels hidden and order shuffled so you judge the output, not the version you wrote last.
  • See exactly which criterion created the gap, and by how much.
  • Export the decision and reasoning as Markdown, JSON or CSV.
Buy Now → $19 once

Runs locally in your browser. No account, no API key, no subscription, nothing uploaded.

Webinar invite email testBlind review
Sample B
Role + audience + constraints
4.85
Winner
Sample C
B plus one-shot example
4.45
−0.40
Sample A
Plain instruction
3.35
−1.50
RecommendationShip Sample B. It clears the 75% threshold at 96.3% and leads by 0.40 weighted points — a real but modest gap, so re-check it on a fresh input before standardising.
Built for marketers running prompt tests without an engineering evaluation stack.· Local-first · Deterministic scoring · Blind review built in

From three tabs of output to one defensible call

The variant that reads best is not always the one that wins.

Reading outputs side by side rewards whichever one you wrote most recently. A rubric with weights, applied blind, removes that: the same five questions, in the same order, against every variant.

01 · WEIGH

A rubric that matches the job

Accuracy, relevance, brand voice, format compliance and actionability by default. Rename them, re-weight them, add your own. Weights must total 100% before the result is trusted.

02 · JUDGE BLIND

Hide the labels while you score

One click hides variant names and shuffles the order across the scoring panel and the side-by-side view. You score the output, not the prompt you are attached to.

03 · DECIDE

Rank, deltas and a plain-English call

Weighted totals, rank order, pass or fail against your threshold, the criteria that created the gap, and a recommendation that says outright when the margin is inside judgement noise.

Included in your one-time purchase

Everything needed to run the test and record the decision.

Unlimited saved experimentsTitle, task, audience, model, date, threshold and notes on each.
Two to four prompt variantsPrompt text, context or input, and pasted output for each.
Editable weighted rubricFive criteria to start, 1–5 scoring, per-criterion notes.
Blind review modeLabels hidden, order shuffled, scoring unaffected.
Decision logWhat changed, what it did, and the next hypothesis to test.
Markdown, JSON and CSV exportPlus validated JSON import and a print-friendly layout.
Works offlineInstallable as an app; everything stays in your browser.
Light and dark themesKeyboard-accessible throughout, responsive on phones.

Stop rerunning the same argument

One price. Every experiment you run.

Pay once and keep the tool. No seats, no usage limits, no monthly bill. Your prompts, outputs and scores never leave your device.

$19
One-time purchase · permanent access Buy Now →

All sales are final.