Prompt testing · without an eval stack
You are guessing which prompt actually works.
Three versions, three outputs, and a gut feeling about which one reads better. Prompt Experiment Scorecard turns that into a weighted score, a ranked winner and a decision you can show someone.
- Score two to four variants against an editable weighted rubric.
- Review blind: labels hidden and order shuffled so you judge the output, not the version you wrote last.
- See exactly which criterion created the gap, and by how much.
- Export the decision and reasoning as Markdown, JSON or CSV.
Buy Now →
$19 once
Runs locally in your browser. No account, no API key, no subscription, nothing uploaded.
Webinar invite email testBlind review
Sample B
Role + audience + constraints
4.85
Winner
Sample C
B plus one-shot example
4.45
−0.40
Sample A
Plain instruction
3.35
−1.50
RecommendationShip Sample B. It clears the 75% threshold at 96.3% and leads by 0.40 weighted points — a real but modest gap, so re-check it on a fresh input before standardising.
Built for marketers running prompt tests without an engineering evaluation stack.· Local-first · Deterministic scoring · Blind review built in
From three tabs of output to one defensible call
The variant that reads best is not always the one that wins.
Reading outputs side by side rewards whichever one you wrote most recently. A rubric with weights, applied blind, removes that: the same five questions, in the same order, against every variant.
01 · WEIGHA rubric that matches the job
Accuracy, relevance, brand voice, format compliance and actionability by default. Rename them, re-weight them, add your own. Weights must total 100% before the result is trusted.
02 · JUDGE BLINDHide the labels while you score
One click hides variant names and shuffles the order across the scoring panel and the side-by-side view. You score the output, not the prompt you are attached to.
03 · DECIDERank, deltas and a plain-English call
Weighted totals, rank order, pass or fail against your threshold, the criteria that created the gap, and a recommendation that says outright when the margin is inside judgement noise.
Included in your one-time purchase
Everything needed to run the test and record the decision.
✓Unlimited saved experimentsTitle, task, audience, model, date, threshold and notes on each.
✓Two to four prompt variantsPrompt text, context or input, and pasted output for each.
✓Editable weighted rubricFive criteria to start, 1–5 scoring, per-criterion notes.
✓Blind review modeLabels hidden, order shuffled, scoring unaffected.
✓Decision logWhat changed, what it did, and the next hypothesis to test.
✓Markdown, JSON and CSV exportPlus validated JSON import and a print-friendly layout.
✓Works offlineInstallable as an app; everything stays in your browser.
✓Light and dark themesKeyboard-accessible throughout, responsive on phones.
Stop rerunning the same argument
One price. Every experiment you run.
Pay once and keep the tool. No seats, no usage limits, no monthly bill. Your prompts, outputs and scores never leave your device.
$19
One-time purchase · permanent access
Buy Now →
All sales are final.