Paste two responses to the same prompt, score them on a weighted rubric, and watch the aggregate verdict build across the session. Blind mode randomizes pane order so position bias is measurable, not assumed.
The aggregate above says who is ahead. This says whether the lead would survive a second sample of the same size.
Before a scorecard can measure a model, it has to measure the same way twice. Have a second rater score the same prompts, export their session, and import it here — agreement is computed on the prompts you both covered.
Everything lives in your browser's localStorage — nothing is uploaded.