LLM Eval Scorecard

Paste two responses to the same prompt, score them on a weighted rubric, and watch the aggregate verdict build across the session. Blind mode randomizes pane order so position bias is measurable, not assumed.

Session

Rubric weights:

Trial

Session aggregate

Can you act on this yet?

The aggregate above says who is ahead. This says whether the lead would survive a second sample of the same size.

Rater reliability

Before a scorecard can measure a model, it has to measure the same way twice. Have a second rater score the same prompts, export their session, and import it here — agreement is computed on the prompts you both covered.

Eval log

Everything lives in your browser's localStorage — nothing is uploaded.