Evaluation setup
Purpose and limits
The ten task briefs
V1 generation system configuration
Rubric, severities and status thresholds
V1 evaluation
One task's full record, end to end: its exact input, its raw output, and both assessments, so you can see the shape of a single evaluated record. View 3 aggregates counts and patterns across all ten; this view stays at the single-record level.
Inspect one complete record
V1 dashboard
Filters
Status counts
Click a count to see which tasks contributed to it.
Affected outputs and error points by dimension
Click a row to see the subtype breakdown for that dimension, then click a subtype to see the specific tasks and annotations.
All tasks, one row each
Every task matching the current filters, with both assessments side by side. Click a task ID to open its full record in View 2.
Human evaluator vs. LLM-as-a-judge status disagreements
Every task where the human evaluator's assessment and the LLM-as-a-judge assessment reached a different status (respects the content-category filter above; not affected by the status/dimension/source filters, since a disagreement inherently compares both).