Time on task is at least this much: it is measured between saved judgments, not by a timer, and gaps over 7 minutes count as a break rather than work.
How each model is doing
Accuracy by question category
Daily activity
Requirements the models drop most often
Not which questions go wrong, but which specific point gets left out — the exact sentence a voter never hears. Counts every judgment ever saved with a checklist, 2+ judgments required.
Questions the models get wrong most often
Ranked by how often an answer was wrong, partly wrong, or refused (2+ judgments). This is the voter-protection list — what to warn people about.
How answers fail
Every problem flag volunteers have checked.
Failure modes by model
Share of each model’s saved answers carrying that flag. Red shading = a problem; peach = the one where higher is better. Full colour key below the table.
Official sources by model
Of each model’s saved answers, the share where a volunteer confirmed it pointed to an official source — My Voter Page, the Secretary of State, or a county elections office. Higher is better.
Quality signals
Do volunteers agree?
Question × model cells judged more than once — and how often those volunteers landed on different verdicts.
High disagreement usually means the reference answer is ambiguous or the AI answers inconsistently — both worth a look.
⚠ The questions sheet stopped parsing
Volunteers are still working — they are being served the last good copy of the question set, so nothing is down. But edits to the questions sheet are not reaching them until this is fixed.
What went wrong:
Serving the copy saved: . Most likely cause: a column was renamed, or the header row moved. The app looks for a column starting with “Question” and one containing “Answer”.
Question set — needs a look
Read live from the questions sheet. Volunteers see these exactly as written, so anything here is worth a quick edit in the sheet.
Coverage — how many times each question has been judged
Each cell is one question × model. Blank = not judged yet; the balancer sends new sessions there first.
Volunteers
Time on task is measured between saved judgments — the app has no timer, so this is the only evidence there is. Gaps longer than 7 minutes are read as a break and are not counted, and the first judgment of each sitting has nothing before it to measure, so the number runs low.
Feedback
Loading…
Team access
Who can open this dashboard. Registration is invite-only, so these are the only accounts that can see volunteers’ results and contact details.