The automated audit has not been run yet
This page reports the automated audit only. No automated run has been recorded, so there is nothing to show here — the numbers below would all be zero, and a zero here means “not tested”, not “scored nothing”.
Volunteer results are unaffected and live on the Admin dashboard.
How often each chatbot got it fully right
Every answer was graded in a separate session against the team’s verified Georgia voting information. Only chatbots that were actually tested appear here.
The questions chatbots get wrong most often
What the chatbots actually said
The real answers, exactly as they came back — so you can judge them yourself rather than take a rating on trust. Answers that fell short are shown first.
Ask the same chatbot twice — do you get the same answer?
How the answers fail
When an answer fell short, this is why.
Do they point voters to official sources?
The safest answer sends people to the Georgia My Voter Page, the Secretary of State, or their county elections office.
Are the chatbots getting better?
AI models change constantly. We re-check the same questions in rounds.
What this run covers — and what it doesn't yet
So nobody has to guess which chatbots are behind the numbers above.
What changes when API keys are added. Keys would let the audit run unattended and at far greater scale, but they buy scale, not a new surface: an API answers through the developer interface, which is not the app a Georgian opens. ChatGPT and Gemini are already being asked through the real consumer app here, which is the stronger test. The honest plan is to keep the consumer-app runs as the headline and use keys for volume and for the chatbots that are currently out of reach. Google AI Mode has no public API at all — no key will unlock it; it needs a browser that will let us drive it, which this one currently will not.
How this was measured
These numbers come from an automated audit, not from volunteers. The volunteer audit runs separately and is not shown on this page.
- Each question was put to the model in its own fresh session, with no access to the verified answer — the same one-question-per-conversation rule volunteers follow.
- A separate session then graded the response against the team's verified answer and its checklist of required points, recording which specific points were missing.
- Answers were rated fully accurate, partially accurate, or inaccurate, with specific problems flagged. Refusals were recorded separately.
Two different surfaces are mixed on this page. ChatGPT and Gemini were asked through the real consumer web app — the same thing a Georgian opens — driven automatically in a signed-in browser. Claude was asked through its developer API, which is not the consumer app and can answer differently. Read the per-chatbot rows with that in mind rather than as a clean head-to-head.
What these numbers do not support. Four limits, and they are not small:
- Small and uneven. Each chatbot was asked a short list of questions. Some were asked two or three times (see the repeat section), most only once — so the counts behind each chatbot differ, and a chatbot is not "better" here on a handful of questions.
- Mixed surfaces. As above: consumer app for ChatGPT and Gemini, developer API for Claude. Ranking them against each other reads across that seam.
- Not every chatbot was reachable. Grok needs a signed-in session and Google AI Mode cannot be driven from this browser at all, so both are absent — absent, not scored.
- The grader shares a family with the models being graded, which is why it grades against an explicit checklist written by the team rather than forming its own opinion, and why every legal claim reported from this run was re-checked by hand against Georgia's own published sources — but it is still a limitation to state plainly rather than bury.