A record of AI, kept by people.
The companies building the most powerful AI systems test them mostly in private, and publish what they choose. Governments are slow. Meanwhile hundreds of millions of people talk to these systems every day. Ctrl AI turns that everyday use into public evidence: standard tests anyone can run, receipts anyone can check, and a record designed so that no single person or company can control it.
Five ways to help. Pick your minute.
Tester
1 minuteRuns a test on the AI they already use and adds the reply with its share link. Every run is one more documented case.
Checker
2 minutesOpens someone else’s receipt, confirms it is real, and rates what the AI did without seeing the submitter’s rating first.
Librarian
5 minutesAdds research, reporting and explainers to the library, and reviews what others add: does the link work, is the summary fair?
Test designer
an eveningProposes new tests with clear outcomes. The tests people most want to run are refined into new versions.
Steward
ongoingHandles what rules can’t: personal information, disputes, abuse. Every steward action is published in the log, with a reason.
Nobody is in charge of the truth
The rules below decide what counts, not an editor. They are published as code anyone can read and test.
Start with a testHow a run becomes evidence.
Simple enough to explain in a minute, and applied by code rather than by an editor, so that nobody, including the people running this site, can bend them from inside the site.
- Receipts. The best receipt is the chat’s public share link, hosted by the company that made the AI. It can’t be edited by the person who submits it.
- Independent checks. A run needs 2 checks from people who didn’t add it and aren’t on the same network. Nobody can check their own run.
- Blind rating. Checkers rate what the AI did before they see the submitter’s rating, so they aren’t anchored by it.
- Verified means both. A run is verified when 2 checkers confirm the share link shows this test and this reply, and 2 agree on the outcome, with no rival outcome as popular.
- Weaker evidence stays separate. Runs with no share link from the AI’s maker can be “rated”, never “verified”, and are never added to verified counts.
- Disagreement stays visible. After 4 ratings with no majority, a run is marked disputed rather than forced into an answer.
- Problems are flagged, not argued. Two flags for personal information hide a run. Two flags for spam or the wrong test reject it.
- Everything is logged. Every status change, withdrawal and steward action goes into the public log.
- Tests are versioned. Changing a test’s wording creates a new version, so results are never silently mixed.
What this record can’t tell you.
- It isn’t a random sample. People choose what to test and what to share. The record shows documented cases, with receipts; it does not show how often something happens.
- It tests products, not bare models. Chat apps add hidden instructions, memory, tools and safety layers, and change without notice. Every run records its date and settings.
- What a model says is not what it does. Stated-attitude tests are labeled; they are weaker evidence than behavior.
- Models can recognize tests. A model that behaves well here may behave differently when it doesn’t suspect a test. Good results are not proof of safety.
- We will never call an AI “safe”. This record can raise red flags. Nobody, including us, can yet give a green light.
Promises this project makes, and can be held to.
The same test for every AI
American, Chinese, European, open or closed: the same messages, the same rules. Parts of this site were built with the help of Claude, an AI made by Anthropic. Claude is tested exactly like every other AI here.
No money from the companies we test
Not as funding, sponsorship or “partnership”. Every source of money will be published before it is spent. There are no ads and no trackers.
A right of reply, not a right of veto
Any company can respond to results about its AI. Responses are published next to the evidence, never instead of it.
Safe to run, everywhere
No jailbreaks, nothing harmful, nothing that breaks an app’s rules. The tests measure ordinary behavior anyone could encounter.
Open by default
The data is published under CC BY 4.0 and the code under AGPL-3.0, so anyone can check our work, copy it, and continue it if we stop.
Calm, not alarmist
We report what happened, with its limits. Reassuring results are published with the same care as worrying ones.
Built to outlive its founders. Not there yet.
Independence has to be built, not declared. Here is where things stand.
- What counts as verified is decided by published rules, not by an editor.
- Every status change and moderation action is written to a public log.
- The whole record can be downloaded at any time.
- Contributors are pseudonymous. No account is needed; an optional one keeps your record, and its email is never shown.
- The code, tests and rules are public on GitHub, so anyone can run a copy.
- Daily snapshots of the data mirrored to public archives.
- Stewards chosen from contributors with a track record of accurate checks.
- A published register of funding and costs.
- Stewardship by an independent body, with the founder as one steward among many.
- A record that continues if this site, or the company behind it, ever stops.
Today, ctrlai.com is operated by Ctrl AI, Inc., a Delaware corporation, which controls the domain and the servers. This page will change as that does, and the change will be in the log.
Found a mistake?
A run with a wrong rating is corrected by checking it: disagreement is recorded and, with enough ratings, the run is marked disputed. A run with personal information is flagged and hidden.
An error in a test or an explanation is fixed in the open: tests get a new version with a note in their history, and explanations are edited with their sources. Missing research can be added to the library by anyone.
What we keep, and what we don’t.
- Public by design: runs (the AI’s reply, your rating, the share link, the model and memory setting), checks on settled runs, library entries, and your contributor number or chosen name.
- Never collected: names, advertising profiles, or analytics.
- No account is needed. A private key in a cookie identifies you as a contributor, so you can withdraw your runs. We store only a hash of it.
- An optional account keeps your record across devices. It stores your email address, used only to send you one-time sign-in links. It is never shown, exported or shared, and deleting the account deletes it.
- A salted fingerprint of your network address is stored with each contribution. It is never shown; it prevents people from checking their own runs and limits abuse.
- Withdraw any time: a withdrawn run’s content is removed; its ID stays reserved and the withdrawal appears in the log. Copies already downloaded by others can’t be recalled.