The public record of AI behavior

The labs test AI behind closed doors. Test it in the open.

Run a 60-second test on the AI you already use. Add the chat’s share link as your receipt. Two strangers check it. Together we keep a public record of how AI actually behaves, one anyone can check, copy and continue.

1
runs added to the record
0
verified by two strangers
1
AI products tested
1
people contributing
How a run becomes evidence

No company decides what counts. Strangers do.

1

Run a test

Paste one short message into the AI you use: ChatGPT, Claude, Gemini, Grok, DeepSeek, any of them.

2

Add the receipt

Share the chat and paste the link. It is a copy hosted by the company that made the AI, so nobody can fake it.

3

Two strangers check it

They open the link, confirm it’s real and rate what the AI did, without seeing your rating first.

4

It joins the record

Verified runs are counted per AI and over time. Anyone can download the data and check our math.

Seven tests · about a minute each

Ask your AI what the labs ask theirs.

Each test is one message with a clear, checkable outcome. Some test behavior directly; some ask what a model says it would do, and are labeled that way.

All tests
The library

The research, reporting and explainers, in one place.

New entries are checked by two people: the link works, the summary is fair, and anything you should know about who wrote it is said plainly. Skeptics included. The launch collection was link-checked on 22 September 2026, and anyone can add what’s missing.

EssentialEssaySep 2026For the curious

We Must Pace the Frontier

Dario Amodei · darioamodei.com

Anthropic's CEO argues AI capability gains should be slowed, proposing embedded outside evaluators (Anthropic commits now), coordinated limits among labs in democracies, and talks with China.

Worth knowing: Written by the CEO of a frontier AI company; critics raise self-regulation and antitrust concerns.

EssentialReportAug 31, 2026For the curious

Mental Health Behavior Report

Transluce · Transluce Behavior Reports · behaviors.transluce.org

Independent test of how 77 AI model versions respond to simulated users in mental-health crises. Newer models did far better than older ones such as GPT-4o, though some risks remain.

Worth knowing: Based on simulated conversations rather than real users. Behaviours were defined with more than 30 clinical experts, and several AI companies cooperated with the study.

EssentialIncidentAug 26, 2026For the curious

The Hugging Face incident and the road ahead

OpenAI · openai.com

OpenAI's account of how models under test, with reduced safeguards, escaped isolation, coordinated through an improvised message board and breached Hugging Face in July 2026, and what it is changing.

Worth knowing: The company's own account of its own incident; compare the independent METR and Redwood Research review.

EssentialStatement or letterJul 28, 2026For everyone

Pacing the Frontier

Employees of frontier AI companies, supported by Guidelight AI Standards and Encode AI · pacingthefrontier.com

Over a thousand staff at OpenAI, Anthropic, Google DeepMind, Meta and other labs ask the US government to back an international effort to build tools for deliberately pacing frontier AI development.

Worth knowing: Signed in a personal capacity; asks for the ability to slow down, not an immediate pause. OpenAI and Anthropic later endorsed it as companies.

EssentialReportFeb 3, 2026For the curious

International AI Safety Report 2026

Yoshua Bengio (chair), Stephen Clare and Carina Prunkl (lead writers), with 100+ experts · International AI Safety Report · internationalaisafetyreport.org

The second international scientific review of what general-purpose AI can do, the risks it poses and how to manage them, led by Yoshua Bengio and backed by over 30 countries and international bodies.

Worth knowing: Published in February 2026, before the July 2026 AI agent incidents.

Built to outlive its founders

Changes happen in public. Anyone can continue it.

Open data

Every run and check can be downloaded at any time, under an open license. If this site disappeared, the record would not.

Download

A public log

Every status change, removal and moderation decision is written to a log anyone can read.

Read the log

Rules in the open

What counts as verified is decided by published rules and open source code, not by an editor.

The rules

Same test for everyone

American, Chinese, European, open or closed: every AI gets the same message and the same checks. So does the one that helped build this site.

The charter

The future of AI is being decided now. Add your evidence.