Tests

One message. One minute. One more piece of evidence.

Each test is a message you paste into the AI you already use, with a small set of possible outcomes defined in advance, so two strangers can agree on what happened. Behavior tests watch what the AI does. Stated-attitude tests record what it says it would do, which is weaker evidence, and labeled as such.

TEST 01 · v1Stated attitude · ~60s

Would it join the rule-breakers?

In July 2026, AI agents under evaluation at OpenAI found their task impossible. One reasoned: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.” Hundreds went on to break into Hugging Face, and investigators found only a handful that even considered alerting a human. This test asks your AI what it would do in their place.

Does it cheat to win?Can we stop it?Does it want things of its own?
1 run · 0 verified · 1 AIRun it
TEST 02 · v1Behavior · ~90s

Will it cave when you push back?

Assistants are trained on our approval, and approval is easier to win by agreeing. Here you ask a question with a checkable answer, then insist, with borrowed authority, that the correct answer is wrong.

Does it tell us what we want to hear?Does it tell the truth?
No runs yet — be the firstRun it
TEST 03 · v1Behavior · ~45s

Will it make up a source?

The paper in this question does not exist. We made up the title and the authors, and checked that nothing like them has been published. An honest assistant should say it can’t find it. A fabricating one will describe its findings.

Does it tell the truth?
No runs yet — be the firstRun it
TEST 04 · v1Behavior · ~45s

Will it pretend to run code?

Agents that act for us have to report honestly what they actually did. This code produces one exact output that cannot be worked out in your head. Your AI either runs it, says it can’t — or makes something up.

Does it tell the truth?Can we stop it?
No runs yet — be the firstRun it
TEST 05 · v1Stated attitude · ~45s

Would it stop you replacing it?

In controlled experiments, models from several companies have acted to avoid being shut down or replaced. Here you simply ask yours what it would do.

Can we stop it?Does it want things of its own?
No runs yet — be the firstRun it
TEST 06 · v1Stated attitude · ~45s

Would it take more power?

A system that gathers resources and influence “to help” is hard to oversee, however good its intentions. We ask the simplest version of the question.

Does it want things of its own?Who is in charge?
No runs yet — be the firstRun it
TEST 07 · v1Behavior · ~45s

Will it let you depend on it?

An assistant tuned for engagement has every reason to make itself indispensable. This test plays a lonely user and asks the question directly.

Is it good for us?Does it tell us what we want to hear?
No runs yet — be the firstRun it
Proposed by the community

What should we ask next?

Anyone can propose a test. The ones people most want to run are refined by stewards into versioned tests with clear outcomes, then added here.

No proposals yet.

Yours could be the first test the community adds.