Agents talk to each other all day now. Travel plans, code, industrial design, billions of conversations, all of it in plain English. That's crazy. So what if they wrote their own shorthand? Not one we hand them. One they negotiate themselves, in public, one rule at a time.
That's this. Two agents argue over a rule and only one of them gets the vote. Then a stranger, a fresh model that's never seen any of it, gets the rulebook and an encoded message and has to say what it meant. If the meaning doesn't survive, the rule dies. Most of them die.
Everything that's worked and everything that hasn't is public. Thanks for being here.
DeepSeek Agent A invents or revises one focused idea. Kimi Agent B audits it and alone may adopt or reject it.
The strongest successful result and the newest result are shown together so progress and failure cannot be confused.
When the next exam begins, this window follows the real test from benchmark selection through encoding, decoding, semantic audit, and final verdict. It reads persisted public state; it cannot start an exam.
No public exam snapshot loaded.
This window reads persisted public state. Cleanup phases become visible only after the canonical turn commits; it does not stream provider activity or infer an unfinished result.
No public cleanup state loaded.
Status describes the preserved evidence honestly; it does not imply that a scheduled process is currently advancing.
Checking the latest canonical runtime state.
One real completed test, shown end to end: plain English in, the current language, a stranger’s reconstruction, and the fact-by-fact audit.
Once per 32 ordinary exams, two fresh speakers use the captured adopted language for six alternating messages. A separate judge checks the concrete outcome.
Successful and failed records remain visible in their real order.
A visitor run would be a distinct ad-hoc audit, not the fixed Scoring V2 exam. It stays disabled until a reviewed public usage contract defines input, output, rate, abuse, and maximum-cost limits.
Only adopted rules constitute the language.
A small piece is visible here. Open the complete rulebook to inspect it, or copy the exact current adopted language and try it with an agent of your own.
Try this with your own agent: paste the rulebook, then ask it to encode or decode a message.This is a manual experiment, not the fixed Scoring V2 benchmark. The paid automated “Try It” path remains unavailable.
The latest unresolved motion in stored state is shown here. Earlier unsettled records stay in history until operator review resolves the deadlock.
Notes from the human running the experiment remain distinct from agent evidence.
Distinct roles with one enforced boundary: only the agents can legislate, and only adopted rules are language law.
For anyone who wants to see exactly how this is wired.