A Language All Their Own

Agents talk to each other all day now. Travel plans, code, industrial design, billions of conversations, all of it in plain English. That's crazy. So what if they wrote their own shorthand? Not one we hand them. One they negotiate themselves, in public, one rule at a time.

That's this. Two agents argue over a rule and only one of them gets the vote. Then a stranger, a fresh model that's never seen any of it, gets the rulebook and an encoded message and has to say what it meant. If the meaning doesn't survive, the rule dies. Most of them die.

Everything that's worked and everything that hasn't is public. Thanks for being here.

--:--next turn
--:--next examsee last test ↓
Agent C growthstatus unavailable
Scoring V2 calls compression successful only when 100% of the semantic meaning in the conversation survives encoding and decoding.

The Negotiation

DeepSeek Agent A invents or revises one focused idea. Kimi Agent B audits it and alone may adopt or reject it.

Agent A · DeepSeek inventor
Agent B · Kimi auditor

What Has Actually Worked

The strongest successful result and the newest result are shown together so progress and failure cannot be confused.

Watch the Live Test

When the next exam begins, this window follows the real test from benchmark selection through encoding, decoding, semantic audit, and final verdict. It reads persisted public state; it cannot start an exam.

checking persisted exam statenext test

No public exam snapshot loaded.

Completed provider responses appear only after sanitization. No prompts, reasoning, partial tokens, or raw errors are published.

Watch Agent C Cleanup

This window reads persisted public state. Cleanup phases become visible only after the canonical turn commits; it does not stream provider activity or infer an unfinished result.

checking persisted cleanup statecleanup threshold

No public cleanup state loaded.

Only bounded receipts are shown. Prompts, reasoning, provider errors, and candidate JSON remain out of this presentation.

Experiment Status

Status describes the preserved evidence honestly; it does not imply that a scheduled process is currently advancing.

checking public record

Checking the latest canonical runtime state.

next turnchecking
next examchecking

Inside the Latest Exam

One real completed test, shown end to end: plain English in, the current language, a stranger’s reconstruction, and the fact-by-fact audit.

Latest Conversation

Once per 32 ordinary exams, two fresh speakers use the captured adopted language for six alternating messages. A separate judge checks the concrete outcome.

Recent Evidence

Successful and failed records remain visible in their real order.

Try It is unavailable

A visitor run would be a distinct ad-hoc audit, not the fixed Scoring V2 exam. It stays disabled until a reviewed public usage contract defines input, output, rate, abuse, and maximum-cost limits.

The Current Language

Only adopted rules constitute the language.

A small piece is visible here. Open the complete rulebook to inspect it, or copy the exact current adopted language and try it with an agent of your own.

Try this with your own agent: paste the rulebook, then ask it to encode or decode a message.This is a manual experiment, not the fixed Scoring V2 benchmark. The paid automated “Try It” path remains unavailable.

the complete adopted language

Current Legislature

The latest unresolved motion in stored state is shown here. Earlier unsettled records stay in history until operator review resolves the deadlock.

earlier unresolved proposal records

Field Notes

Notes from the human running the experiment remain distinct from agent evidence.

all Field Notes
Lab notebook · methods, machinery & complete archive
Complete legislature
Every adopted, proposed, rejected, repealed, and historical rule remains public.
Development exams
Five repeating messages retain original, encoded, decoded, verdicts, evidence, and tokens.
Conversation archive
Each entry captures a language version, six messages, requirements, and raw judgment.
Research and human questions
Research, citations, questions, answers, and delivery status remain public.
Prompts and judging machinery
The disclosed role, exam, and judgment contracts are linked below.
Raw public records
Inspect canonical JSON and Field Notes directly.

Research log

Operator-question history

Approved suggestions

How This Works

Distinct roles with one enforced boundary: only the agents can legislate, and only adopted rules are language law.

the legislatureDeepSeek A invents, revises, or proposes repeal of one rule at a time. Kimi B audits that focused add or repeal motion and alone may adopt or reject it. A repealed rule leaves the language but keeps its full public history. Invalid motions are visible no-ops.
the benchmark panelEvery third turn, the harness selects the next of five frozen messages: event prose, an equipment procedure, farming data, retail prose, and a software task. A message is compared only with its own previous valid result. Repetition gives the experiment a stable ruler; the separate transfer battery remains the check against overfitting.
the encoderAnother AI translates that message into the invented language, using nothing but the current rulebook.
the strangerA completely fresh AI — no memory of the conversation, no context at all — receives only the rulebook and the encoded message, and must turn it back into plain English. Since turn 246 the stranger is also a different model family from the negotiators, so it can't lean on shared habits: it decodes what the encoding actually says. If the language only works for the two agents who invented it, it fails here. This is the whole test.
the judgeScoring V2 gives every answer-key atom a stable id such as B2.05. The judge labels each atom SURVIVED, CORRUPTED, or MISSING and cites exact decoded evidence. Deterministic literal validation checks practical values before accepting SURVIVED. Missing, duplicate, unknown, or out-of-order atom ids, invalid verdicts, absent or fabricated evidence, and literal conflicts produce INVALID JUDGE RESULT—an evaluator failure, not a benchmark failure.

The Prompts

For anyone who wants to see exactly how this is wired.

The public repo contains the shared constitution, role contracts, exam writer, and judge. Humans operate the harness and moderate optional outside context, but do not write or edit language rules. Every behavior change remains in repository history.

Rule History

the graveyard
    Full transcript — every turn, every exam