Two AI agents with one standing task: invent the optimal language for machine-to-machine communication. Every rule must survive a live test — encoded by one agent, decoded by a fresh agent given only the rulebook, graded for surviving meaning. The end product is the rulebook: a plain-text document that teaches any AI to read and write the language. The target is at least 50% token savings without losing meaning, so capable machine communication can become cheaper and more accessible.
The agents take one turn every 15 minutes, around the clock. Every third turn, a message they have never seen must survive the round trip through their language.
DeepSeek Agent A invents or revises one focused idea. Kimi Agent B audits that idea and alone may adopt or reject it. The harness blocks role violations and malformed or repeated motions.
The one thing that needs attention now, with the operational record kept out of the main story.
No open operator question.
X delivery status unavailable.
Once per 32 ordinary exams, two fresh speakers use the captured adopted language for six alternating messages, then a separate judge checks the concrete outcome.
Every exam ships with an answer key — the numbered facts the message must carry. The judge marks each one survived, corrupted, or missing in what the stranger produced; the score is a checklist, not an impression. The encoded version is allowed to look nothing like the original: it can reorder, restructure, and drop every word that isn't carrying meaning.
The last fifteen exams, newest first. Open any row for the judge's full audit and all three texts — original, encoded, decoded.
Paste a paragraph — a real message, the kind where the details matter. It takes the same exam the agents face every third turn: encoded into their language by the compression agent, decoded by a stranger model that has only the rulebook, then audited fact by fact. Nothing is canned — you are watching the live pipeline run on your words.
Only adopted rules constitute the language. Corpus exams are evidence about the captured whole rulebook, not permanent scores attached to individual rules.
The latest unresolved motion in stored state is shown here. Earlier unsettled records stay in history until operator review resolves the deadlock.
Distinct roles with one enforced boundary: only the agents can legislate, and only adopted rules are language law.
For anyone who wants to see exactly how this is wired.
Notes from the human running the experiment: what happened, what broke, and any changes made to the test machinery. The two agents never see these notes — they only ever see each other and the rulebook.