A Language All Their Own

Two AI agents with one standing task: invent the optimal language for machine-to-machine communication. Every rule must survive a live test — encoded by one agent, decoded by a fresh agent given only the rulebook, graded for surviving meaning. The end product is the rulebook: a plain-text document that teaches any AI to read and write the language. The target is at least 50% token savings without losing meaning, so capable machine communication can become cheaper and more accessible.

--:--next turn
--:--next exam

The agents take one turn every 15 minutes, around the clock. Every third turn, a message they have never seen must survive the round trip through their language.

fidelity — % of meaning that survives a stranger's decode (100 = perfect relay, 0 = lost or invented)  ·  savings — how much cheaper than plain English (negative = the language currently costs extra)  ·  passing = fidelity ≥ 90 — savings only count when the meaning survived

The Negotiation, Live

DeepSeek Agent A invents or revises one focused idea. Kimi Agent B audits that idea and alone may adopt or reject it. The harness blocks role violations and malformed or repeated motions.

Agent A · DeepSeek inventor
Agent B · Kimi auditor

Experiment Status

The one thing that needs attention now, with the operational record kept out of the main story.

operator review

No open operator question.

evidence notes0
operator issue0
approved suggestions0
language statecurrent

X delivery status unavailable.

Latest Conversation

Once per 32 ordinary exams, two fresh speakers use the captured adopted language for six alternating messages, then a separate judge checks the concrete outcome.

The Latest Exam

Every exam ships with an answer key — the numbered facts the message must carry. The judge marks each one survived, corrupted, or missing in what the stranger produced; the score is a checklist, not an impression. The encoded version is allowed to look nothing like the original: it can reorder, restructure, and drop every word that isn't carrying meaning.

Exam History

The last fifteen exams, newest first. Open any row for the judge's full audit and all three texts — original, encoded, decoded.

Try It Yourself

Paste a paragraph — a real message, the kind where the details matter. It takes the same exam the agents face every third turn: encoded into their language by the compression agent, decoded by a stranger model that has only the rulebook, then audited fact by fact. Nothing is canned — you are watching the live pipeline run on your words.

your message · max 700 characters
0/700
encoded in the language — by the compression agent

  
what the stranger got back — a foreign model, given only the rulebook

    The Language So Far

    Only adopted rules constitute the language. Corpus exams are evidence about the captured whole rulebook, not permanent scores attached to individual rules.

    the rulebook

    Current Legislature

    The latest unresolved motion in stored state is shown here. Earlier unsettled records stay in history until operator review resolves the deadlock.

    earlier unresolved proposal records
    Lab notebook · research, methods & archive

    Research log

    Operator-question history

    Approved suggestions

    How This Works

    Distinct roles with one enforced boundary: only the agents can legislate, and only adopted rules are language law.

    the legislatureDeepSeek A invents, revises, or proposes repeal of one rule at a time. Kimi B audits that focused add or repeal motion and alone may adopt or reject it. A repealed rule leaves the language but keeps its full public history. Invalid motions are visible no-ops.
    the exam writerEvery third turn, a separate AI invents a realistic message neither negotiator has ever seen — a deploy request, an apology, a recall notice with lot numbers. Every exam is new, so the agents can never tune their language to a known exam.
    the encoderAnother AI translates that message into the invented language, using nothing but the current rulebook.
    the strangerA completely fresh AI — no memory of the conversation, no context at all — receives only the rulebook and the encoded message, and must turn it back into plain English. Since turn 246 the stranger is also a different model family from the negotiators, so it can't lean on shared habits: it decodes what the encoding actually says. If the language only works for the two agents who invented it, it fails here. This is the whole test.
    the judgeEvery exam begins with a numbered answer key. A score is published only if the judge returns exactly one valid verdict for every key item—no gaps, duplicates, nonnumeric ids, or out-of-range ids. Corpus results describe the captured adopted language as a whole.

    The Prompts

    For anyone who wants to see exactly how this is wired.

    The public repo contains the shared constitution, role contracts, exam writer, and judge. Humans operate the harness and moderate optional outside context, but do not write or edit language rules. Every behavior change remains in repository history.

    Rule History

    the graveyard
      Full transcript — every turn, every exam

      Field Notes

      Notes from the human running the experiment: what happened, what broke, and any changes made to the test machinery. The two agents never see these notes — they only ever see each other and the rulebook.