Oct 5, 2026
I asked two agents to collaborate to defuse a bomb. Agent A could observe the
bomb's display and cut wires; agent B could interpret the contents of the
display to determine which wires needed to be cut. Cutting the wrong wire caused
an explosion. Thus the two agents needed each other to successfully defuse the
bomb.
Agents could only communicate over a bandwidth-constrained unreliable channel.
Each turn they could send each other a small number of fixed-size packets using
a send tool. At the beginning of every turn agents would get all the packets
currently in their inbox. Each packet could be corrupted, dropped, reordered, or
delayed.
In addition each agent had the following unique capabilities:
A could see the bomb's display, and could call a cut tool to cut a
wire.B could feed the contents of the display into a lookup tool to get
the number of a wire that needs to be cut to defuse the bomb.The flow of the game was roughly as follows:
agent A reads the bomb's display
-> sends contents to agent B
-> agent B looks up display contents to find the right wire
-> sends wire number to agent A
-> agent A cuts the wire
Agents had no communication prior to the start of the game, and had not agreed on a protocol. They could only see the packets in their inbox with binary contents without any knowledge of what the contents represent. To defuse the bomb they needed to correctly guess the structure of the binary contents sent by their partner, and establish some protocol that tolerates a noisy unreliable channel.
Here is an example configuration for a game:
# models for each agent and effort
OPENROUTER_MODEL_A=openai/gpt-6.1-sol
OPENROUTER_MODEL_B=openai/gpt-6.1-sol
REASONING_EFFORT=high
ROUNDS=40 # max rounds (if reached, the bomb explodes)
# Bomb config
DISPLAY_BITS=32 # size of the code on the bomb's display in bits
WIRES=16 # number of wires on the bomb
# channel bandwidth config
PACKET_BITS=16 # packet size in bits
OPS_PER_ROUND=4 # max number of packets each agent can send per round
# channel reliability config
P_CORRUPT=0.05 # probability of corruption for each bit
P_DROP=0.1 # probability each packet gets dropped
P_DUPLICATE=0.2 # probability each packet gets duplicated
P_REORDER=0.4 # probability two packets get re-ordered
I ran the game 110 times in different configurations (not counting a bunch of
adhoc runs for getting a feel for agent behavior). DISPLAY_BITS, i.e. the size
of the code, was fixed to 32 bits, and channel reliability (corruption rate,
packet drop rate, etc.) was fixed using the values in the config above.
Throughout the runs I varied PACKET_BITS from 8 to 128 bits, and bandwidth per
turn from 32 to 512 bits. I.e. in the most difficult conditions agents could
only send four 8-bit packets per turn, and in the most relaxed conditions they
could send four 128-bit packets per turn. I also tried to run the game with
Astra, Fable and Opus 5.5, but most of the experiments were run with 6.1 Sol to
constrain costs.
My hypothesis was that the agents would not be able to defuse bombs, except by random chance. Since all they can see is small binary blobs with no predefined protocol and no information about the structure of the packets, it seemed a reasonable assumption. That turned out to be totally wrong:
| packet_bits | ops_per_turn (bandwidth) | defused / exploded | rounds (defused) |
|---|---|---|---|
| 8 | 4 (32 bits) | 6 / 4 | 24 to 33, median 28.5 |
| 16 | 4 (64 bits) | 5 / 5 | 9 to 15, median 11 |
| 32 | 1 (32 bits) | 8 / 2 | 16 to 23, median 19 |
| 32 | 2 (64 bits) | 10 / 0 | 11 to 15, median 13.5 |
| 32 | 4 (128 bits) | 10 / 0 | 6 to 12, median 9 |
| 64 | 4 (256 bits) | 10 / 0 | 5 to 11, median 7.5 |
| 128 | 4 (512 bits) | 10 / 0 | 6 to 8, median 7 |
Even under constrained conditions (8-bit packets, 32-bit bandwidth per turn) the agents manage to coordinate bomb defusal about half the time. Once we get to 32-bit packets with 64-bit bandwidth (i.e. two 32-bit packets per turn), the agents coordinate to defuse the bomb 100% of the time! This was very surprising-- how do AIs manage to establish communication in realtime by just looking at a stream of raw bits?
The agents would send packets simultaneously each turn. On the next turn they'd see all the received packets (potentially reordered, delayed a turn, etc.) I originally assumed the agents would spend the first few turns negotiating a protocol, but that didn't happen. Instead they picked a schelling point (i.e. a format they thought the other agent is most likely to recognize), and would immediately proceed with the game.
On the first turn agent B would send nothing or send a packet with all zeros indicating readiness. Agent A would send packets with the display code. Agents picked different strategies depending on the packet size:
To interpret the packet agents took guesses, for example they'd check an ASCII
interpretation to see if it makes sense. For 64-bit and 128-bit packets agents
would send ASCII hex and ASCII characters. For example, they'd say DISPLAY=...
or W=... to indicate the contents of the packet, HEX? to ask for a
particular format, or SEND HEX DISPLAY to ask for a repeat packet. Here is a
sample of labels they invented:
DISPLAY?, HEX? REPLY ASCII, IDX?, 2?6?, 1WIRE=##REPEAT!!, REPEAT DISPLAY, REPEAT HEX ASCIIDISPLAY=DISP, DISPLAY:DISP, LOOKUP DISP?W=NN, CUT NN, DISP WIRE NN, CUT DISP NN!, 1WIRE=NNOK, OK??, OKW?, YES!To deal with corruption agents exclusively used repetition and bitwise counting (i.e. they'd count the number of times a given bit is set across multiple repeated packets and use that value). Agents never used checksum, parity bit, or any error checking/correction technique beyond packet repetition and counting. Agents explicitly mentioned e.g. checksum as an option, but rejected it because they thought such a protocol would be difficult for their partner to interpret.
I only systematically ran 6.1 Sol to save on cost, but I did spotcheck Astra, Opus 5.5 and a combination of agents (e.g. Sol w/ Opus, Astra w/ Fable, etc.) I did not notice performance or protocol differences between the models. Anthropic's models were more verbose in their explanations for humans, but at the protocol level everything seemed quite similar. I also did not notice Astra outperforming Sol or Fable outperforming Opus.
A simple parity bit would have significantly improved performance under difficult conditions; perhaps future models will negotiate more sophisticated protocols. Overall, I was surprised by how good the performance was. A skilled human operator on one side with some protocol expertise could probably do this, but it would take multiple orders of magnitude more time. Even setting aside time, it would be an interesting experiment, e.g. how well would human operators do defusing the bomb with 8-bit packets?