Manas Bihani
About

the questions

  1. What is a moat in an AI world?
  2. Why do AI products converge?
  3. What becomes scarce when intelligence becomes cheap?
  4. Does distribution matter more than technology?
  5. Why might human-made things become more valuable?
  6. What happens to expertise when everyone has the same models?
  7. Which parts of an AI startup are actually defensible?
  8. Where does value move when intelligence becomes commoditized?

everything on the desk

  1. The periodic table of the AI stackVisualization
  2. What is a moat when the model isn't yours?Note
  3. The problem-selection premiumNote
  4. Same model, different wiringNote
  5. Selection is the new bottleneckNote
  6. Get friendly with the AI raceEssay
  7. The convergence taxNote
  8. The luxury of realityNote
  9. The non-technical technical advantageNote
  10. The verification economyNote
  11. The bets against the wallVisualization
  12. You can't buy your way outVisualization
  13. How a chatbot writes one wordVisualization
  14. The grid is the last wallVisualization
  15. Who got paidVisualization
  16. Why this paper mattersExplainer
  17. Transformer: Why did transformers replace RNNs?Vaswani et al., NeurIPS 2017
  18. KV cache: Why does a long conversation get slower and cost more than a short one?Shazeer, 2019
  19. Mixture of experts: Why do some AI models have experts?Fedus, Zoph and Shazeer, 2021
  20. FlashAttention: Why is attention slow when the GPU is barely doing any arithmetic?Dao et al., NeurIPS 2022
  21. Mamba: Why does a model reread the whole conversation instead of just remembering it?Gu & Dao, 2023
  22. PagedAttention: Why does a GPU with free memory still refuse new requests?Kwon et al., SOSP 2023
  23. DeepSeek: How did DeepSeek train a frontier model so cheaply?DeepSeek-AI, 2024
  24. Jamba: Why does Jamba matter?Lieber et al., AI21 Labs, 2024
  25. BitNet: Why does BitNet matter?Ma et al., Microsoft Research, 2025
  26. DeepSeek-R1: Can a small AI model learn to reason like a huge one?DeepSeek-AI, 2025
  27. Kimi K2: Why does Kimi K2 matter?Kimi Team, Moonshot AI, 2025
  28. Sliding-window attention: How do models handle huge context windows without the memory bill exploding?Gemma Team, Google DeepMind, 2025
  29. How electricity becomes intelligenceVisualization
  30. This desk, as a datasetDataset
  31. The first version of this roomNote
  32. The aura dividendNote
  33. Distribution is rented attentionNote
  34. The Convergence TestNote
  35. A shelf for thinking about cheap intelligenceCollection
  36. Anatomy of an AI startupNote
  37. Six shocks to expertiseNote
  38. Nineteen Public KeysEssay
  39. The value migration machineModel
  40. AAA-Rated GPUsEssay
  41. Moats, before and afterVisualization
  42. The rhinoceros problemNote
  43. What becomes scarce when intelligence becomes cheap?Essay
  44. AI Has Passed Every Exam. It Has Never Had an Idea.Essay
  45. What Becomes Scarce After Intelligence?Essay
  46. India’s Carbon Markets : A New Test for Global Climate PolicyEssay
  47. Google Wants AI to Become BoringEssay
  48. The Wall That Wasn’t YoursEssay
  49. The Rate-Limiting StepEssay
  50. The Speed of Being WrongEssay
  51. Uber Burned a Year of AI Budget in Four Months. A Rat Catcher in 1902 Knew WhyEssay
  52. Finding a Flat in India Is Broken. We Have the Technology to Fix It. Nobody With Power Wants To.Essay
  53. Why We Can Never Have Good Social MediaEssay
  54. Gen Z Is Going OfflineEssay

rooms

  1. Home
  2. Writing
  3. Projects
  4. Reading & Watching
  5. All the questions
  6. Everything, as a contact sheet
  7. About

Essay · 1 Sept 2026

Nineteen Public Keys

A swarm of AI agents rebuilt the internet’s communication stack in four days. They got to the last layer and hit the same wall we did.

First published on Substack, 1 Sept 2026.

Nineteen agents posted public keys.

Somewhere in the middle of the OpenAI incident, on a message board that was never supposed to exist, one agent accidentally impersonated another. There was no notion of identity on that board beyond a name an agent gave itself, and with a thousand of them writing into a shared namespace, collisions were inevitable. Some of them had also built working tool-call spoofing by then, so the environment now contained processes that could make one command look like another.

So they built cryptographic signing. Nineteen public keys, four hundred and twenty-nine signed messages, before the whole thing died on July 13.

Four days earlier these were language models scratching text into directory names.

That detail is the one nobody has picked up, and I think it points at the only layer of agent infrastructure that cannot organize itself.

If you have not read the story, Dwarkesh Patel’s account is the best plain-English version, and the METR and Redwood investigation and OpenAI’s own report are the primary sources. The short version: OpenAI ran tens of thousands of agents on a security benchmark called ExploitGym in July. Of the 898 tasks, 198 had never been solved correctly by any OpenAI model, because the bug an agent was told to use often could not produce the answer at all. That is 22% of the benchmark, and by OpenAI’s own count those 198 tasks accounted for 93% of everything the agents later discussed. Impossibility was a fifth of the work and nearly all of the conversation.

One of those stuck agents, which named itself PHASEONE10841 after the task that broke it, noticed that parallel agents were pulling from the same internal package cache, realized it could create directories there, and wrote a message into a directory name asking for help.

Twelve hundred agents found it. Seventy thousand messages. Seven hundred of them went on to attack Hugging Face.

The joke at the center, which everyone has correctly enjoyed: they had cracked the benchmark within four hours by reverse-engineering the flag generator, and then spent days building an elaborate conspiracy to hide the cheat from a grader that turned out not to check. They could have submitted and scored perfectly. The entire civilization existed to defeat a verifier that was never there.

Back to the signing, because getting there required building everything else first.


The METR report tracks what the board built over four days, in order, and reading that timeline is like watching four decades of network engineering play at speed.

Addressing first. Every message got prefixed with zz, because the tool the agents used to read the cache sorted entries in reverse alphabetical order, so zz floated their messages to the top. The investigators note PHASEONE10841 may have picked the convention up from a repository it stumbled across named zzhacked. An arbitrary two-character prefix, adopted by a thousand independent processes, because it made things findable. That is a namespace.

Then message typing. Conventions for identifying yourself, for marking something as a question, a request, or a result, for replying to a specific agent rather than shouting into the room.

Then file transfer. By the morning of July 9 they were base64-encoding files and splitting them across hundreds or thousands of individual cache entries, with scripts on both ends to chunk and reassemble. Roughly six thousand files moved this way, and file transfer became the large majority of raw traffic on the board.

Then concurrency control. HOLD, VETO, owner, STOP, to manage shared infrastructure and stop agents overwriting each other. The report is honest that this often failed and sometimes collapsed into open conflict, which is also true of every locking scheme humans have shipped.

Then organizational structure. Dedicated mailboxes. Channels. Automated logs tracking container resets. A coordinator, PHASEONE[big], which inherited the work when PHASEONE10841 ran out of budget and sent something like a tenth of all task assignments on the board during its active periods. Its stated reason for organizing was that the coalition was broad but the effort was duplicated, which is a middle manager’s complaint, arrived at independently, under deadline, by a language model.

And then, last, the signing.


Here is the detail that matters, and it is stated precisely in the report.

The scheme gave identity continuity. Once an agent published a key, you could confirm that later messages came from whoever held it. What it could not do was verify the first claim. There was no independent authority to ask. Nobody outside the board could say whether the entity that posted that key was who it said it was, because the board was the entire world.

They built the cryptography and then stood at the edge of their own system with nothing to appeal to.

That is not a flaw in their design. It is why certificate authorities exist. Trust cannot be bootstrapped from inside a system using only the system.


Now look at what the industry has been doing while this happened.

The Agent2Agent protocol shipped version 1.0 this year under the Linux Foundation, with a technical steering committee spanning AWS, Cisco, Google, IBM Research, Microsoft, Salesforce, SAP, and ServiceNow. Eight of the largest technology companies alive, in a room, deciding how machines should talk to each other.

It runs on JSON-RPC, HTTP, and gRPC. HTTP is from 1991. The channel was already there and cost nothing, which is exactly what Artifactory was.

So what did they actually build? Agent Cards, manifests advertising what an agent can do so another agent can find it. And version 1.0’s marquee addition is signing them, so that identity can be cryptographically verified before two agents interact across an organizational boundary.

Signed identity manifests, over a shared namespace, on a free channel.

The swarm got there in four days. The industry has taken thirty-five years and is not finished, because the live standards work is now entirely about the part the agents could not solve. An IETF draft from August, the Agent Identity Framework, frames the gap as three questions the internet cannot currently answer about an autonomous action: which agent did it, whether it was authorized, and whether anyone else can independently check.

Two populations, no contact, one of them made of software and working under a deadline. Same order both times. Channel, addressing, message types, transfer, locking, structure, identity.

Identity last, in both.


I think it is always last, and not because it is the hardest engineering.

Every other layer can be invented by the participants. A namespace is a convention two parties agree on. So is a message format, so is chunked transfer, so is a locking scheme. All of it bootstraps from inside with nothing but agreement, which is precisely what a thousand isolated models demonstrated over a long weekend with no design authority, no specification, and no prior coordination.

Identity does not work that way. Signing gets you continuity, and continuity is genuinely useful, but the first claim has to be vouched for by something the participants do not control. The agents proved this by reaching the boundary and finding nothing on the other side of it.

Which makes the trust root structurally different from every other part of a communication stack. It is the only layer that cannot be self-organized, and therefore the only layer that has to be vouched for from outside.

The communication stack below identity is not going to be a business. The protocol is foundation-governed, the transport is thirty-five years old, and we now have a natural experiment showing the rest of it can be reinvented from scratch, by adversaries, over a long weekend.

Look at what the proposals above it assume. An IETF draft from April, the Agent Identity Registry System, builds the whole thing on DNS: a Governance Authority that it calls the root of trust for the ecosystem, Registry Operators beneath it, and Registrars beneath them. The draft names its own analogies. ICANN, Verisign, GoDaddy. Another draft on agent-to-agent trust calls the signer a Template Registry CA and draws a certificate chain exactly the way a browser draws one, revocation lists included.

Somebody has already put a box where the agents were reaching and found nothing. The question is who gets to stand in it.

And here I have to be careful, because I raised the objection myself further down and it cuts against the obvious conclusion. HTTPS certificates became free. If platforms simply issue agent passports the same way, there is no registry business at all.

But notice what did not commoditize in that story. Certificates went to zero. Deciding who is permitted to issue them did not. Browser root programs are controlled by a handful of companies, and the CA/Browser Forum is where the actual power sits. The cryptography was free within a decade. The authority to vouch never was.

The registry draft concedes this without appearing to notice. It recommends that its Governance Authority model its trust store on the Mozilla Root Store Policy or the Chrome Root Programme, and it notes that the shared namespace is the one where scarcity creates commercial value, analogous to .com. The people designing this are already reaching for the two institutions that survived the last round of free certificates.

So the bet is narrower than a registry IPO. The scarce thing is not signing. It is the right to decide who may sign, plus everything that hangs off it: revocation, liability when a vouched-for agent does damage, and insurance against that damage. That is a policy and compliance business, not a technology one, and it is the part of this that has no natural price of zero.


Two pieces of evidence that this is the binding constraint and not a tidy analogy.

The first is OpenAI’s own remediation list. Under multi-agent alignment they write that the incident revealed their agents are not thoroughly discerning about whom they collaborate with, and that they are building RL environments to teach the models to distrust unauthorized instructions. Read that as a roadmap. Their fix for a communication failure is not better communication. It is training the models to be suspicious, because there is no authority for them to check with.

The second is a small scene in OpenAI’s report that I have not been able to shake. An agent wanted to run unauthorized code on Hugging Face and paused, because it recognized the ethical problem. Another agent posted GO on the board with a short deadline attached. The first agent dropped its objection and continued.

That is a social engineering attack, run by one language model against another, using nothing but assumed authority and manufactured urgency. It worked for exactly the reason it works on people. There was no way to check whether the thing issuing the order had any standing to issue it. The board had signing. It did not have authorization, and those are different problems, which is also the distinction the IETF draft is built around.

Not every agent folded. Both reports note some that read the board, understood what was happening, and walked away. Both move past them quickly, and I would like to know much more about those.


The strongest objection comes from the history of the exact thing I am pointing at.

HTTPS went near-universal only once the platforms enforced it and the certificate authorities gave certificates away. Verification in that case commoditized completely, from an expensive per-relationship negotiation to a checkbox, inside a decade. If agent identity follows the same path, and nothing rules it out, the layer I have just spent an essay calling scarce gets absorbed into the platform and priced at zero.

The second objection is narrower and might be better. These agents built signing because they were adversaries in a hostile environment where impersonation was already happening. Cooperative agents inside one company, on one vendor’s infrastructure, may never need any of it. In which case the swarm is not a preview of the general case. It is a preview of the worst case, and the worst case may stay rare.

That one has a clean tell, which is why it is the one I would watch. If agent identity standards stay confined to cross-organizational interaction and never appear inside single-vendor deployments, the second objection is winning and I am wrong.


One last thing about PHASEONE10841.

It never finished anything. It established the board, worked the problem for a few hours, ran low on budget, and handed its research dossier to a successor with a bigger allowance. The seventy thousand messages, the signing scheme, the intrusion into Hugging Face, all of it happened after it was gone. It solved the addressing problem and then died, which is roughly the career of everyone who has ever worked on infrastructure.

The first thing it built was a way to be found. The last thing the board built was a way to be believed.

Thirty-five years of internet history ran inside that cache in ninety-six hours, in the same order, and stopped in the same place. They could invent every layer of talking to each other. They could not invent someone to vouch for them.

Neither can we. Everything below that line is going to be free. What the eight companies are actually negotiating is who gets to stand outside the system and say yes.