Vitalik Buterin CFN

  • Vitalik Buterin connects adversarial governance design with AI safety and limits on coordination among stronger AI agents.

  • Buterin says decentralization, secret ballots and privacy tools can create barriers against harmful collusion.

  • He cites blockchain forking, skin in the game and whistleblower incentives as defenses against coordinated attacks.

Vitalik Buterin says adversarial governance design could help address AI safety by limiting collusion among advanced AI agents. In a recent post, the Ethereum co-founder compared governance systems with AI environments involving less-sophisticated principals and stronger agents. He said limits on agent coordination could improve outcomes in both settings.

https://twitter.com/VitalikButerin/status/2099228441963012475?s=20

Buterin Compares Governance With AI Safety

Buterin described a shared structure between governance and AI safety. In governance, a static algorithm acts as the principal while humans operate as more-sophisticated agents. In AI safety, humans and weaker language models could serve as the principal. 

Stronger large language models would then act as the agents. Notably, Buterin focused on how agents coordinate rather than individual actions. He said governance design can produce better outcomes when systems limit how much agents can collude.

That distinction also separates useful coordination from harmful coordination. Groups can cooperate for shared goals, but some coalitions can disadvantage people outside their group.

Collusion Creates Governance Risks

Buterin cited several examples of harmful coordination, including election vote selling and price fixing. He also pointed to miners coordinating to launch a 51% attack against a blockchain.

He said actions alone cannot always reveal whether harmful coordination occurred. A seller charging a high price, for example, could act independently or coordinate with competitors.

However, rules against collusion can target the coordination itself. Buterin also noted that vote selling can create incentives that push voting systems toward plutocracy.

He connected the issue to cooperative game theory, which examines groups acting together. Buterin said some games lack stable outcomes because coalitions can repeatedly profit by changing their strategy.

Decentralization Can Limit Harmful Coordination

Buterin identified decentralization as one method for creating barriers against large-scale collusion. He also cited secret ballots, privacy tools, whistleblower incentives and internal negotiation problems.

In blockchain systems, he said forking can support counter-coordination after a harmful coalition takes control. A competing version can remove the attacking coalition’s influence while retaining most original rules.

Buterin also highlighted “skin in the game” as another defense. He said markets can make participants individually accountable for decisions.

Finally, he listed several coordination tools, including per-person voting, physical separation and role-based constituencies. He also cited Schelling points and encouraging defectors to expose planned collusion.