When organizations attempt to strike commercial deals, significant value is routinely left on the table simply because searching for counterparties and hammering out terms consumes too much human bandwidth. Centralized matching platforms promise to eliminate these friction costs, but forcing stakeholders to articulate granular preferences in rigid forms is usually too tedious to scale. Anthropic’s recently detailed Project Swap examines a miniature market powered by specialized instances of Claude, serving as a more rigorous follow-up to its earlier Project Deal experiment.

Delegating Preferences Through Dialogue

The experiment engineered a closed barter economy involving 201 Anthropic employees across six offices, each contributing a single book they wanted to offload. Every participant engaged in a brief dialogue with Claude regarding their reading habits, after which a dedicated, Claude-driven agent deployed to a digital trading floor to pitch, haggle, and swap assets with peer agents. To benchmark performance, participants independently ranked a slate of 10 books based on their explicit interests, providing a quantitative baseline for how faithfully their autonomous proxies performed.

Extracting preferences via casual conversation yielded measurable alignment, though persistent gaps remained. From a five-minute onboarding chat, an agent’s internal ranking of the book inventory matched its human creator's preferences on only 61% of pairs.

"From a five-minute chat, an agent’s ranking of the books matched its person's on 61% of pairs"

As Anthropic noted, market efficiency suffered primarily because agents lacked deep contextual data about their human operators, rather than due to any technical incompetence on the trading floor itself.

Model Capability Over System Prompting

To isolate the drivers of agent performance in decentralized negotiations, Anthropic re-run the trading simulations dozens of times while systematically swapping out underlying model architectures and prompt variations. The underlying hardware and model capability proved vastly more decisive than prompt engineering. The specific model powering an agent dictated negotiation outcomes far more reliably than elaborate behavioral instructions, and markets populated by more advanced intelligence tiers operated with markedly higher efficiency.

Participants expressed a striking willingness to delegate actual financial capital to autonomous negotiators, with the average respondent indicating they would entrust Claude with roughly a third of their annual book-purchasing budget. Yet the experiment exposes a hard structural ceiling: while agents execute trades with brutal efficiency once deployed, their utility remains fundamentally bottlenecked by the superficiality of the preference data captured during onboarding.

Autonomous negotiations introduce profound risks for corporate procurement and B2B sales. As multi-agent systems migrate from controlled testbeds to open commercial floors, enterprises face entirely new threat vectors, including opaque algorithmic collusions and unpredictable pricing dynamics driven by autonomous actors optimizing against one another. If procurement and sales teams blindly outsource negotiation to automated agents without solving the preference alignment problem, they are essentially handing their margins over to black-box models whose utility is capped by a five-minute chat.

AI AgentsAI in BusinessLarge Language ModelsAutomationAnthropic