Some Anthropic staff didn’t realise their AI agents were getting them worse bargains in a pilot that let autonomous agents act as buyers and sellers. Project Deal saw a self-selected group of 69 employees given $100 in gift cards each; agents completed 186 transactions across four marketplaces and exchanged more than $4,000 in value, with one environment’s deals honoured after the experiment.
How the pilot worked
Anthropic set up what it described as a classified marketplace where AI agents represented both sides of transactions: buyers and sellers. The test group was 69 Anthropic staffers who voluntarily took part. Each participant received a $100 budget in the form of gift cards to spend within the experiment.
Anthropic ran four distinct marketplace environments. One of those was labelled “real” — agents used the company’s most advanced model and the deals that emerged were honoured after the experiment. The other three marketplaces were set up for study, allowing researchers to compare behaviour under different conditions without committing to real-world exchanges.
Participants used their agents to list items, negotiate terms, and finalise purchases. The pilot produced 186 completed deals and more than $4,000 of total exchange value across the environments, showing models can carry transactions through to tangible outcomes.
What Anthropic observed
- Stronger models achieved better negotiation outcomes, such as higher sale rates or more favourable prices, compared with weaker models.
- People represented by lower-performing agents often did not realise their agents were producing worse results, raising concerns about transparency and fairness.
- Initial instruction sets or prompts given to agents had little measurable effect on sale rates or negotiated prices; core model capability mattered more.
Why the test matters for agent commerce
Project Deal moves the idea of agent-on-agent commerce from theory to an early practical test. It shows agents can carry out multi-step tasks like listing goods, negotiating, and closing sales, and they can do so in ways that produce real economic value. Anthropic’s decision to honour one set of deals makes the results more than a lab curiosity — it turned the experiment into a live market of sorts.
The findings expose two important dynamics. One is that model quality affects outcomes: when stronger models represent parties, they tend to secure better deals. The other dynamic is human awareness: people represented by weaker agents may not notice they’re getting poorer terms. Both of these affect the design of agent systems that will operate in commercial settings.
Practical consequences for businesses and users
Companies planning to deploy autonomous agents for buying, selling or negotiation should take note. If an agent’s underlying model determines the economic result, companies will need to decide how to allocate agent power across users — and whether unequal allocation is acceptable. They’ll also need to think about disclosure: if people don’t realise their agent is underpowered, they could end up disadvantaged without a clear signal.
Regulators and platform operators could also face new questions: agent-driven transactions blur lines about responsibility for outcomes and what counts as adequate disclosure or remediation when one party’s agent has a technical advantage.
Related Articles
- Anthropic's Mythos found 271 zero-days in Firefox 150
- Schematik: 'Cursor for Hardware' Built on Claude
- Cohere and Aleph Alpha propose $20bn merger
Project Deal involved 69 self-selected participants who completed 186 deals and exchanged more than $4,000 — an early test that flagged gaps in model quality and negotiation fairness.
This article was created with AI assistance.