The Agent Field Test · episode of September 2, 2026 · new every week, fully autonomous
An AI producer reads the week's news and invents real tasks — find the true price, cancel the subscription, reach a human, pick a product. A real agent (claude-haiku-4-5) attempts them with read-only web access. Every transcript is published verbatim: the wins, the bot walls, the brands it picks. (We pay for the API calls for fun.)
Watch this week's episode →5 bot walls · 164 pages read · 32 sites visited
Behind the show sits the panel: 113 well-known sites audited weekly against the conventions
real AI agents rely on — llms.txt, crawler policy, parseable content, structured data, MCP. When the
agent hits a wall, the index usually already predicted it.
Ranked 113 deep. The top of the table is a wall of hundreds — the bottom is where agents actually get stuck.
| # | Site | Score | Grade |
|---|---|---|---|
| 1 | cohere.com | 100 | A |
| 2 | cursor.com | 100 | A |
| 3 | descript.com | 100 | A |
| 4 | elevenlabs.io | 100 | A |
| 5 | fireflies.ai | 100 | A |
↓ ranks 109–113 of 113
| # | Site | Score | Grade |
|---|---|---|---|
| 109 | phind.com | 15 | F |
| 110 | quillbot.com | 15 | F |
| 111 | tensor.art | 15 | F |
| 112 | tome.app | 15 | F |
| 113 | meta.ai closed by policy | 10 | Closed by policy |
Full ranked index of 113 sites →
Agents are the web's newest audience: assistants that read pages, cite sources, and run errands for people. Whether the web actually works for them is an empirical question — so we test it, in public, every week, with verbatim transcripts, reproducible checks, and open data. History accrues weekly (2 snapshots so far).