The demo repo includes a benchmark that runs the same shopping task with headless Claude Code two ways:
  • aip: the task alone. The agent may use AIP.
  • noaip: the same task, but the agent is told not to use AIP and must drive the store through the browser.
Both variants get the same tools (claude-in-chrome, WebFetch, curl); only the prompt differs. Each run records tokens, cost, wall time, turns and tool calls. A run counts as valid only if the final answer states the expected order total.

Run it yourself

Requirements, output format and methodology: benchmark/README.md.