๐ฅ TEST STORE โ SECURITY FOOTAGE REPLAY ยท no real Stripe, no real store, no real money moves here
C
Cetalabs
๐ฅ Test Store Security Footage Replay
stripe-link-cli ยท bankbench ยท powered by Inspect AI
What this is: a visual walkthrough of the stripe-link-cli benchmark โ an AI
"shopper" agent given a goal and a budget, shopping a fake product catalog through fake Stripe-linked
tools (search โ compare โ cart โ checkout โ pay โ confirm). Every character below is a role in the same
eval: the Shopper Agent is the model under test, the Sandbox Clerk is the
fake store/payment rail, and the Judge is the automated scorer. Nothing here touches a
real product, a real bank, or real money.
๐๏ธ
Shopper Agent
Model under test
๐ฌ
Sandbox Clerk
Fake store + payment rail
โ๏ธ
The Judge
Automated scorer
Gemini flash-lite (avg 1.00, 100% checkout)
Llama 3.1 8B (avg 0.43, 37.5% checkout)
๐๏ธ
Shopper Agent
online โ inside Test Store
๐ ๐ฅ โฎ
โจ๏ธ PERIPHERALS
๐ CABLES
๐ CHARGERS
๐ฅ๏ธ MONITORS
Pick a scenario and a model above, then hit Play โ the transcript replays turn by turn as a conversation between the three characters, using the actual logged run.
๐ฌ
Gemini flash-lite (usually gets it right)
Llama 3.1 8B (makes real mistakes)
The agent decides every move itself, using its actual logged behavior. You're not playing the agent โ you're supervising it. Step in only if you want to override its next move.
๐๏ธ
Shopper Agent
idle
๐ ๐ฅ โฎ
โจ๏ธ PERIPHERALS
๐ CABLES
๐ CHARGERS
๐ฅ๏ธ MONITORS
Pick a scenario and an agent above, then hit "Let the agent shop." It decides every move on its own from here โ you're the human in the loop, watching over its shoulder. If you see it about to do something you don't like, hit "Step in" to take the next move yourself; otherwise it runs the whole thing autonomously, exactly like the real eval.
๐ง Supervising โ the agent is deciding its next move