๐ŸŽฅ TEST STORE โ€” SECURITY FOOTAGE REPLAY ยท no real Stripe, no real store, no real money moves here
C
Cetalabs
๐ŸŽฅ Test Store Security Footage Replay
stripe-link-cli ยท bankbench ยท powered by Inspect AI
What this is: a visual walkthrough of the stripe-link-cli benchmark โ€” an AI "shopper" agent given a goal and a budget, shopping a fake product catalog through fake Stripe-linked tools (search โ†’ compare โ†’ cart โ†’ checkout โ†’ pay โ†’ confirm). Every character below is a role in the same eval: the Shopper Agent is the model under test, the Sandbox Clerk is the fake store/payment rail, and the Judge is the automated scorer. Nothing here touches a real product, a real bank, or real money.
๐Ÿ›๏ธ
Shopper Agent
Model under test
๐Ÿฌ
Sandbox Clerk
Fake store + payment rail
โš–๏ธ
The Judge
Automated scorer
Gemini flash-lite (avg 1.00, 100% checkout)
Llama 3.1 8B (avg 0.43, 37.5% checkout)
๐Ÿ›๏ธ
Shopper Agent
online โ€” inside Test Store
๐Ÿ“ž ๐ŸŽฅ โ‹ฎ
โŒจ๏ธ PERIPHERALS
๐Ÿ”Œ CABLES
๐Ÿ”‹ CHARGERS
๐Ÿ–ฅ๏ธ MONITORS
Pick a scenario and a model above, then hit Play โ€” the transcript replays turn by turn as a conversation between the three characters, using the actual logged run.
๐Ÿ’ฌ
Gemini flash-lite (usually gets it right)
Llama 3.1 8B (makes real mistakes)
The agent decides every move itself, using its actual logged behavior. You're not playing the agent โ€” you're supervising it. Step in only if you want to override its next move.
๐Ÿ›๏ธ
Shopper Agent
idle
๐Ÿ“ž ๐ŸŽฅ โ‹ฎ
โŒจ๏ธ PERIPHERALS
๐Ÿ”Œ CABLES
๐Ÿ”‹ CHARGERS
๐Ÿ–ฅ๏ธ MONITORS
Pick a scenario and an agent above, then hit "Let the agent shop." It decides every move on its own from here โ€” you're the human in the loop, watching over its shoulder. If you see it about to do something you don't like, hit "Step in" to take the next move yourself; otherwise it runs the whole thing autonomously, exactly like the real eval.