ARL — Agent Runtime Laboratory
An interactive, deterministic walkthrough of a single agent run — from the harness that receives the task, through context construction, tool calls and authority checks, to the human approval step. Everything runs in your browser as a teaching model.
Scope: execution is scripted, checks are calculated and authority is simulated. Nothing here establishes real-world verification or universal agent safety, and all actions stay inside the simulation.
What it covers
- Context construction, tool use, authorization and evidence evaluation.
- Nine scenarios: revenue report, prompt injection, bad evidence, temporary tool failure, persistent timeout, wrong calculation, missing evidence, context overflow and budget exhausted.
- Rewind, comparison and step-by-step inspection of the same run.
- Agent Runtime 101, plus the assumptions and evidence behind every step.