Practice a decision.
Inspect its consequence.
AgentDrills is a bounded experiment in short debugging practice. It is free and account-free. It does not execute a model or evaluate your real agent.
What a mission does
Read a goal, a tool contract and a fictional trace. Identify a failure, select a repair, and inspect a written replay. A replay branches from the point described in its option; it is not a prediction of a real system. Optional hints explain where to look. A separate badge-creation scenario asks you to combine ideas in a new setting.
All examples and explanations are original. They are not copied from the courses or research below. The same versioned mission data used by the interface is available for inspection, including answers. This makes the practice reproducible, not a secure examination.
Why these problems?
- Anthropic: Demystifying evals for AI agents distinguishes transcripts from outcomes and describes code-based checks. That supports inspecting evidence and end states. Our fixed scenarios cannot capture the variation of real agent runs.
- Microsoft Research: Agent Skills Can Be Harmful (August 2026) reports functional failures and efficiency regressions in its benchmark analysis, including excessive verification. This motivates a bounded stopping-rule case; it does not imply that verification is generally harmful.
- Hugging Face’s Context Course combines concepts, runnable projects and quizzes. It is adjacent evidence for practice-oriented learning, with substantially more technical depth and setup than these browser exercises.
Format evidence, not proof of demand
DataCamp’s Introduction to AI Agents displays 140K+, 27 exercises, 25,555 reviews and a 90-minute duration, as checked October 9, 2026. These are vendor-reported signals for an adjacent learning format, not independently verified usage or evidence of AgentDrills retention. Its individual and business plans are adjacent monetization models; AgentDrills has no paid features, certification or commercial partnership.
What completion does not establish
A completed mission means you reached a defined replay and reviewed its lesson. It does not demonstrate educational mastery, skill improvement, real-world safety or transfer to your own systems. Feedback makes later attempts easier. The separate challenge is still a small, authored scenario, not a validated assessment.
The experiment we would test
Hypothesis: a meaningful share of consenting starters will explore three guided missions and then complete the transfer challenge without hints or retries. We have no observed product results yet.
Proposed decision rule—not an industry benchmark: review after four weeks and at least 100 non-synthetic, consented tab sessions that start a guided mission. Consider continuing if at least 25% reach three guided completions and at least 40% of those eligible sessions finish the transfer on a first unhinted pass. Below either threshold, investigate comprehension and acquisition quality before revising or stopping this format.
Low traffic is inconclusive. Consent, tab-session measurement, acquisition channel, prior knowledge and repeat visits limit interpretation. These thresholds are provisional product choices, not learning-science claims.
Measurement boundaries
Optional analytics counts a small allowlist of fixed events after consent. It does not record choices or traces. A transfer event records only two booleans: first pass without hints, and whether three guided missions were completed in that consented tab session. Only starts after consent qualify; there is no retroactive collection. Replays do not inflate per-session completion counts. Synthetic checks must be kept separate from organic traction.
Optional analytics measures aggregate practice actions on the production hostname after consent. Synthetic launch checks are excluded from traction. Search indexing and educational outcomes remain unverified.