Read an A/B test in the order that protects the decision: assignment health first, then the primary metric, then guardrails and secondary metrics with a multiple-testing correction. The default load is the synthetic onboarding test onb_agent_templates, which carries a planted logging bug.
Size a test before it starts. The variance factor models CUPED: with a pre-period covariate correlated at rho, the variance and the required sample shrink by (1 - rho²).
Certified metric definitions as SQL, running in DuckDB inside the browser over a synthetic event warehouse (about 9,000 signups, 67,000 event rows, 22 weekly cohorts). Pick a metric, read its definition, edit the SQL, or ask in plain English.
Developer adoption of ElevenAPI and ElevenAgents, read from public package downloads. Downloads count CI runs and mirrors as well as people, so trends and shares carry more signal than levels.
API SDKs, weekly downloads
Agents SDKs, weekly downloads
npm migration: share of JS SDK downloads on the scoped package
Version lag, JS SDK (npm, last 7 days)
Version lag, Python SDK (PyPI, last 7 days)
Version lag is weighted by last-week downloads. A long tail on old versions means fixes and new model ids reach developers slowly, and it sets the ceiling for any feature-adoption metric that needs SDK support.
The Voice Library is a two-sided marketplace: creators share voices, users clone and generate with them. The public shared-voices endpoint exposes per-voice adoption (clones, characters generated over 7 days and 1 year, creation date). Without a key it returns 3 voices; with any ElevenLabs API key it pages through the catalogue.
The key goes from this browser straight to api.elevenlabs.io and is not stored or sent anywhere else.
What each tab is for
The synthetic warehouse
Generated in the browser from a fixed seed. It has three planted truths so the SQL can be checked against a known answer: each product line has one week-0 behaviour that drives retention, the onboarding test lifts agent deployment and slightly lowers Creative generation, and half of the Android treatment assignments are dropped by a logging bug. Plan names and prices in it are placeholders. Nothing in it is ElevenLabs data.
Tests
The statistics are checked against scipy reference values (normal and t tails, chi-square, Welch, sample size, Holm) and the generator against its planted effects: node tests/run.mjs.