A series about an internal tool I built so a team could understand test results, not just produce more of them. Tests got cheap. Figuring out what the results mean did not.

  1. Beyond the test suite — why a test framework, not just another suite
  2. Rules first, LLM second — triaging hundreds of failures without fooling yourself
  3. Comparing responses across a flag — “tests passed” is not “behaves the same”
  4. Making the test framework agent-native — one backend, a UI and an agent
  5. Dashboards for the whole team — different doors for QA, developers, leads, and PMs
  6. From laptop tool to team service — restart-safe, and safe for more than one person