Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service

My nightly run went red, and my tooling confidently filed most of the failures under one cause. It was the wrong cause. The tests had done their job by failing; nothing, human or machine, could tell me why without digging. (The full story is in article 2.)

I now think that gap is where QA engineering is going to live.

The thesis: tests got cheap, understanding didn’t

Writing a test used to be the expensive part. You read the docs, craft the request, assert the response, wrangle the data. Today an AI will draft that in seconds, and it is usually decent.

So the bottleneck moves. My Playwright API suite has 2,300+ tests. Every night it produces a pile of results, and the pile is the real product. The questions that matter are not “can we write more tests?” but:

  • Which of these hundreds of failures are one problem wearing a different hat each?
  • Is this failing test new, or has it been red for a while?
  • Do we actually cover this endpoint?
  • Does the migrated code behave like the old code?
  • Who needs to know about this, and in what form?

Cheap tests mean abundant results. Abundant results with no understanding layer are just noise with a CI badge.

So I built the understanding layer. I call it the test framework.

How it grew (it was not planned)

I would love to say I designed a platform. I did not. It accreted, one annoyance at a time.

  1. A report on a shared page. Playwright already makes a good HTML report. I published it somewhere the team could open, with several runs visible by build and date. “Send me the report” stopped being a sentence.
  2. An HTML site around the reports. People wanted history and trends, not just the latest run, so I wrapped the reports in a small site.
  3. A React application. Every new idea (coverage, performance, a runner, an analyzer) needed a home. A pile of static pages stopped scaling, so the site became an app.
  4. Scripts became tools. A command-line “AI failed-test analyzer” I ran by hand became a UI tool. A standalone script that listed first-failed assertions was later replaced by a test framework tool too.
  5. From my laptop to a container. It lived on my machine for far too long. Eventually it became a hosted container on the company’s cluster, with state kept in a cloud object store so a restart loses nothing.
  6. An MCP server on top. The same capabilities, exposed so an AI agent can drive them in conversation.
flowchart TD
    A["Shared report"] --> B["History site"]
    B --> C["React app"]
    C --> D["Scripts become tools"]
    D --> E["Hosted container"]
    E --> F["MCP server"]

Today it is about eleven tools in the UI: a React 19 + Vite + TypeScript + Tailwind front end of roughly 40k lines, and a Node + Express back end of roughly 17-20k lines, run directly by a TypeScript runner with no compile step. It sits beside the suite, and when it needs to run tests, the server shells out to Playwright.

Why build instead of subscribe

There are commercial products for test analytics. Here is how I see the trade.

A subscription gives you someone else’s idea of a dashboard. It is a good idea, for the average customer. My questions are oddly specific: “which tests are tagged as known bugs while their ticket is already resolved?” No vendor ships that, because no vendor has my data model.

Building gives you your questions, answered with your data. And AI changed the economics: adding a feature is now a matter of hours, not a sprint.

The honest costs: you maintain it, and you must keep scope tight. My server grew into a monolith of several thousand lines, and I am splitting it into route modules, which is the normal growing pain of anything that gets used. I will take that bill over a renewal invoice for a dashboard that can’t answer my questions.

The rule that shaped everything: the third time, build a tool

When I catch myself asking an AI to produce the same report a third time, I turn it into a tool. Not because prompts are bad, but because a recurring analysis done as a prompt:

  • re-spends tokens every time, on work that didn’t need a model;
  • varies from run to run;
  • can’t be trusted to count, since models are fluent, not arithmetically honest;
  • is slow.

As a tool, the same analysis is deterministic, cached, fast and testable. The model keeps the part it is good at: judgement. And an AI agent can call the tool, so the thing I stopped asking the AI to do becomes something it does better. (Article 4 is the full story.)

flowchart LR
    A["Prompt you keep retyping"] --> B["Deterministic tool"]
    B --> C["The AI calls it"]

A tour of the tools

Grouped by the job they do rather than as a list of eleven names. Articles 2 to 5 go deep on most of these.

”Why is it red?” Understanding failures

The AI failure analyzer is the flagship. Upload a run, or pull one from CI, with hundreds of failures. It groups them into root causes, runs a deeper analysis per group, and offers a trace viewer (requests, responses, a cURL), a copyable re-run command and a ticket text generator. A “context” box lets you tell the AI what you already know; what you write there is treated as authoritative. Deterministic rules run first and the LLM only sees the leftovers (article 2).

The failure report is the plain sibling: the first failed assertion per failed test in a sortable table, with CSV export. No AI. Sometimes a table is the right answer.

Test performance is the landing page: duration, pass rate and flakiness from the nightly runs, a failure calendar, a persistent-failures panel, and a base-run versus compare-run view. When older data lacks a number, it says “unknown” rather than showing a confident, wrong zero. The suite’s pass rate sits in the low-to-mid 90s percent and the flakiness target is under 5%; this page is how I know where we stand.

API coverage shows which endpoints have tests and where the gaps are, by domain, with history.

The bug tracker reconciles the @bug tags in the tests with tickets in the ticketing system and the latest results, and digs up stale tags: ticket resolved, tag still there.

”Did the migration change anything?” Comparison

Response comparison runs the same suite with a feature flag on and off, captures every request/response pair, and diffs them in two passes: structure, then values. “The tests pass both ways” is a much weaker claim than “the responses match both ways” (article 3).

”Can it take the load?” Performance under pressure

The load-test tool wraps k6, so anyone can launch scenarios, watch runs and read the report without a terminal. It has been used for real: one event pushed 1,600+ orders through in five minutes in production, and an earlier test showed that 80-120 concurrent account-creation calls failed. That is the kind of finding you want from a test, not from a customer.

”Who needs to know what?” People and process

The dependency alerts tracker lists open security alerts and automated update PRs, politely: it refreshes on a schedule and never hammers the source host. The ticket and PR reader turns a ticket and its pull requests into clean markdown, and feeds a QA-readiness queue that checks each sprint ticket against a checklist using a small, cheap model. Meeting notes turns a pasted transcript into per-ticket notes. All of these are the subject of article 5.

”How do I run this thing?” The workbench

The Playwright runner runs tests from a UI: workers, a tag or grep filter, or a pasted raw command that a hand-written parser cleans up. It streams output, cancels, archives and restores reports, and browses traces.

Plus context and roadmap pages: markdown “context documents” that are literally the system prompts for the AI features, editable in the UI. Your prompts are documentation, and your documentation is prompts. I rather like that.

One principle that runs through every tool

A green result means nothing until you know how many things were examined. Early on, a type-check on the server printed “clean” while looking at zero files. That is why every tool in the test framework states its denominator: how many tests, how many days of history, how many failures were categorized, how many were left over. The idea comes back in almost every article.

Where to next

Test cases are now cheap. Next: Rules first, LLM second, where the expensive part gets cheaper too. The rest of the series is in the line at the top: comparison, the agent layer, dashboards for the whole team, and finally how a laptop tool became a team service.