Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service
At some point I noticed I was typing the same sentence to an AI assistant every morning: “Here’s last night’s failing run. Group the failures, tell me which ones are data problems, and count them.”
It does a decent job. It also does a slightly different job each day, burns tokens re-reading the same shape of data, and can miscount a bucket. Counting is not what language models are for.
That sentence is the whole argument of this article. When you catch yourself asking an AI for the same report a third time, stop prompting and build a tool. Then let the AI call the tool.
The front end is optional
The test framework started as a React app. Every capability, such as analyzing a run, re-running failures or comparing captures, lives behind an HTTP route on a Node server. The UI is one client of those routes.
An MCP server is just a second client. MCP (Model Context Protocol) is the open standard that lets an AI agent discover and call tools. I put one in front of the same routes, and suddenly the front end became optional. I can sit in a conversation with an agent and say “pull last night’s run and tell me what broke”, and it drives the same backend the UI does. I stopped thinking of the UI as the product and started thinking of the backend as the product, with two ways to talk to it: buttons for humans, tool calls for agents.
The thin-wrapper principle
The MCP server exposes 21 tools. Every one is a thin HTTP wrapper over a test framework route. That is deliberate, and it is the rule I would defend hardest:
flowchart TD A["Agent"] --> B["MCP tool"] B --> C["HTTP route"] C --> D["Test framework logic"] D --> E["Tests, storage, LLM"]
- No QA logic is duplicated. The rules, the LLM defences, the remediation mapping: all of it lives in the test framework. If I fix a bug in triage, the UI and the agent both get the fix at once.
- No secrets in the MCP server. Keys live only on the test framework server. The MCP process holds no keys, so adding an agent did not add a new place to guard.
- Tools are boring on purpose. A wrapper that is just “call route, return JSON” is almost impossible to get wrong, and when it does break, you know the bug is in the route.
There is a corollary I like. Where the UI does heavy parsing, the agent often does not need the parsing at all. For response comparison, the tools stay thin and the agent reads raw trace files itself instead of me porting the UI’s parser into a second place. Agents are good at reading messy data. They are bad at being the only thing standing between you and a wrong number.
Why a recurring analysis becomes a tool
Back to the morning sentence. Here is what changes when it becomes a tool:
| Prompt, every day | Tool | |
|---|---|---|
| Tokens | Re-spends context on the raw data each time | The data never enters the model; only the result does |
| Consistency | Varies run to run | Same input, same output |
| Counting | ”Approximately”, sometimes wrong | Exact, because code counted |
| Speed | Seconds to minutes of model time | Cached, mostly instant |
| Testability | You can’t unit test a mood | You can test it like any function |
The model still does what it is good at: reading the result, noticing that two categories are really one story, and deciding what to do next. The tool does the part that has a right answer.
This is the same instinct as the triage pipeline in article 2: deterministic first, judgement second. The tool is the deterministic layer. The agent is the judgement.
Read-only versus approval-gated
Not all 21 tools are equal. Some just look, and some change the world. Every tool’s description says which, because the description is what the agent reads when deciding what to call.
The read-only side is the large majority: list CI runs, pull a run, load a run, categorize failures, persistent-failure stats, draft a remediation, check whether a test account’s credential still works (a login attempt, nothing more), and verify a product is in stock (a quote probe only). Plus the dependency-alert tools and the archive and response-comparison tools.
The mutating side is small and loud:
- Apply fix changes test data, so it requires explicit approval from the user.
- Re-run failures executes real tests. It takes a bounded subset and has a five-minute cap, so a curious agent cannot launch the whole suite by accident.
- Create ticket makes a real ticket in the ticketing system on every call. The description says so in plain words.
The loop that ties these together is the one I would put on a poster:
flowchart LR A["Propose"] --> B["Human approves"] B --> C["Apply"] C --> D["Re-run to verify"]
The draft-remediation tool only proposes. Each proposal carries a “certain / not certain” flag, a reason and a requires-human-review marker. Nothing changes until a person says yes. Then apply runs, and the re-run gives a per-test verdict, so “I fixed it” becomes “I fixed it and here is the test passing”.
One detail worth stealing: for some fix kinds, apply never edits the data file itself. It returns a snippet and the calling agent pastes it. The server stays out of your source tree, and the change shows up in your normal diff and review flow.
Context documents are editable system prompts
The test framework’s AI features need domain knowledge: the vocabulary of the backend, the conventions of the test suite, the format a ticket title must follow, and the gotchas that only people who have been bitten know about.
I keep this in markdown “context documents”: an overview, conventions, backend domain knowledge, and ticket-writing format. They are not documentation that happens to sit near the prompts. They are the system prompts. The server loads them when it builds a prompt, and they are viewable and editable in the UI.
That has a nice consequence. When the AI gets a domain detail wrong, the fix is not a code change and a deploy. It is a sentence in a document. A QA engineer who has never touched the server can improve the AI’s behaviour by editing prose, and the same documents ground the agent in conversation.
Teaching the system a new failure pattern
This is the recipe I wrote down, because the first time I did it from memory I forgot a step.
Say a new, repeatable failure shows up: one specific business error that currently gets lumped in with a generic bucket.
- Add a rule above any broader overlapping rule. First match wins, so order is the whole game. A specific rule must sit above the generic one it would otherwise be swallowed by.
- Add the matching remediation branch. Decide what a fix looks like for this pattern, and whether it is fixable at all. “Not fixable, needs a human” is a legitimate answer.
- Update the tool descriptions. The agent only knows what the descriptions tell it. A capability nobody described may as well not exist.
- Verify end to end on a real run. Not a fixture. Pull a real run, confirm the new rule claims the right tests, confirm the old buckets shrank by the same amount, and confirm nothing vanished.
Step 4 is the denominator check from article 2 in practice: a green result means nothing until you know how many things were examined, so count what the new rule claimed.
What it looks like in conversation
Here is an invented, generic session. The names and numbers are made up, but the shape is real.
Me: Pull the latest nightly run and tell me what’s broken.
Agent: (calls list CI runs, then pull a CI run) Pulled the most recent run. 41 failures. Categorizing now.
(calls categorize failures)
Agent: Three groups. 22 look like one business error returned by the same operation, 12 are failed logins for what looks like a single test account, 7 are an out-of-stock product in cart scenarios. The rules claimed 34; the other 7 went through the model and were sanity-checked. All 41 are accounted for.
Me: Draft fixes. Don’t change anything yet.
(calls draft remediation)
Agent: Proposals:
- Out-of-stock (7): swap the product for an in-stock candidate from the resolver pool. Certain, I probed the candidate.
- Failed logins (12): the account’s credential check fails. Not certain whether to reset or reuse; needs your call.
- Business error (22): not a data problem. A recurring operations task appears not to have been run. Not fixable by editing tests, so I’ve left it alone.
Me: Apply the stock swap. Skip the rest.
Agent: Applying the swap for the 7 stock failures. (calls apply fix) Re-running just those tests, capped at five minutes. (calls re-run failures) 7 of 7 pass. Here’s a command you can paste to verify independently. Want me to open a ticket for the business error?
Me: Yes.
Agent: Created. It’s one real ticket; I won’t create another unless you ask.
Notice what the agent never did: it never edited anything without my yes, and it never counted anything itself. The numbers came from the tools, and the judgement came from the agent.
One more guardrail
In hosted mode there is a per-user daily ceiling on AI requests. The remaining budget is reported in the health endpoint and in response headers, so an agent can see its own allowance rather than discovering it by hitting a wall. A friendly agent with no budget awareness is a very efficient way to spend money.
I also keep the repo’s skill and process documents pointing at the test framework tools, and the recipe above lives next to them. The agent reads those documents, so the written recipe is both for me and for the next conversation.
The takeaway
Build the tool the third time you ask the AI for the same report. Keep it thin, share it between the UI and the agent, make read-only and mutating tools obvious, and make humans approve every mutation. The model stays good at judgement because you took the counting away from it.
Next: 5. Dashboards for the whole team: the load-testing reports, trend views and trackers that make the test framework useful to people who never open a test file.