Building Fiuto · Part 12 of 12 · 2026
LinkedIn EmailFiuto: visual QA with a swarm of agents
Reviewing a delivered feature that crosses several repositories can take a team days: someone walks the journey, records what breaks and where, and finally writes tickets developers can pick up. With agentic workflows, the user journey can be reviewed by swarms of agents working in real browsers, who can then write a report just like a human QA team would. The report they return still has to be read and decided on before any of it becomes a ticket.
Context
When a large feature ships there will be more issues than I can review as a solo builder. In this case, the delivered “AI generated decks from insights” feature contained too many issues for me to list manually. To map and fix all issues as fast as possible, the build orchestrator fanned out a review of the feature to ten agents working in a Playwright browser. Their final report returns after minutes, it is reviewed and converted into a Linear project to be actioned later.
Reviewing a delivered feature
The review of the analysis section had to cover four things before any fix could start.
- The full journey, from generating an analysis to opening and editing the deck, since a defect like the stretched deck only shows at the end of it and a review that starts from the deck misses what the analysis handed it.
- Each aspect of the implementation in the rendered product, because the defects here are in what renders and in what fails when you interact with it, and a read of the source does not show either.
- One report: what passed, what failed, what the run cost and every bug found, so the result can be read once and compared with the next run after the fixes.
- The source location of each bug, because with a separate generator and package the same visible fault can belong to any of three codebases, and a ticket without that location goes to the wrong one.
The first two are a tester’s days, the last two a developer’s triage. In the recording they are one agent’s mapping run, ten agents’ browser sessions, a report and a project.
01The pipeline
The build orchestrator fanning QA out
The build orchestrator is the written procedure that moves feature work through delivery, the one I showed in the planning and implementation part of this series. Here I give it the issues from the day before and have it fan the review out, so that the same procedure that implemented the feature is the one that reviews it.
MapOne agent walks the full journey from generating an analysis to opening and editing the deck, and turns it into the aspects to review.
Fan outTen agents each take one aspect and review it visually in a Playwright browser.
ReportThe findings come back as one report: a summary, the validation failures and successes, the cost of the run and every bug found.
Map to sourceAn agent locates each bug across the app, the generator and its package, and composes the tickets into a project.
ImplementThe build implement agent picks the project up and fixes it.
The agents return a report with a list of issues ranked by severity. After human review these issues are added as tickets to a new build project in Linear, so that a fresh build session can action them sequentially.
02The session
What is in this video
Eleven agents did the work in this recording, one mapping the journey and ten each reviewing one aspect of it in a browser, and what they reviewed is the analysis section. The overview is a cross-study view across three studies, all of them on synthetic data that exists for testing, and inside one analysis sit a summary under the title, the main findings, a full analysis page and the deck. At 2:20 I open the deck I had worked on the day before: it renders stretched vertically, editing it fails, and there are more bugs than I go through on camera. The deck’s origin comes up at 3:07, a separate report that generates the slides through its own package, and I say what that means for the review: with features this large, checking the implementation on my own would take days.
At 3:56 the report is on screen. It summarises the run, and the failures are not only in the generation: three validation failures against 15 successes, the calculated cost of the run, and the list of every bug the agents found. I point at the cost because the next piece of work I had named was modelling the cost of each AI feature, and a report that already prices its own run gives that work a first figure. The eleven agents produced this in a session, and the result is more specific than what I would have written after days on it.
The session. A stretched deck, ten browser reviews, one report, five minutes, with the tickets still being composed into a project as the video ends.
As the recording ends, an agent is mapping the source for each bug and composing the tickets into a project, and the next step is the build implement agent. When the recording ends the project is still being composed, so there is no fix outcome to report, and the ten aspects are not named on camera, the cost figure is not printed, and the three failed validations are not identified. Fiuto soft launched after this run, so staging and production no longer carry the same weight, and two narrower QA procedures now sit beside the swarm as alternatives I choose between. /test signs synthetic users into the product as personas and measures what each journey costs. build-qa executes checks written during planning once ordinary verification is green. The swarm is the one I reach for when the defect set is open and crosses repositories.
Conclusion
Twelve parts have now followed features through the whole loop inside the same coding agent: discovery, research included, then design, then delivery, and with this one the review that comes back after delivery, the agent checking its own implementation visually and returning a list of work to the orchestrator it started from. The reusable point of this last part is the two columns the report carries beyond the bug list. One is a source location for every bug, without which a fault in a deck rendered by a separate package gets ticketed against the wrong codebase. The other is the cost of the run, which gave the next piece of work its first figure before that work had started.
Eleven agents, one to map the journey and ten in a browser, returned three validation failures, 15 successes, a cost and a list of bugs in a session, with the project to fix them still being composed when the recording ended.
Contact
Get in touch
Product design lead across research, design systems and handover, with coding agents in the loop. Contract, outside IR35, or a permanent lead role.
- Available
- Now
- Location
- London, remote or hybrid, UK and EU
- Right to work
- UK settled status, EU citizen
- RolesLead Product Designer, Principal Product Designer, Design Engineer
- ScopeResearch, design systems, prototyping, handover, QA with agents
- Email[email protected]
- LinkedInlinkedin.com/in/dariocodipietro