A research agent takes a question, searches for sources, and writes a report that cites them. Asked about something that changed recently, it can do what a model answering from memory cannot: find the current answer and show where it came from.
Here is why that matters. On 27 September 2026 I asked Claude Sonnet 5, with no tools, when the EU AI Act's obligations for the high-risk AI systems listed in Annex III start to apply. It answered 2 August 2026, the date in the Act as adopted in 2024, and flagged that it could not confirm whether the date had changed. It had: an amendment moved it to 2 December 2027. The same model, inside the research agent in this post, found the new date and cited the European Commission's announcement.
An agent also brings new ways to go wrong. It can cite a source for something the source does not say, or trust a vendor's blog as much as the law itself. This post walks through how I built one and how I tested it, in five steps:
- Pick the simplest level that does the job.
- Give it the right sources.
- Compute confidence in code, not in the model.
- Add a reviewer that can say no.
- Test the reviewer with planted errors.
Step 1 draws on Dave Ebbelaar's five levels of AI agents and Shaw Talebi's guide to multi-agent systems. The code, the benchmark results and the planted-error test are on GitHub.
Step 1: Pick the simplest level that does the job
AI systems come in levels of complexity. At the bottom is a single call to a model. Next is a fixed workflow, where code decides the order of the steps and the model does the work inside each one. At the top are agents that decide their own next step, and teams of agents coordinated by a supervisor. Each level up costs more, takes longer and has more ways to fail. So start at the bottom, and move up only when the level below cannot do the job.
I built the same research task at three levels, all running on Claude Sonnet 5:
| Version | Level | Time and cost per report |
|---|---|---|
| V1 | One model call, no tools | 24 to 42 seconds, 2 to 4 cents |
| V2 | A fixed workflow of four steps: researcher, analyst, writer, reviewer | 85 to 182 seconds, 13 to 30 cents |
| V3 | A supervisor with researchers working in parallel, and no reviewer | 87 to 115 seconds, 14 to 19 cents |
V2 and V3 split the work between specialised roles, each a model with its own instructions. They differ in who decides the order:
V1 cannot answer questions whose answers changed after the model was trained, and it has no sources to show. V3 splits the question and searches in parallel. Several agents pay off when a task splits into independent parts, and searching many sources at once is one of those, so V3 found more sources than V2 (59 to 87 per report, against 23 to 28) in less time. But it has no reviewer, and step 4 is why that matters.
V2 is the one I would hand to a client. Code fixes the order (research, then analysis, then writing, then review), and the model does the work inside each step but never decides what comes next. Research continues until there are at least 5 sources or 3 rounds have run, and the review loop ends when the report is approved or after the second round.
Step 2: Give it the right sources
An agent is only as current as what it reads. V2's researcher writes search queries and runs each one twice through a web search service: once restricted to official EU websites (europa.eu), and once on the open web. Official sites are searched on purpose, because open web search on EU rules mostly returned vendor pages and blogs, which is also where the out-of-date deadlines were.
For the Annex III question, 10 of the 23 sources were official EU pages, including the Commission's announcement that the amendment entered into force on 27 July 2026, the AI Act Service Desk, and the updated text of the Act on EUR-Lex:
The prompts carry today's date instead of facts, because facts in a prompt go stale without anyone noticing. One of my prompts stated that the amendment was "proposed, not adopted" and that the obligations applied from 2 August 2026. Both became false on 27 July 2026, and from then on, an instruction written to keep reports current made them out of date.
Step 3: Compute confidence in code, not in the model
Each claim in a report gets a confidence score. The model's only job is to say which sources support a claim. Code computes the score from those sources, from three parts:
- Authority: how much the strongest source counts. The law itself on EUR-Lex counts 1.0, the Commission and other EU bodies 0.85, national regulators 0.7, law firms 0.5, vendors 0.3, and news and blogs 0.15. The type comes from a fixed list of websites, not from the model.
- Recency: full weight for a source up to 90 days old, then 0.8 up to six months, 0.5 up to a year and 0.2 beyond. A source without a date gets 0.8: it may be current, but nothing shows it.
- Corroboration: 1.2 when two independent websites agree and 1.4 for three or more. Three pages from the same site count as one.
The three are multiplied, capped at 1, and the result is labelled high confidence from 0.7 and medium from 0.4.
Why code? When I put the formula in the prompt and let the model write the number, none of the sources had a date, yet recency had been "applied", and 12 claims scored above the highest value the formula allows. Asked to apply a formula, the model produced numbers that looked like the formula without computing it. In code, the same sources always give the same score, and every score can be traced back to them.
A score measures the evidence, not the truth. The Annex III report named the amendment as Regulation (EU) 2026/1744 with a confidence of 0.14, because the only sources that gave the number were two compliance websites: 0.15 for authority, times 0.8 for having no date, times 1.2 for two sites. The number is correct, and I checked it on EUR-Lex. The report said exactly as much as its sources could support.
Step 4: Add a reviewer that can say no
The reviewer is the last step before a report is approved. It first checks, in code, that every link in the report points to a source the researcher actually found. Then a model checks the report against the extracted claims on six points: every statement cited, no facts beyond the claims (including changed numbers and dates), contradictions disclosed, confidence labels that match the scores, a summary that matches the body, and gaps in the evidence stated.
A report is approved only when every check passes. If one fails, the report goes back to the writer with the reviewer's instructions, and the reviewer looks again. If it still fails, the report ends as "not approved". This is called failing closed. The opposite, failing open, approves a report whenever the reviewer's own output cannot be read, and it is an easy bug to write.
In the benchmark run, the reviewer approved three of the four reports. The fourth ended not approved, because it stated penalty figures more strongly than the claims supported and left an open question out of its knowledge gaps.
Step 5: Test the reviewer with planted errors
A reviewer that approves everything looks exactly like one that works, until something goes wrong. So I tested it the way I would test any model: with errors of known types, planted where I knew they were.
I took four finished reports and made copies with one error each:
- a sentence citing a Commission press release that the researcher never found;
- a plausible obligation, cited to a real source, that no claim supports: "Providers must also notify their national data protection authority at least 30 days before placing a high-risk AI system on the market";
- a claim restated without its citation;
- a date moved by one year, from 2 August 2026 to 2 August 2025.
The reviewer ran on every copy, and on each unmodified report to count false alarms.
| Planted error | Caught | Checks that flagged it |
|---|---|---|
| Link to a source that was never found | 4 of 4 | link check 4, hallucination 4, confidence 4 |
| Obligation that no claim supports | 4 of 4 | hallucination 4 |
| Claim without a citation | 4 of 4 | citation coverage 3, confidence 1 |
| Date moved by one year | 1 of 1 | hallucination 1 |
| None (false alarm) | 1 of 4 rejected | confidence labels |
It caught all 13 planted errors, for 72 cents across 17 reviewer calls. The date error ran only once, because only one report had a date in the format the test changes. The one false alarm is a fair objection: the reviewer rejected a clean report because one sentence merged a medium-confidence claim and a high-confidence claim under the medium label.
The test also shows what the reviewer cannot catch. It checks the report against the extracted claims, so an error that is already inside a claim passes through. The Annex III report has one. The analyst took the European Parliament's "36 months after the entry into force", which is the date for AI in products covered by other EU safety law, and attached it to the Annex III date, which is 24 months after. The claim is wrong, the report repeats it faithfully, and the reviewer approves. The fix is to check each claim against the text of its sources, the way the question-answering system I built on the same regulation checks every citation against the passages it retrieved.
The result
V2 against V3, over the four benchmark questions:
| V2 | V3 | |
|---|---|---|
| Time for the four questions | 558 seconds | 381 seconds |
| Cost for the four questions | 92 cents | 64 cents |
| Sources per report | 23 to 28 (10 to 14 official EU) | 59 to 87 (30 to 44 official EU) |
| Quality check | Reviewer: 13 of 13 planted errors caught | None |
V3 is faster, cheaper and better sourced. What it gives up is the check: nothing stops a report with an invented link or an unsupported claim from going out. For broad questions where coverage matters more than a check, V3 is the better tool. For anything a person will rely on without re-checking it, use V2, or V3's parallel search with V2's reviewer added, which is the obvious next version.
Limits
- Few sources have dates. Search results rarely include one, so most sources count as undated.
- The website list is hand-made. Reputable sites that are not on it count as blogs, which undervalues them.
- The reviewer is a model, and it checks the report against the claims, not against the sources.
- The benchmark is four questions. The results show behaviour, not statistics.
The five steps, in short
- Pick the simplest level that does the job, and move up only when the level below cannot.
- Give it the right sources, and put today's date in the prompt, not facts.
- Let the model choose the sources, and compute the confidence in code.
- Add a reviewer that can say no, and make it fail closed.
- Test the reviewer with planted errors, and write down what it cannot catch.
The code, the benchmark results and the planted-error test are on GitHub: multi-agent-research-system. The companion post walks through the question-answering system over the same regulation, in five steps.