Skip to main content

AI agents

Multi-Agent Research System with a Human Checkpoint: Cadenza

How we built Cadenza, our multi-agent AI research system: agents plan, search and write, a person approves, and each claim is checked against its source.

Mohammad Ashraful Islam8 min read
question to cited brief in our test run (machine time)
39 s
model cost of that run, on our own API key
$0.03
claims matched word for word to their cited source
5/5
research runs for our own business so far
50+

Market research is the slowest part of our week. Someone has to search, open twenty tabs, read, compare and write it up, and a single AI chatbot does it faster but invents numbers that look exactly like real ones. So we built Cadenza, a multi-agent research system: a planner splits the question, three researchers read the web in parallel, an analyst sums up, a person approves the direction, then a writer drafts the brief and a critic checks every claim against the page it cites. In our measured test, a question went to a cited brief in 39 seconds for about 3 cents of model usage.

Cadenza is our own system, not client work. One of our AI engineers built it for our internal market research, and we've run it more than 50 times for our own research so far. It is not perfect yet, and this case study says where.

At a glance

  • What it is: a multi-agent AI research system that turns one question into a short, cited market brief
  • Built by: Devs Core, for our own research; one AI engineer, with the first full pipeline working in nine days (23 June to 1 July 2026)
  • Who uses it: our team; 50+ research runs for our own business development so far
  • The agents: a planner, three parallel researchers, an analyst, a writer and a critic, plus one human approval step
  • Try it: a recorded demo run is live at cadenza.devs-core.com; live runs on your own API key are coming next
  • Models: Claude or GPT with your own key, with a cheaper model routed to the mechanical steps
  • Stack: LangGraph, FastAPI with server-sent events, Next.js with React Flow, Redis, PostgreSQL, Tavily search and Firecrawl

The problem: AI research that sounds right and isn't

We research markets, competitors and prices for our own business development every day. Doing it by hand takes hours. Doing it with one AI prompt is fast, but it fails in ways that are expensive:

  • Confident, invented numbers. A single model asked for a market size will give one, cited or not, and the made-up figure reads exactly like the real one.
  • No way to see how it got there. You get an answer and a list of links, and no way to tell whether the answer actually came from those links.
  • The open web fights back. A research agent reads pages anyone can write, including pages with hidden instructions aimed at the AI. That's called prompt injection, and it's the top risk on OWASP's list for AI applications.
  • Wrong direction, found too late. If nobody checks the direction until the brief is finished, you're proofreading a fluent document, which people are bad at.

A research agent is only useful if you can check it. Speed without a source is just a faster way to be wrong.

What we built: a research team, not a prompt

  • Planner: breaks the question into focused sub-questions and shows why it chose them
  • Three researchers, in parallel: each searches the web for one sub-question and reads the best page it can, with every page screened for injection before any model sees it
  • Analyst: sums up what the pages actually say and proposes a direction for the brief
  • Human checkpoint: the run stops and waits. A person approves the direction or adjusts it in plain words
  • Writer: drafts the brief from the approved sources only, with inline citations, and follows the person's note
  • Critic: checks each key claim against the source it cites, sends unsupported claims back to the writer, and flags anything that still doesn't hold
  • Live console: a graph of the agents updates as they work, beside a plain-English explainer and a log of every decision

Cadenza: AI agents that research, debate and fact-check, with the agent pipeline from planner to critic

Everything runs through one orchestration graph, and the graph on screen is the same shape as the one in the code, so what you watch is what's actually running.

For the person asking: one question in, a brief you can check

The Cadenza console: research question, model and key settings, the live agent graph and the event log

  • Ask in plain words. Type the question and press Run. The planner turns it into three sub-questions, so the researchers don't all chase the same thing.
  • Watch it work. The graph lights up agent by agent, the log shows each search and each page read, and a running meter shows steps, tokens, estimated cost and time.
  • Steer before anything is written. At the checkpoint you see the analyst's proposed direction. In our test it drifted towards banking and healthcare, so we typed one line asking it to stay on Malaysian distributors, and the writer followed it.
  • Get a brief with its receipts. Every section cites its sources, the byline says how many claims were verified, and each run gets a shareable link.

For trust: four things that separate a real agent system from a toy

Why it's hard: prompt-injection defence, decision transparency, claim verification and human-in-the-loop control

  • Prompt-injection defence. Every fetched page is screened by a rule-based guard before a model reads it. A page is passed, cleaned of suspicious lines, or quarantined, and the screen shows which.
  • Decision transparency. The planner's reasons for its sub-questions and the critic's verdict on each claim are shown, not hidden in a log file.
  • Claim verification. The critic looks for the claim's key figure or phrase word for word in the cited page. In one test run it caught 4 of 10 claims that weren't in their sources and sent the draft back.
  • Human control. The run can't write anything until a person approves, and the run's state is saved while it waits.

For cost: cheap by design, and you see the number

Multi-agent runs use more tokens than a single prompt, so cost is built in from the start:

  • Model routing: a stronger model plans, analyses and checks; a cheaper one handles the mechanical steps
  • Hard limits: every run stops at 250,000 tokens or 24 steps, and the critic can retry at most twice
  • Bring your own key: visitors run on their own provider key, so a public demo costs us nothing; without a key they get the recorded run
  • The real number: our measured run used 14,734 tokens, about $0.03 by Cadenza's own cost meter

Under the hood: LangGraph, bring-your-own model, model routing, FastAPI and server-sent events, React Flow, the injection guard, Redis and Postgres

How we built it

The checkpoint goes in the middle

We put the human step after the analysis and before the writing. Too early, and you approve a plan with no evidence. Too late, and you're proofreading a fluent draft. In the middle, the evidence is in and nothing is written, so changing direction costs one sentence. Our test showed why: the analyst's first direction was generic, and one adjustment fixed it.

Testing it for this case study found real bugs

Before writing this, we ran Cadenza end to end on a real question and measured it. The first runs failed, and we fixed 12 bugs in one day. Two of them matter to anyone buying an agent:

  • A silent fallback. When the researchers couldn't read their pages, the writer quietly used built-in example content instead of stopping. Now a run with nothing readable stops and says so.
  • An over-generous label. When claims failed every retry, the brief still said "claim-verified". Now it shows the real count, for example "4/5 claims verified", and flags the rest.

Both looked fine in a demo, and both would have mattered in production. That's why we test agents on real questions with real keys, not only on mocked runs.

Tests that don't burn money

The orchestration runs against a mocked model in the test suite, so the graph, the checkpoint, the retry loop and the injection guard are all tested on every change without paying for a model call. The backend has 84 automated tests, plus front-end tests and an end-to-end browser test.

The results

Before Cadenza With Cadenza
Hours of searching, reading and writing up A cited brief in under a minute
One AI answer you can't check Every claim checked against the page it cites
Web pages trusted as they are Each page screened for injection before a model reads it
Direction checked only at the end A person steers before anything is written

Our measured test run, on 4 October 2026, with Claude Sonnet and cost routing on:

  • 39 seconds from question to finished brief, 30 of them before the human checkpoint
  • 3 pages read and screened, with 0 injection flags
  • 5 of 5 claims grounded word for word in their cited source
  • 14,734 tokens, about $0.03, by Cadenza's own meter

What it doesn't do well yet

  • Global sources for local questions. Asked about Malaysia, the researchers found strong global market reports rather than Malaysian ones. Better search queries are next.
  • Strict checking. The critic matches text exactly, so a true claim written in different words can be flagged. That's safe, but strict.
  • Two providers today. Live runs work with Claude and GPT. Gemini, Llama and Mistral are next.

What we'd do again

  • Put a person in the middle, not at the end. It's the cheapest place to catch a wrong direction.
  • Screen everything the web hands you before a model sees it, and show the screening to the user.
  • Make the agent prove its claims against its own sources, and say honestly when it can't.
  • Test with real keys on real questions. Mocked tests keep the logic right; real runs find the bugs that matter.

Frequently asked questions

What is a multi-agent research system?

It's a set of AI agents that each do one job, such as planning, searching, analysing, writing and checking, and pass work to each other. In Cadenza, five kinds of agent and one human checkpoint turn a single question into a cited brief.

How much does a multi-agent AI research run cost?

Our measured Cadenza run used 14,734 tokens, about $0.03 of model usage by its own meter. Routing the mechanical steps to a cheaper model and capping every run at 250,000 tokens keep the cost predictable.

How do you stop a research agent from making things up?

Make it cite, then check. Cadenza's critic looks for each claim's key figure or phrase word for word in the page it cites, sends unsupported claims back for revision, and flags any that still fail rather than hiding them.

How do you protect an AI agent from prompt injection?

Treat every fetched page as data, never as instructions. Cadenza screens each page before a model reads it, removes or quarantines anything that looks like instructions to the AI, and shows the result in the run log.

Why does Cadenza stop for human approval?

Because the direction is the cheapest thing to fix and the hardest to spot in a finished draft. Cadenza pauses after the analysis, so a person can approve or redirect the brief before a word of it is written.


Want an agent like this working inside your own business? Book a call, or see how we approach AI agent development.

One email when we publish something worth your time

No cadence, no drip sequence. We write when we have learned something running agents in production, which is not weekly.