Skip to main content

Package

Fixed scope4–6 weeks

Agent Pilot — one process, proven on your live data in four to six weeks

Not a demo on sample data. One real process, running against your real systems, with scoped permissions and a review queue on every write — then run beside the people who do that job today so the agreement rate is measured rather than claimed. At the end you have a working system, a number you can take to a board, and the source code.

$6,500 from · 4–6 weeks · scope agreed in writing

Scope agreed before we start
From $6,500Scope agreed before we start
To a measured result
4–6 weeksTo a measured result
Your systems, not a sandbox demo
Live dataYour systems, not a sandbox demo
Source code, in your repository
100%Source code, in your repository

What you receive

6 named artefacts, not a summary call

A pilot that ends in a slide deck is a failed pilot. These are the things that exist on the last day, and all of them are yours.

  1. 01

    A working agent on one real process

    Running against your systems with production data, handling the routine path end to end and routing everything else to a person. Deployed in your environment, not ours.

    You getRunning system in your infrastructure
  2. 02

    Connectors for the systems that process touches

    Typically two to four: the system of record, the channel the work arrives on, and wherever the result has to land. Written as reusable adapters rather than one-off glue, so the second agent inherits them.

    You getDocumented adapters with their tests
  3. 03

    A permission model and a review queue

    Per-system, per-operation scopes — the credential can do the one thing the agent needs and nothing else — plus a queue where a named reviewer sees every proposed write with its evidence and a one-click accept or change.

    You getPermission matrix and the reviewer interface
  4. 04

    The agreement-rate report

    The parallel run, measured: how often the agent and your team reached the same answer, broken down by case category, with every disagreement examined rather than averaged away. This is the number that decides what goes unattended.

    You getPer-category agreement figures and the disagreement log
  5. 05

    Cost and failure instrumentation

    Per-run token cost, latency, retry and failure rates on a dashboard you can read without us. Agent costs are lumpy and nobody should discover that from an invoice.

    You getDashboards and alerting, in your accounts
  6. 06

    The repository and the runbook

    Full source, commented, with deployment instructions, the rules as versioned code, and a written runbook for the failure modes we actually hit during the pilot.

    You getRepository and operations runbook

How it is done

Why the parallel run is the whole design

Every agent demo works. The question a pilot has to answer is whether it works on the cases your team finds hard, often enough to be trusted — and the only honest way to answer that is to run both and compare, on live data, for long enough to see the tail.

  • Writes are reviewed from day one, not retrofitted

    The review queue exists before the agent does anything useful. Adding approval to a system that was built to act freely is a rewrite; building the queue first costs nothing and is the reason the pilot can touch live systems safely.

  • Unattended is earned per category, not switched on

    Routine orders might clear 99% agreement in week three while exceptions sit at 70%. Those are different decisions. You release the first and keep reviewing the second, and you can move that boundary yourself afterwards.

  • Rules live outside the model

    Thresholds, limits, policy and tolerances are code with tests and a version number. The model reads, drafts and proposes; the rules decide what may pass. That is what makes a decision explainable nine months later.

  • The team doing the job is in the room

    They define the exceptions, they review the queue, and they see the agreement numbers. A pilot that lands on a team as a surprise gets worked around within a fortnight, however good the technology is.

Week by week

What happens, and what it costs you in hours

The last column is the one nobody publishes: how much of your own team’s time each stage takes. It is the question every buyer has and almost nobody asks out loud.

  1. Week 1

    Lock the scope and the rules

    Exactly which process, which cases count as routine, which must always stop for a person, and what agreement rate would justify going unattended. Written down and signed off before anything is built.

    Your time2–3 hours, process owner plus one engineer
  2. Week 1–2

    Build the connectors

    Authentication, scoped credentials, the reads and writes the process needs, idempotency and retry behaviour. Against a sandbox where one exists, with the production path configured but not yet enabled.

    Your timeIT access, then on call
  3. Week 2–4

    Build and shadow

    The agent runs on live data in shadow mode: it produces an answer for every real case and writes nothing. We compare daily against what your team actually did and fix what differs.

    Your time2 hours a week reviewing disagreements
  4. Week 4–5

    Parallel with the queue open

    Writes are enabled through the review queue. Your reviewer approves or corrects each one, which both protects the systems and produces the corrections the agent learns from.

    Your time20–40 minutes a day, one reviewer
  5. Week 5–6

    Measure, release and hand over

    The agreement report per category, a decision on what goes unattended, repository handover, runbook walkthrough and the monitoring in your accounts.

    Your time2 hours total

Not included

What this package does not cover

Published rather than discovered. A fixed price only means anything if the edge of it is written down.

  • A second process — the pilot is deliberately one, to keep the result readable
  • Rebuilding a system of record, or fixing data quality upstream of the process
  • Unattended payments, refunds or money movement of any kind
  • Model or infrastructure costs, which are billed to your own accounts at cost
  • A support retainer, unless you separately want one
  • Change management across a wider team than the one running this process

Is this the right package

Who this is for, and who it is not for

We would rather lose the sale here than three weeks in. If the right-hand column describes you, say so on the call and we will point you somewhere better.

A good fit when

  • You can name one process and say roughly how many cases a week it handles
  • It has a right answer that two people would agree on most of the time
  • You can give a reviewer twenty to forty minutes a day for two weeks
  • Your systems have an API, a database or a scheduled export
  • You need evidence for a wider decision, not a proof of concept for its own sake

Not this package when

  • The process is pure judgement with no checkable answer
  • Volume is a handful of cases a month — the measurement will not mean anything
  • Nobody internally can review the queue during the parallel run
  • You want five processes at once; that is the Access Layer, not a pilot
  • You want it live and unattended in week two, before anything has been measured

The money

$6,500 from — what that covers

From $6,500, quoted as a number before we start

The figure depends on how many systems the process touches and how awkward they are. You get a fixed quote from the audit findings or a scoping call — not an hourly rate with an estimate attached.

What moves the price

Each additional system beyond three, a system with no API that needs a file-based path, and any requirement to run models on your own infrastructure rather than a hosted API. All of it is priced before the work starts.

Model and hosting costs are yours, at cost

Model keys and cloud accounts are in your name, so you see the real bill with nothing marked up through us. We size the expected monthly run cost during week one and put it in writing.

Questions

Agent Pilot — before you commit

Ask us something else

What if the pilot fails?

Then you have found out in six weeks for a known price, with a written account of why — usually data quality, an integration limit, or a process that turns out to need redesigning rather than automating. You keep the connectors, the permission model and the report. That is a considerably better outcome than a year-long programme discovering the same thing.

What agreement rate should we expect?

It varies by process and we will not quote a number before seeing yours. What matters is that the number is measured per category rather than averaged, because a headline 95% can hide an exception class at 60%. The threshold for going unattended is yours to set, not ours.

Can the agent act without review during the pilot?

Only in categories you have explicitly released after seeing their agreement rate, and usually only in the last week or two. Everything starts in shadow mode, writing nothing at all.

Who owns the code and what happens if we stop working with you?

You own it. Source in your repository, infrastructure in your accounts, model keys in your name, no runtime of ours in the path. If you stop working with us the system keeps running and another team can pick it up — the runbook is written for that case.

Do we need the audit first?

No, if you already know which process and your systems are straightforward. Yes, in practice, if you are choosing between several candidates or you have not tested whether your systems can be integrated — the audit fee comes off the build, so it is usually the cheaper order.

What if we want to change the process mid-pilot?

A small clarification is normal and absorbed. Changing to a different process is a new pilot, and we will say so rather than quietly delivering less. Scope drift is the most common way a fixed-price pilot turns into an unhappy one.

Which models do you use?

Whichever fits the task and your constraints — a hosted frontier model where quality matters most, a cheap one for classification, an open-weights model on your own infrastructure where data cannot leave. The architecture keeps that swappable, because the right answer changes every few months.

Can you run this alongside our own engineers?

Yes, and it is usually the better outcome. We will pair, review and hand over deliberately so your team can extend the system rather than inherit something they are afraid to touch.

Tell us which process is costing you.

Thirty minutes, no preparation, no deck. You describe what keeps going wrong and we tell you whether AI is the answer — including when it is not.

Prefer email? contact@devs-core.com