The integration was never scoped
The demo ran on a spreadsheet and nobody checked whether the real system would cooperate. The most common diagnosis by a distance.
An agent designed around a write endpoint the ERP does not expose.
Capability
For the AI pilot that impressed everyone in the demo and then stalled before production. We take it over, find what it is actually missing, and either get it live on a fixed scope or tell you to stop.
From $3,500 · 2–3 weeks to diagnosis · fixed scope, agreed in writing
What your team gets
Not a strategy deck. A running system, the permissions around it, and the code in your repository.
What we build
We have yet to find a seventh. The diagnosis is usually one of these, and occasionally two at once.
The demo ran on a spreadsheet and nobody checked whether the real system would cooperate. The most common diagnosis by a distance.
An agent designed around a write endpoint the ERP does not expose.
It works, it is accurate, and it stalled at read-only because there was no mechanism for a person to review what it wanted to change.
A finance lead who cannot see or undo what the agent would do, so it never gets to.
It seems good. Nobody can say how good, because there is no held-out set and no agreed definition of correct. Without that, it cannot be approved and cannot be improved.
A classifier everyone likes and nobody can defend in a meeting.
A prompt chain in a notebook, one contractor who understood it, and no way to tell when the output drifted after a model update.
Output quality that quietly changed three months ago and nobody noticed.
It works, and at production volume the model bill exceeds the saving. Usually recoverable by routing cheap steps to cheap models, sometimes not.
A workflow costing more per document than the person it replaced.
Built well, aimed badly. The automated step was not the expensive one, and nobody measured before starting.
A ten-minute task automated while the two-day reconciliation continues by hand.
How we build it
A rescue is diagnosis before treatment. We do not start rewriting anything until we can say, in writing, what is actually wrong — and sometimes the finding is that the existing build is sound and the blocker is organisational rather than technical.
The code, the prompts, the logs if there are any, and the history. Much of the answer is usually visible in the first day, in what was never written down.
A real read and a real write against the systems it was meant to reach. This is where the majority of stalled pilots turn out to have been stopped.
A held-out set built from real cases, and a number where there was previously an impression. You cannot approve what you cannot measure, and neither can your risk team.
Fix, rebuild, or stop — each with a cost and a timeline. Sometimes the existing work is 80% there; sometimes it is a demo that cannot become a system, and saying so is the value.
If it is worth shipping, we ship it on a fixed scope. If it is not, you get the findings and owe nothing further.
Prompts and the domain knowledge inside them are usually worth keeping — someone spent real effort learning what the model needed to be told. Evaluation cases, where any exist, are gold. The orchestration layer is the part most often rewritten, because notebook-era code rarely survives contact with retries, concurrency and logging.
We recommend stopping when the automated step was not the expensive one, when the model cost per run exceeds the saving and cannot be routed down, or when the system it depends on genuinely cannot support the write. That verdict has cost us build work more than once, and it is the reason clients come back with the next project.
Most rescues involve someone’s existing work, sometimes an in-house team and sometimes another vendor. We write the diagnosis about the system rather than about the people, and we would rather hand them a plan they can execute than take the work. That is not always what happens, but it is where we start.
Where it lands
A defect-detection pilot with no measured accuracy and no route to the line.
An ordering agent that stalled the moment it needed to write to the ERP.
A support assistant that is accurate in testing and unapproved in production.
A document pipeline blocked by audit requirements nobody designed for.
A member chatbot built on an export that is now six months stale.
An extraction pilot with no confidence threshold and no human check.
Technology
We pick the boring option unless there is a reason not to — the framework is the part most likely to be abandoned before your system is.
How the engagement runs
Every phase ends in a named deliverable. You can stop after any of them and keep what has been built.
What was built, what it does today, and where it stopped. We say whether a rescue is the right shape or whether you are better starting again.
The code, prompts and logs reviewed, and the integration tested properly against the systems it was meant to reach — usually for the first time.
A held-out set built from real cases and an accuracy number produced, so there is something to approve or reject rather than an impression.
The written diagnosis, what is worth keeping, and each path — fix, rebuild, stop — priced with a timeline.
The write path, the permission model, evaluation in CI, logging and the deployment — on a fixed scope quoted from the diagnosis.
Why Devs Core
Fixed price for the diagnosis, and no obligation to continue. We would rather be right about what is wrong than be hired to rebuild it.
The integration. It is the most common cause of a stalled pilot and the least often checked before the build began.
A held-out set and a real accuracy number, because unmeasured accuracy is why a working pilot cannot get approved.
When the step was not expensive enough, the unit economics do not work, or the system cannot accept the write. That verdict has cost us build work.
Prompts and evaluation cases usually carry real knowledge. Starting from zero is rarely the right answer and always the more profitable one for a vendor.
Whether we build the fix or not. Several rescues have ended with the client’s own team executing the plan.
Questions
AI build rescue is taking over an AI project that was built, works in some form, and has not reached production. It begins with a fixed-price diagnosis: reviewing what exists, testing the integration it depends on, measuring accuracy against real cases, and producing a written go / no-go with the cost of each path — fix, rebuild, or stop.
Almost always one of six reasons: the integration was never scoped, nobody would approve the write, accuracy was never measured so it cannot be approved, there is no owner or tests or logs, the cost per run exceeds the saving, or it automated a step that was not expensive in the first place. The diagnosis establishes which.
AI Build Rescue starts at $3,500 for a diagnosis delivered in two to three weeks. If the recommendation is to proceed, the production work is quoted separately as a fixed scope from the findings. If the recommendation is to stop, you keep the findings and owe nothing further.
No. The diagnosis is written so your own team or another vendor could act on it, and several rescues have ended exactly that way. We are paid for the diagnosis regardless of who executes the fix.
Where they are still involved, yes, and we write the diagnosis about the system rather than about the people. Handing an in-house team a plan they can execute is a better outcome than taking the work, and we start there.
We will say so, and say what is worth carrying forward. Prompts and evaluation cases usually contain real domain knowledge worth keeping even when the code is not; notebook-era orchestration rarely survives contact with retries, concurrency and logging.
Usually. Code, prompts and whatever logs exist tell most of the story, and a week is normally enough to establish what it does and where it stops. Absent documentation is itself a finding, because a system nobody can explain is one nobody can safely operate.
The audit decides what to build before anything exists. The rescue diagnoses something already built that has not shipped. If you have a working pilot that stalled, this is the right one; if you are still deciding where to start, the audit is.
Thirty minutes, no preparation, no deck. You describe what keeps going wrong and we tell you whether AI is the answer — including when it is not.
Prefer email? contact@devs-core.com