Skip to main content

AI agents

How to Choose an AI Agent Development Company

How to choose an AI agent development company: the proof to ask for, the integration test to insist on, 10 questions for the first call, and the red flags.

Mohammad Ashraful Islam8 min read
Choosing an AI agent development company: four checks between a demo and an agent in production, from proof and integration to permissions and ownership

Choose the AI agent development company that can show you an agent already running on real data, will test the connection to your own systems before quoting the build, can tell you exactly what the agent is allowed to change and who approves it, and hands you the code at the end. Most other differences between vendors matter less than those four.

This guide is for founders, operations leads and IT managers comparing vendors for a first AI agent project. It covers what to check, the questions to ask on the first call, and the warning signs. We build AI agents for businesses ourselves, so where it helps we use our own projects, including what went wrong, as examples. There is no ranking of companies here; a vendor ranking its own industry is not a source you should trust, us included.

Key takeaways

  • Ask for an agent in production with a number attached, not a demo.
  • Insist on a test against your real systems before the build is priced.
  • Get the agent's permissions in writing: what it may read, what it may change, and who approves each change.
  • Ask what went wrong on their last project after launch. A good vendor has an answer.

What proof should an AI agent development company show you?

Ask to see an agent that real users rely on today, with one number that shows it works, and ask how that number is measured. A demo shows that a model can answer a question. Production shows that the whole system survives real customers, messy data and the second month.

A good answer sounds like a specific system, a specific metric and a method. For example, an AI support agent we built for a US direct-to-consumer brand now handles 40% of its support volume on its own, out of 2,000–3,000 conversations a day. The number is measured by what reaches the human team, not by what the agent thinks it answered.

Be wary of these answers:

  • "Everything we've built is under NDA." Some of it will be. All of it should not be. A vendor can anonymise a client and still show the system and the numbers.
  • A metric with no method. "95% accuracy" means nothing until you know what was counted, over how many cases, and who checked.
  • Only internal tools. A vendor's own products are real evidence of skill, but they are not the same as delivering for a client with a deadline and a budget. Ask which is which.

Can they connect the agent to the systems you already run?

This is where most AI projects fail, so it is the most important thing to test. An agent is only useful if it can read your real data and act in your real systems: the ERP, the CRM, the helpdesk, the database nobody documented. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and in our experience the integration that nobody scoped is a common reason.

So ask the vendor to prove the connection before they price the build: read access to a sandbox or a copy of your data, a working test of the integration, and a written list of what your systems will and will not allow. That is what our AI Readiness Audit does in two weeks, and its findings belong to you whether you build with us or not.

A vendor who quotes a fixed build price without ever touching your systems is quoting a guess.

What will the agent be allowed to change, and who approves it?

Get this in writing before anything is built. An agent that only reads data is easy to approve. The moment it needs to create an order, change a price or close a ticket, someone in your business will ask what it did and how to undo it, and the project needs a good answer.

The answer we use, and the one to look for, has three parts:

  • Scoped permissions. The agent is granted named operations on each system, never a login with full access.
  • Reads pass, writes wait. Queries run straight away. Anything that changes a system of record goes to a review queue, where a named person approves or rejects it.
  • Everything is logged. Who asked, what ran, what came back, and who approved it.

We call this an access layer. Whatever the vendor calls it, ask them to show you where the approval happens and what the log looks like.

How do they handle the agent being wrong?

Ask what went wrong on their last project after launch, and what they changed. Every agent is wrong sometimes. What separates vendors is whether they planned for it.

Here is ours. On the support agent above, we designed the message categories from historical tickets, and they worked well in testing. After launch, real customers phrased the same problems in many more ways than the history showed, and the agent met questions its categories could not place. Because the category also decided which content to search, a wrong category meant a weaker answer. We refined the categories and retrieval rules from real conversations in the weeks after go-live.

The first weeks of real traffic teach you more about an agent than any test set. Budget for a tuning phase after launch.

Also ask what the agent does when it is not sure. The right answer is that it hands over to a person with the context attached, rather than guessing.

Will they tell you roughly what it costs?

A vendor who builds agents regularly can give you a range on the first call. The range depends mostly on how many systems the agent must connect to and how much it is allowed to change. For reference, a first agent covering one process typically starts in the mid four figures to low five figures, and a support agent connected to a helpdesk, a database, an app and a website sits nearer $20,000–$30,000.

Ask for two more numbers, because they matter as much as the build:

  • Running costs. Model and API usage, hosting and the vector store, per month, at your volume.
  • Maintenance. Who keeps the agent working when a model changes or a connected system updates, and what that costs.

"Every project is unique, let's schedule a discovery call" is fine as a next step. As the only answer to a price question, it usually means the price depends on what they think you can pay.

Who will actually build it, and who owns the code?

Ask who will be on the project by role, and ask for an engineer on the second call. A small, senior team is normal for this work. Our support agent was built in 45 days by four people: a system designer, a project manager and two AI engineers.

Then ask who owns the result. The answer you want is that the code, prompts, workflows and documentation sit in your own repository and accounts, and you can take them to another team. Watch for agents that only run on the vendor's platform, licence fees that never end, or prompts and workflows the vendor keeps.

Specialist, large consultancy or freelancer?

Each fits a different situation. Match the choice to the size of the risk, not the size of the logo.

  • Specialist AI agent company: one to three processes, existing systems to connect, and you want a working agent in weeks with a small senior team.
  • Large consultancy: a company-wide programme across many countries and departments, with a budget and timeline to match.
  • Freelancer: a narrow, well-defined build where you have someone in-house to own the architecture and the maintenance.
  • Off-the-shelf platform: your use case is standard (common support questions, scheduling) and the platform already covers most of it.

If a packaged product already does 80% of what you need, buy it. Custom work earns its cost where your process, data or systems are what make you different.

10 questions to ask on the first call

  1. Can you show me an agent in production, and the number that shows it works?
  2. How was that number measured?
  3. Will you test the connection to our systems before quoting the build?
  4. What will the agent be allowed to read, and what will it be allowed to change?
  5. Who approves the changes, and where do we see the log?
  6. What does the agent do when it is not sure?
  7. What went wrong after launch on your last project?
  8. Roughly what will it cost to build, to run each month and to maintain?
  9. Who will build it, by role, and can an engineer join the next call?
  10. Who owns the code, prompts and workflows when you finish?

Red flags when hiring an AI agent developer

  • No live system to show, only demos and videos
  • A fixed build price before anyone has seen your systems
  • "Full access" for the agent, with no approval step for changes
  • No answer to what happens when the agent is wrong
  • A long strategy phase before any working software
  • No mention of running or maintenance costs
  • The agent can only run on their platform, and the code stays with them

One flag is worth a follow-up question. Three is a reason to walk away.

Frequently asked questions

How much does it cost to hire an AI agent development company?

Most first agent projects for small and mid-size businesses cost from the mid four figures to around $30,000, depending on how many systems the agent connects to and how much it can change. Large consultancy programmes cost far more. Always ask for monthly running and maintenance costs as well as the build price.

How long does it take to build an AI agent?

A first agent covering one process usually takes four to eight weeks, including the integration work. Our support agent, connected to four systems, took 45 days from system design to deployment. Plan a few weeks of tuning after launch.

Should I hire an offshore AI agent development company?

Location matters less than proof, integration testing, permissions and ownership. Many good teams work across time zones. Check who your day-to-day contact is, which hours overlap with yours, and which country's law governs the contract.

What is the difference between an AI agent and a chatbot?

A chatbot answers questions. An AI agent can also act: look things up in your systems, create or update records, and hand work to people. That ability to act is why the integration and the approval process matter so much when you choose who builds it.


Comparing vendors for a first agent? Start with an AI Readiness Audit, or book a 30-minute call and ask us the 10 questions above.

One email when we publish something worth your time

No cadence, no drip sequence. We write when we have learned something running agents in production, which is not weekly.