The short answer: compare the first year's cost, then make them show their work
Choose the firm that can show you an agent people use in production, puts the first year's full cost in writing, and answers plainly about ownership, permissions, and what happens when the agent gets something wrong. To compare quotes, put each on a 12-month total: the build plus model usage, hosting, monitoring, maintenance, and changes. Ask every firm the same seven questions and drop the ones that can't answer. And if you can't yet name the workflow an agent should run, don't hire anyone yet.
I sell custom AI agent development myself, so this checklist comes from the other side of the table. My own answers are at the end.
Four ways firms price the work, and who pays when it runs long
Most quotes for a custom agent use one of four structures, sometimes two in a row. The structure decides who pays for an overrun, so read it before the number.
| Structure | Who pays for an overrun | Your risk | Get in writing |
|---|---|---|---|
| Fixed price for a defined scope | The firm | Anything missing from the scope becomes a change order | The definition of done, what’s excluded, and the rate for changes |
| Paid discovery or audit first | Nobody, if the fee and end date are fixed | A study that turns into a sales pitch, or never ends | An end date, a deliverable you keep, and whether the fee counts toward the build |
| Time and materials with a cap | You, up to the cap | Hitting the cap with the agent half built | Weekly hours against the cap, and what happens when it’s hit |
| Monthly retainer | You, if nothing ships | Paying for time instead of results | What ships each month, a short notice period, and no minimum term |
None of these is a red flag by itself, and time and materials is fair when nobody can know the scope up front, if the cap is real. My bias, since it's how I price: a short fixed-fee discovery, then a fixed price, so the scope is on paper before the big number.
Compare AI agent development pricing on one 12-month total
The build price is the number on page one, and rarely the whole cost. Ask each firm to fill in the same first-year lines, in writing, and add them up.
| Cost line | Count it as | Ask |
|---|---|---|
| Discovery or audit | One-time | Is it credited toward the build? |
| Build | One-time, fixed or capped | What’s in scope, and what’s out? |
| Model usage | Monthly estimate at your volume, × 12 | Whose account, who pays the bill, and is there a markup or a cap? |
| Hosting | Monthly × 12 | Whose cloud account, and what does the fee include? |
| Monitoring and logs | Monthly × 12 | Who reads them, and who gets the alert? |
| Maintenance | Monthly or yearly | Does it cover a model retirement, or a change in a connected system? |
| Change requests | Changes you expect × the rate | What’s the rate, and are any changes included? |
| Platform or license fees | Per seat or per agent, × 12 | What stops working if you stop paying? |
| First-year total | The sum of every line | Which lines did the quote leave blank? |
Usage is the easiest line to leave blank, and you pay it either way. OpenAI and Anthropic both bill API use per million tokens, input and output priced separately, and Anthropic counts an agent's tool definitions and tool results as tokens too, so an agent that looks things up before it acts costs more per task than a chat. Get an estimate at your real volume. On your own model account you see every bill; on the firm's, ask about the markup.
Quotes for what sounds like the same agent can land far apart, mostly because the scopes differ. The cost of AI automation covers what moves a build's price, with my own numbers.
Seven questions to ask an AI agent developer, me included
Ask them on the first call, and get the answers into the contract. Good answers are specific. A vague one is an answer too.
- 01
What have you shipped that people use every day, and can I see it?
A demo runs on clean data with its builder watching. Ask about an agent in production: who uses it, what it does without a person checking, how often someone steps in, and what happened when it last got something wrong. An NDA can hide a client's name, not whether the system exists.
- 02
Who owns the code, the prompts, the data, and the accounts?
Paying for software doesn't make it yours. Under U.S. copyright law, code an outside firm writes starts out as the firm's or its programmer's. The Copyright Office's guide to works made for hire says a commissioned work only counts as made for hire, making you the owner, if it's one of nine listed kinds and both sides sign an agreement saying so. Custom software isn't on the list. Otherwise, a transfer to you counts only if it's in writing and signed by the owner. So the contract should assign you the code, prompts, and configuration, bar the firm from reusing your data, and put the cloud, model, and code-hosting accounts in your company's name. General information, not legal advice: have your lawyer read that clause.
- 03
What can the agent read and change, and what waits for a yes?
Give an agent the least access the job needs. The OWASP Top 10 for LLM applications lists excessive agency (more functions, permissions, or autonomy than the task requires), and the fixes it gives include minimum access, a person's approval before high-impact actions, and logs of what the agent did. The list's first entry, prompt injection, shows why: text in a web page or file the agent reads can change what it does. Get the answer in writing, system by system: read, write, or ask first.
- 04
How will you know it got something wrong, and how do you undo it?
Anthropic's guide to building effective agents says an agent's autonomy means higher costs and mistakes that can compound, and recommends heavy testing in a sandbox, with guardrails. Ask what the agent is tested against before launch (real past cases from your business with known right answers beat a demo script), what score it must reach, what gets logged, who gets alerted, and how a bad action is reversed.
- 05
Which model does it run on, and what happens when that model is retired?
Model makers retire old models on a schedule. Anthropic's deprecation page says requests to a retired model fail, and it gives at least 60 days' notice for publicly released models. OpenAI's lists at least six months for generally available ones. Ask who watches for the notice, who tests the replacement, who pays for it, and whether the agent could switch providers without a rebuild.
- 06
Where does our data go, and who can see it?
List every service your data passes through (the model maker, hosting, logging, any search index) and read the terms. Anthropic says it doesn't train its models on API inputs or outputs by default; OpenAI says API data isn't used for training unless you opt in, though it keeps abuse-monitoring logs for up to 30 days by default. Then ask what the firm itself keeps, where, and for how long.
- 07
What happens after launch, and if we part ways?
Ask what support costs, how fast it responds, and what it covers: bugs, model retirements, changes in connected systems. Then ask about leaving: you should walk away with the code, documentation, test cases, and accounts, in a state someone else can run. If leaving means starting over, you were renting.
Six red flags that should end the conversation
One of these is reason enough to keep looking.
- No production references. Only demos and slides, or an NDA where a working system should be.
- An accuracy guarantee. OWASP lists misinformation, false output that looks credible, among the top ten LLM risks. Nobody can promise a rate on your data before measuring it; the honest offer is a measured score on your own cases and a plan for the misses.
- A platform you can’t leave. If the agent runs only on the firm’s proprietary platform, you’ll pay its fee for as long as you use the agent.
- No answer to what the agent may do. If they can’t list what it reads, writes, and asks permission for, it hasn’t been designed yet.
- No plan for usage costs. No estimate, no cap, no alert. OWASP puts unbounded consumption, including runaway pay-per-use bills, on the same top ten, with rate limits, quotas, and monitoring among the fixes.
- No price until the fourth meeting. A firm that has built similar agents can give you a range on the first call. You can’t compare a number you haven’t seen.
When you shouldn't hire an AI agent developer yet
Anthropic's agent guide advises finding the simplest solution that works, and says that might mean not building an agent at all. I agree. Three signs you're there:
- You can't name the workflow. “Use AI somewhere” isn't a project. Price the manual work first: what to automate first has the arithmetic.
- A product already does the job. Missed calls, booking, reminders, website questions: common jobs have prebuilt tools at a monthly price, and your current software may already include one. Try that first.
- A chatbot would do. If all you need back is an answer, and a person takes it from there, you need a chatbot, and that's a subscription, not a build. AI agents vs. chatbots draws the line.
How I answer the same seven questions
Hold me to the same list. Where my published terms answer a question, here it is.
- Shipped work. The production systems I've built are written up in my case studies.
- Ownership. On a custom build you own the code, infrastructure, and documentation. It's deployed on your infrastructure, with no license fee on your own workflow.
- Permissions. What the agent may read and write is mapped in the design step, before any code, and you set the permissions for each system. Its actions are logged.
- Price. Work starts with a one-week operations audit at a fixed $4,500. If it doesn't surface savings worth more than its cost, you don't pay, and if you build, the fee is credited toward it. Builds are fixed scope from $45,000. Ongoing work is from $8,500 a month, month to month.
- After launch. Your team gets documentation and training. Support is available but not required, and nothing is held hostage to a retainer.
- The rest. The model, the test cases, and who pays usage depend on your workflow. Ask me, and get the answers in writing before the build starts.