How to Choose an AI Development Company: 15 Questions to Ask Before You Sign
Short answer
Choose an AI development company by testing for production experience, not demo quality. Ask how they evaluate accuracy, handle model failures, secure data, monitor systems after launch, and price the work. Insist on code and IP ownership, direct access to engineers, a paid prototype on your data, and references for systems that are live today.
Key takeaways
- A great demo proves little. Ask for evidence of systems running in production with real users.
- The best teams talk about evaluation, error handling and monitoring before they talk about models.
- Insist on owning the code, prompts and documentation, with no lock-in to proprietary platforms.
- Start with a small paid prototype on your own data before committing to a full build.
Two years ago, "AI development company" described a small group of specialist firms. Today almost every software agency, freelancer platform and consultancy uses the phrase. Some have deep production experience. Many have built a few chatbot demos on top of an API.
From the outside, they look the same. Websites promise "custom AI solutions." Demos are impressive. Prices range wildly.
This guide gives you 15 questions that cut through that. They're the questions we'd ask if we were hiring an AI partner ourselves, and they're organised so you can turn them into a scoring sheet.
Why is choosing an AI development company harder than hiring a regular dev shop?
Because AI systems fail in ways normal software doesn't. A regular web app either works or throws an error. An LLM-based system can look like it's working while quietly giving wrong answers, leaking data or running up a large API bill.
That means the skills that matter most are not visible in a demo:
- Measuring whether answers are correct, at scale
- Handling unpredictable model output safely
- Protecting against prompt injection and data leakage
- Monitoring cost, latency and quality after launch
Industry research reflects this gap. Gartner has predicted that a significant share of generative AI projects would be abandoned after proof of concept, citing poor data quality, unclear business value, inadequate risk controls and escalating costs. Our article on why AI pilots fail to reach production goes into the causes. The right partner is one who plans for those failure modes from day one.
What questions should you ask about production experience?
1. "Which of your AI systems are running in production today, and can we speak to the client?"
This is the most important question on the list. A demo or proof of concept proves the model can do something once. A production system proves the team can make it work every day, for real users, under load. Ask how long systems have been live, how many users or transactions they handle, and what broke after launch.
2. "What went wrong on a past project, and what did you change?"
Experienced teams have war stories: a rate limit at peak traffic, a document format nobody anticipated, a prompt change that broke an edge case. If a vendor can't name a failure, they either haven't shipped much or aren't being straight with you.
3. "What does your standard production architecture look like?"
Listen for specifics: how state is stored, how failures are retried, how logs and traces are collected, how secrets are managed, how deployment works. Vague answers ("we use the latest AI stack") suggest each project starts from scratch.
How do they measure whether the AI actually works?
4. "How will we know the system is accurate before launch?"
The right answer involves an evaluation set: a collection of real examples with known correct outputs, scored automatically each time the system changes. Our guide to evaluating LLM applications explains the approach. If the answer is "we'll test it manually" or "the model is very good," that's a warning sign.
5. "Will you promise an accuracy number?"
Counterintuitively, a good partner usually won't — at least not before seeing your data. They'll offer to measure accuracy on a sample during a prototype phase and agree targets from there. Anyone guaranteeing "99% accuracy" in a sales call is guessing.
6. "How do you monitor quality after launch?"
Models change, user behaviour shifts and data drifts. Ask how they track answer quality, cost and latency in production, and who gets alerted when something degrades. See LLM observability for what good looks like.
How do they handle security, data and risk?
7. "How do you protect the system against prompt injection and data leakage?"
Any system that reads user input or external documents is exposed to prompt injection. A credible team can explain their mitigations: input and output filtering, least-privilege tool access, separating instructions from data, and human approval for sensitive actions.
8. "Which model providers will process our data, and under what terms?"
You should know whether your data goes to OpenAI, Anthropic, Google, an open-source model you host, or several of them, and what each provider's retention and training policies are. For regulated data, ask about data residency, encryption and whether a business associate agreement (for US healthcare) or data processing agreement is in place.
9. "What can the AI do without a human approving it?"
The answer should be specific and deliberate. Good partners design guardrails and human-in-the-loop steps around risky actions like sending messages, moving money or changing records.
What should the commercial terms look like?
10. "Who owns the code, prompts and models?"
You should. Ask for this in writing: full repository access, prompts, evaluation datasets, infrastructure-as-code and architecture documentation. Be cautious of vendors whose solution only runs on their proprietary platform.
11. "Is this fixed scope or hourly?"
Both models can work, but for a defined first project, fixed scope with clear deliverables protects you from open-ended bills. If the vendor proposes hourly billing, ask for a capped estimate and weekly reporting.
12. "What will this cost to run each month?"
Build cost is half the picture. Ask for an estimate of LLM usage, hosting and maintenance at your expected volume. Our AI agent cost breakdown shows how to sanity-check the answer.
How will you work together day to day?
13. "Will we talk directly to the engineers building the system?"
Layers of account managers slow decisions and lose detail. The best results come when your domain experts and the engineers talk directly, especially during the prototype phase.
14. "Can we start with a paid prototype on our own data?"
A two-week prototype on real data is the cheapest way to test both the idea and the partner. You'll learn how they communicate, how they handle ambiguity and whether the model can actually do the task.
15. "What happens after launch?"
Ask about the handover process, documentation, training for your team, bug-fix warranty and the cost of ongoing support. A partner who disappears at launch leaves you with a system nobody understands.
What are the red flags when hiring an AI agency?
Walk away, or at least slow down, if you see:
- No verifiable production references. Only demos, mockups or "confidential" clients.
- Model-first sales pitch. All talk about which LLM they use, nothing about testing, monitoring or failure handling.
- Guaranteed accuracy before seeing data.
- Lock-in. The system only runs on their platform, or you don't get the source code.
- No questions about your process. A good partner spends the first call understanding your workflow, data and constraints, not presenting slides.
- Unclear pricing. No defined deliverables, no estimate of running costs.
A simple scoring sheet
Score each vendor 0–2 on each of the five areas below (0 = weak or no answer, 1 = adequate, 2 = strong, specific evidence), for a maximum of 10:
| Area | Questions | What a "2" looks like |
|---|---|---|
| Production experience | 1–3 | Live systems, reference calls, specific architecture |
| Evaluation & quality | 4–6 | Evaluation sets, measured targets, production monitoring |
| Security & risk | 7–9 | Named mitigations, clear data flows, deliberate guardrails |
| Commercials | 10–12 | Full ownership, fixed scope, running-cost estimate |
| Collaboration | 13–15 | Direct engineer access, paid prototype, clear handover |
A vendor scoring 8 or more is worth a paid prototype. Anyone weak on evaluation or security should be a "no" for anything customer-facing, regardless of price.
Agency, freelancer or in-house team?
| Option | Best for | Watch out for |
|---|---|---|
| Specialist AI agency | First production system, defined projects, speed | Lock-in, unclear ownership (solve with contract terms) |
| Freelancer | Small, well-scoped tasks or prototypes | Single point of failure, limited ops and security coverage |
| In-house team | AI as a core, long-term product capability | Hiring time and cost; senior AI engineers are scarce |
Many companies combine them: an agency builds and documents the first system, then an in-house team takes over. If you're still deciding whether to build at all, read build vs buy AI.
How we answer these questions at Keyved
We wrote this list partly because it's how we'd want to be evaluated. In short:
- Production first. Our systems run on a shared platform foundation with state management, retries, tracing and alerting built in, and you can see examples on our projects page.
- Evaluation from week one. We build an evaluation set from your real examples during the prototype, and every change after that is measured against it.
- You own everything. Code, prompts, evaluation data and architecture documents are yours, with no proprietary runtime.
- Fixed scope, direct access. Most builds ship in 4–7 weeks on a fixed scope, and you work directly with the engineers. See how we run MVP and prototype engagements.
If you're comparing partners, we're happy to be one of them. Book a call and bring these 15 questions.
Frequently asked questions
What should I look for in an AI development company?
Look for production systems you can verify, a clear evaluation and testing process, strong software engineering fundamentals (security, monitoring, error handling), transparent fixed-scope pricing, full code ownership for you, and direct access to the engineers doing the work.
Should I hire an AI agency or build an in-house team?
An agency is usually faster and cheaper for a first production system or when you need specialised skills for a defined project. An in-house team makes sense when AI is core to your product and you have continuous, long-term work. Many companies start with an agency and hire as the system matures.
What are red flags when hiring an AI developer?
Red flags include no production references, guaranteed accuracy numbers before seeing your data, no mention of evaluation or monitoring, proprietary platforms that lock you in, unclear IP ownership, and open-ended hourly billing with no defined deliverables.
How do I compare AI development proposals?
Compare what is included, not just price: integrations, guardrails, evaluation, monitoring, security review, documentation, handover and post-launch support. Ask each vendor to estimate monthly running costs at your volume.
Is it safe to share company data with an AI agency?
It can be, with an NDA, a data processing agreement, least-privilege access, and clarity on which model providers will process your data and under what retention terms. A good partner will raise these points before you do.