How to Choose an AI Development Partner: Seven Questions Before You Sign

GuidesDevelopment

Most scorecards compare portfolios, team size and references, and every serious vendor passes on all three. What separates two candidates is what happens after the work is delivered: who owns the result, what arrives at handover, and who pays when the model drifts.

You do not have to invent those criteria. The largest software buyer in the world has published its own, and the requirements are public. This piece turns them into seven questions you can run in one call, with the answers that should worry you.

How to Choose an AI Development Partner: Seven Questions Before You Sign

Why the usual scorecard fails to separate two candidates

Portfolio, headcount, references, tech stack. Any team worth a shortlist clears all of it, which is why the shortlist rarely gets shorter.

The problem is structural rather than personal. In Deloitte’s outsourcing research, 70% of executives say their vendor management function is not fully mature, so most buyers are running these decisions without a settled process behind them. The same research finds 70% have selectively brought work back in-house over five years, which means the boundary moves and the exit matters.

Pressure makes it worse. IBM’s study of 2,000 chief executives found 64% invest in AI because they fear falling behind rather than because they understand the value, and half admit they ended up with a patchwork of disconnected technology.

So the question of how to choose the right ai development partner starts somewhere other than capability. The criteria that matter are not the ones that prove a team can build. They are the ones that tell you whether you will be left with a working system or with a dependency. If you are still deciding between an internal team and an outside one, our cost comparison covers that step and this article picks up after it.

The checklist already exists, and the largest buyer published it

In April 2025 the US Office of Management and Budget issued memorandum M-25-22, which governs how federal agencies buy AI. Federal procurement is not a template for a mid-sized company, but it is written by people who buy more software than anyone else and who have to live with what they sign.

Five requirements sit at the centre of it, and they read as a checklist for anyone hiring a team to build this. Agencies must address intellectual property rights, the use of government data, protection against vendor lock-in, continuous testing and monitoring, and notification when new AI components are introduced into a delivered system.

The same memorandum names what reduces lock-in specifically: knowledge transfer, portability of data and models, clear licensing terms and price transparency.

What the memorandum requiresWhat it means when you are the buyer
Intellectual property rightsName who owns the model, the prompts and the tuning work, in writing
Use of your dataSay whether your data may train anything beyond your own system
Protection against lock-inAsk for portability of data and models, and for licensing you can read
Continuous testing and monitoringAgree who measures quality after launch, and against what
Notification of new AI componentsRequire being told when a subcontracted model appears inside your system

That is a usable checklist before anyone writes a proposal, and it carries more weight in a negotiation than anything a supplier writes about itself.

Seven questions to ask before you sign

This is the part no scorecard can do for you. Ask these on a call, in this order. Each one has an answer that should reassure you and an answer that should not.

1. Who owns what we paid for?

A good answer names an explicit assignment of rights in the contract. A poor one says the agreement includes a work-made-for-hire clause, because in the United States that clause does not cover commissioned software: the Copyright Office lists nine categories of commissioned work that can qualify, and software and models are not among them. British law reaches the same place differently, since the author is the first owner unless the contract says otherwise, and the employment exception covers employees rather than contractors.

2. What exactly gets handed over at the end?

A good answer is a list of artefacts, and it should include all four of these:

  • The evaluation set. The cases, the expected answers and the scoring rules, in a repository you control.
  • Your data in a format you can read. Not access to a dashboard that renders it for you.
  • Prompts with versions and the reasoning attached. A prompt without the note explaining why it is worded that way is an incantation.
  • A deployment recipe. What runs where, written down well enough that another team could rebuild it.

A poor answer is “the source code and documentation”, which sounds complete and leaves out everything that makes the system judgeable.

3. Where does our data live, and who can reach it?

A good answer names the regions, the subprocessors and whether your data trains anything. Worth knowing that this question also decides your timeline: in G2’s 2026 survey of 1,038 buyers, security review is the longest delay after a supplier is chosen, affecting 39% of deals and half of large ones. Starting that review late is the most common self-inflicted delay in these projects.

4. Who answers when the model is wrong?

A good answer describes a review step, a threshold above which a human sees the case, and a named owner. There is a legal edge here that surprises buyers: under article 25(1) of the EU AI Act, you can become the provider of an AI system yourself if you put your name on it, modify it substantially or change what it is used for. Obligations you thought sat with your ai development partner can move to you without anyone renegotiating.

5. What does it cost to run, rather than to build?

A good answer splits the number in two: the one-off build, and the monthly cost of model calls, hosting, monitoring and adjustment. A poor answer gives one figure. The second half is where budgets are surprised, because model calls are charged per request and grow with use.

6. Who watches quality in six months, and at whose expense?

A good answer names a person, a schedule and a measurement. This is not a hypothetical risk. Chen, Zaharia and Zou measured GPT-4 on the same set of tasks three months apart: the March 2023 build answered correctly 84% of the time and the June build managed 51%, with no API change and nothing to deploy on the customer side. A system can get worse while nobody touches it.

7. What happens when the model is retired?

A good answer shows the team has read the model supplier’s deprecation policy and planned for it. OpenAI commits to a minimum of six months’ notice before removing a generally available model. Anthropic commits to at least sixty days, and claude-opus-4-1 was marked deprecated on 5 June 2026 and switched off on 5 August, exactly two months later. Those floors are shorter than most annual plans.

What belongs in the contract rather than the call

Seven good answers on a call are worth nothing by themselves. Four of them have to survive into the document: ownership of the result, the list of what is handed over, portability of data and models, and the terms under which either side can stop.

Contract length is its own protection, and buyers have worked this out. In the same G2 research, 70% prefer shorter contracts and 29% want terms under twelve months. A short first term costs a team that intends to stay nothing at all.

Write acceptance around an outcome rather than hours. The market has already moved that way: Deloitte reports 67% of organizations now buy outcomes instead of time, against 45% two years earlier.

What you have to prepare on your side

Half of what delays these projects sits with the buyer, and an honest look at an ai development partner includes a look in the mirror.

Integration is heavier than it looks from the outside. The 2026 Connectivity Benchmark, based on 1,050 IT leaders, counts 957 applications in the average enterprise with 27% of them integrated, and both numbers moved the wrong way over the year.

Then ownership of the data itself. In dbt Labs’ 2026 survey, 41% of organizations have no defined owner for their data, unchanged year on year, which is why an export that was promised in a week takes six. And readiness is rarer than confidence: Cloudera and Harvard Business Review Analytic Services found 7% of enterprises consider their data fully ready for AI while 73% admit difficulty preparing it.

None of this needs fixing before you start. It needs naming, because a team that plans around it gives you a real date, while one that ignores it gives you a date that slips.

When you do not need a vendor yet

Some projects are not ready for procurement, and the evaluation is wasted on them.

A prototype with a month to live, an internal tool three people use, or a process nobody has written down yet. In the last case the order is wrong: describe the process first, because a vendor cannot automate a decision that only exists in somebody’s head, and you will pay them to discover that.

If you are a step earlier than that, and the open question is whether this process wants AI at all, our guide to how to implement AI in business covers where it pays off and where it is too early. Come back to these seven questions once the process is chosen.

We say the same thing on our own site, that we suggest a simpler answer when the task has one. That conversation goes better before an AI project has been announced internally.

What to do next

Take one process that actually hurts and put these seven questions to two or three teams. You are not looking for perfect answers. You are looking for the ones who answer concretely and do not bristle when asked what happens at the end of the work.

Then ask to see a working system on a screen rather than on slides. A live interface with real data, anonymised if it has to be, tells you more about an ai development partner than any portfolio page.

We built our own products before we built anyone else’s. Sorting incoming documents in DocuFlow used to take a full working day and now takes a couple of hours, which means we have answered these seven questions about ourselves at least once.

If you want to go through the list together, we run a free review: fifteen to thirty minutes on your task, your process and your data, with no obligation on either side.

  1. Why the usual scorecard fails to separate two candidates
  2. The checklist already exists, and the largest buyer published it
  3. Seven questions to ask before you sign
  4. What belongs in the contract rather than the call
  5. What you have to prepare on your side
  6. When you do not need a vendor yet
  7. What to do next