Economics and buying Checklist
15 Questions to Ask an AI Development Company Before You Sign
Choosing an AI development company comes down to what you will own, what the firm can prove and how easily you can leave. Fifteen questions, and a scorecard that weights them, separate teams that have run AI in production from teams that have demoed it.
For CTOs, procurement leads and general counsel choosing between shortlisted AI development companies before a contract is signed.
The short answer
To choose an AI development company, ask every shortlisted firm the same questions in five areas and score the answers: ownership of code, prompts, evaluation sets and model artifacts; production evidence over demos; data rights and security; delivery and operations; and regulatory and exit terms. The deciding test is whether another team could run the system next month from what the contract says you receive.
Key takeaways
- Name every AI asset in the ownership clause: prompts, evaluation sets, fine-tuned weights, vector indexes and pipelines, held in your repository and your environment from the first commit.
- Have a system built and put it into service under your own name, and the EU AI Act makes you its provider. The contract has to oblige the builder to supply what you need to comply.5
- Ask for production evidence and operator references. Gartner estimates only about 130 of the thousands of vendors selling agentic AI are real.1
- Weight certifications by how much of your data and production the vendor will hold. None of them shows that the system built for you works.
- Run the exit test before you sign: another team should be able to operate the system next month from what the contract says you receive.
Every firm on a shortlist has a capable deck and a demo that works. Neither shows what you will own when the contract ends, whether the team has kept an AI system healthy in production, or how hard leaving will be. The fifteen questions below test those things, plus data handling and delivery. Ask for written answers and score them on the weighted scorecard below, so procurement, security and legal judge every firm against one bar.
- 130of the thousands of vendors selling agentic AI are real, by Gartner's 2025 estimate1
- 48%of breaches in Verizon's 2026 report involved a third party, up 60% in a year2
- 40%+of agentic AI projects Gartner expects to be canceled by the end of 20271
How to choose an AI development company
Choose the firm whose answers you can verify. Vendor security questionnaires cover part of this, but they were written for buying AI as a service. A build raises more, because the vendor writes code that becomes yours, builds components from your data and may hold production access for years. The questions come in five groups of three, each with a good answer and a red flag. They apply equally to an AI agent development company; because an agent takes actions, weight evaluation and security higher.
Questions about code, IP and model ownership
Ask about ownership first, because it is the hardest term to fix after signature and the one that decides whether you can leave.
1. Will the code live in our repository and run in our environment from the first commit?
- Good answer Yes. You create the repository and grant named engineers access. The system deploys to your cloud account or your data center, bills are in your name, and you can revoke the vendor's access in one step.
- Red flag "We hand everything over at the end." Until then you cannot audit the work or leave mid-project.
2. Who owns the prompts, the evaluation set, the model weights and the vector index?
An AI build produces assets software contracts rarely name: the prompt library, the graded evaluation set, fine-tuned adapter weights, the vector index built from your documents and the pipelines feeding them. They hold most of what you paid for. The US Copyright Office concluded in January 2025 that copyright does not extend to purely AI-generated material, and that prompts alone do not make a person the author.3 Where code and prompts are produced with AI tools, your protection rests on the contract, so the assignment must cover every deliverable however it was made.
- Good answer Each asset is listed as a deliverable, assigned to you as it is created, stored in your environment and documented well enough to rebuild: the chunking rules and embedding model version for the index, the base model and training data for any adapter.
- Red flag "Joint IP," a perpetual license to "learnings," or a carve-out for "aggregated or anonymized data." Each can carry your evaluation set and prompts into work sold to others.
3. What will we still license from you after go-live, and can we change the model?
- Good answer Nothing proprietary: your infrastructure, open standards and components you can license directly. The model sits behind an interface, so moving to another frontier model or to an open-weights model is a configuration change and an evaluation run.
- Red flag "Our platform accelerates delivery." A runtime, orchestration service or console you must keep licensing is a permanent cost and an exit barrier.
Questions that test production evidence
Production evidence means a system running today, measured results from it and people outside the vendor who depend on it. A demo shows one good answer; these questions show whether the team can keep good answers coming for a year.
4. Which of your AI systems are in production, and may we speak to their operators?
- Good answer Named systems, what each does, how long it has been live and at what volume, and a client-side contact who operates it day to day. An engineering or operations lead tells you more than the executive sponsor.
- Red flag Case studies with no system named, references who only saw the pilot, or a demo offered in place of a production walkthrough.
5. Can we see an evaluation report from a live system?
- Good answer A graded set built from real cases, including hard ones; pass rates by category; the threshold that blocks a release; and a release it blocked. For your build, the set is written with your experts, and its threshold becomes the acceptance criterion in the statement of work. Our LLM evaluation guide shows how to build one.
- Red flag "We test thoroughly." Judging quality by whether recent outputs looked right misses the regression a model update brings.
6. What broke in production on your last AI system, and how did you find out?
- Good answer A specific incident, the alert or trace that surfaced it, the fix, and the test added so it cannot recur. Teams that have carried the pager tell these stories readily.
- Red flag Generalities, or "nothing significant." Every production AI system fails in some way; a team with no failure to describe has not operated one for long.
Questions about data rights and security
Data and security answers should name providers, terms and controls, because a build partner is a third party with access to what you guard most closely. Verizon's 2026 Data Breach Investigations Report found a third party involved in 48% of breaches, up 60% on the year before.2
7. Where will our data go, and can anyone train on it?
An AI build sends data where ordinary software does not: model endpoints, embedding services, evaluation tools and the coding assistants the vendor's engineers use. The same Verizon report found frequent employee use of AI tools rose from 15% to 45% in a year, and ranked unapproved "shadow AI" the third most common non-malicious data leakage activity.2
- Good answer Every model provider and sub-processor named, with region, retention and a written no-training commitment, inside your data processing agreement. Under the GDPR, a processor may engage another only with your written authorization, and must delete or return personal data when the work ends.4
- Red flag Unnamed "enterprise-grade providers," unstated retention, or no rule on which AI tools engineers may use with your code.
8. How will your team's access to our systems be granted, logged and revoked?
- Good answer Named individuals on your identity provider, with scoped roles, multifactor authentication and no shared accounts; masked or synthetic data until a stated milestone; every session logged in your systems; access removed the day a person leaves the project.
- Red flag A shared service account, credentials sent over chat, or your production data copied into the vendor's environment.
9. Which certifications do you hold, and what do they cover?
- Good answer A current SOC 2 Type 2 report or ISO/IEC 27001 certificate whose scope covers the team that will touch your work, shared under NDA with the exceptions intact, and ISO/IEC 42001 where the vendor will run AI systems for you.
- Red flag A logo on the website, a Type 1 report presented as Type 2, or a certificate scoped to an office that will not do your work.
Questions about delivery and operations
Delivery answers show who is accountable after signature and after launch, once the sales team has moved on.
10. Who exactly will build this, and who will run it after launch?
- Good answer Named engineers with roles and time allocation, the same people at the evidence session and at kickoff, and a named owner for operations after launch: the vendor under a managed service, or your team after handover.
- Red flag A capability slide in place of names, one team in the pitch and another on the work, or no plan for who is on call.
11. How will we know when quality drops, and what happens when the model changes?
- Good answer Sampled live cases scored against the evaluation set; alerts on quality, latency and cost per task; and a written model-change procedure: run the evaluation set on the new version, compare, release only on a pass.
- Red flag "We check manually when something changes." Quality degrades after the invoice is paid, and an unprepared team hears about it from your customers.
12. What would make you tell us not to build this?
- Good answer Specific conditions: data that cannot support the use case, an error cost the business cannot absorb, or a process that rules or conventional software would handle better. A firm that says this while selling will say it when a milestone slips.
- Red flag No answer. A vendor that has never talked a client out of a project will build whatever you fund, including the version that cannot work.
Questions about regulatory roles, run cost and exit
The last three questions settle who carries the regulatory duties, what the system costs to operate and what happens when the relationship ends.
13. Under the EU AI Act, who is the provider, and what will you hand over?
Usually you. The AI Act's provider definition covers an organization that "has an AI system ... developed" and puts it into service under its own name; a deployer uses a system under its authority.5 For a high-risk system, provider duties include a quality management system, documentation, logs and a conformity assessment.5 Article 25, as amended in 2026, requires the provider and any supplier of models, tools, services or components to specify in a written agreement the information, technical access and assistance the provider needs to comply.56 High-risk obligations now start on December 2, 2027 for Annex III systems and August 2, 2028 for AI in regulated products.6 Our EU AI Act compliance guide covers the duties in full.
- Good answer A written statement of who is the provider, and a schedule of vendor deliverables: technical documentation, logging design, data and evaluation records, test results and support through conformity assessment. The European Commission's model contractual AI clauses, drafted for public procurement and updated in March 2025 in full and light versions, are a sound starting point.7
- Red flag "The AI Act is the client's concern; we only build." Every duty stays with you, and nothing in the contract helps you meet it.
14. What will this cost to run per completed task, and what drives that number?
- Good answer A running-cost model per completed task at your volumes, built from its drivers: model calls and tokens per task, retries, retrieval, the human review rate and infrastructure. It has a budget alert and a named owner, and measured figures replace estimates after the pilot.
- Red flag Silence on running cost, or no idea how many model calls one task makes. Inference is the cost line that grows with adoption.
15. What happens when we leave?
- Good answer Exit terms in the contract: a transition period during which the vendor keeps the system running, a handover package defined as a deliverable, knowledge-transfer sessions, return of every asset and certified deletion of your data from the vendor and its sub-processors.
- Red flag Exit handled by a single termination clause, handover "to be agreed," or deletion left to trust.
A weighted scorecard for AI vendor assessment
Score each firm from 1 to 5 on six criteria, multiply by the weights, and treat a 1 on ownership or on data as a knockout whatever the total.
| Criterion (questions) | Weight | What earns a 5 | What earns a 1 |
|---|---|---|---|
| Ownership and portability (1 to 3) | 25 | Your repository and environment from the first commit; every AI asset named and assigned; the model swappable | Handover at the end; joint IP or a license back; a platform you must keep licensing |
| Production evidence (4, 6) | 20 | Named live systems; operators who confirm volume and incidents; a specific failure and its fix | Demos and pilot-only references; no failure to describe |
| Evaluation and monitoring (5, 11) | 15 | A real evaluation report with a blocked release; live quality monitoring; a written model-change procedure | Manual spot checks; no release threshold; no model-change plan |
| Data rights and security (7 to 9) | 20 | Every provider named, with no-training terms; access on your identity provider, logged and revocable; certifications in scope | Unnamed providers; shared credentials; your data in the vendor's environment; certificate out of scope |
| Team, candor and run cost (10, 12, 14) | 10 | A named team that stays; clear reasons to decline the work; a run-cost model per task | No names; no reason ever to decline; no view of run cost |
| Regulatory roles and exit (13, 15) | 10 | Provider role and deliverables in writing; transition period, handover package and certified deletion in the contract | Regulation left to the buyer; exit left to a termination clause |
Adjust the weights to the system. A bank, an insurer or any buyer whose system may be high-risk under the AI Act should shift weight toward data, security and exit; a buyer of an agent that takes actions, toward evaluation. Score twice: on the written answers, then after an evidence session on a live system, the reference calls and a small paid discovery piece (your evaluation set, a data review, the architecture) delivered into your repository. A score that falls between rounds measures the distance between a firm's sales process and its delivery.
Verify references with production evidence
A reference is worth what it can confirm about a live system, so ask to see one running and to speak with its operator. In June 2025 Gartner estimated that only about 130 of the thousands of vendors selling agentic AI are real, and warned of "agent washing": AI assistants, robotic process automation and chatbots rebranded without substantial agentic capability.1
A demo
Built for the meeting
- The happy path, on inputs the vendor chose
- One good answer from a model
- An interface the team can build
- Nothing about failure, cost or change over time
Production evidence
A system with users and a history
- Pass rates on graded real cases, by category
- Traces, alerts and an incident caught and fixed
- Running cost per completed task at real volume
- An outside operator who confirms all of it
On a reference call, ask the operator five things:
- What does the system do each week, at what volume, and how long has it been live?
- What went wrong after launch, and how quickly did the vendor find and fix it?
- Did the people named in the proposal stay on the work?
- Do you hold the code, prompts and evaluation set in your own environment today?
- Would you hire them again for a system that touches customers or money?
Where certifications matter and where they don't
Certifications matter in proportion to how much of your data, access and production the vendor will hold. A firm building inside your environment, on your identity provider, works mostly within your controls. A firm that hosts your data, runs the system as a managed service or keeps standing production access joins your control environment, and its reports belong in your third-party risk file.
| Assurance | What it shows | What it does not show | Weight it when |
|---|---|---|---|
| SOC 2 Type 2 | A service auditor's opinion on the design and operating effectiveness of controls over security, availability, processing integrity, confidentiality or privacy.8 | The quality of the system built for you, or controls outside the report's scope | The vendor hosts your data, runs the system or holds standing production access |
| ISO/IEC 27001:2022 | A certified information security management system, for the scope on the certificate.9 | Controls outside that scope, or the AI system's behavior | As for SOC 2, once the scope includes the delivery team |
| ISO/IEC 42001:2023 | A certified management system for providing or using AI, from risk assessment to risk treatment.10 | That any given system is accurate, safe or compliant with the AI Act | The vendor will run AI systems for you over years |
Ask for the report itself and read the scope, the period covered and the exceptions found. ISO/IEC 42001 is the first international AI management system standard, and it does not replace laws or regulations.10 For most builds it is a useful signal and a weak filter.
The exit test: could another team run this next month?
The exit test asks whether a capable team that has never met the vendor could operate, fix and change the system next month using only what the contract says you receive. Write this list into the contract as the handover package, and check each item during the build.
What another team needs to take over
- The repository, with full history, in your account
- Infrastructure as code for every environment, in your cloud account or your data center
- Prompts and model configuration, versioned with the code
- The evaluation set, its graded outputs and the code that runs it
- Fine-tuned weights, the vector index and scripts to rebuild both from your data
- Runbooks, dashboards, alerts and incident history
- An inventory of open-source components and their licenses
- Credentials you own, with vendor access removable in one step
Regulated buyers already plan for this. Since January 17, 2025, EU financial entities under DORA must have exit strategies for ICT services supporting critical or important functions, with a mandatory transition period in which the provider keeps delivering while the entity moves to another provider or in-house.11 The same terms suit any AI system your business depends on.
These are the standards we hold our own AI development work to: your repository and environment from the first commit, evaluation sets your experts shape and own, and a handover another team could run.
Questions leaders ask
What is AI vendor due diligence?
AI vendor due diligence is a structured check of a vendor before contract: what you will own, what it can prove in production, how it handles your data, how it delivers and how you can leave. Most published questionnaires assume you are buying an AI product as a service. Commissioning a build adds ownership of code, prompts, evaluation sets and model artifacts, and the EU AI Act provider role, which usually falls to you.5
How do you avoid vendor lock-in with an AI development company?
Keep everything in your environment from the first commit, name every AI asset in the ownership clause, decline proprietary runtime components you would keep licensing, and put the model behind an interface so it can be swapped after an evaluation run. Then write exit terms with a transition period and a defined handover package, and rehearse the handover midway through the build.
Who owns the IP in code written with AI tools?
Ownership follows the contract, to the extent rights exist. The US Copyright Office concluded in January 2025 that copyright does not extend to purely AI-generated material and that prompts alone do not give enough human control; human-authored work and creative modifications remain protected.3 Ask the vendor to assign every deliverable however it was produced, keep it confidential, and disclose which AI tools its engineers use on your code.
Does an AI development company need ISO 42001 certification?
Rarely as a condition of a build. ISO/IEC 42001 certifies an organization's AI management system; it does not show that the system built for you is accurate or compliant, and it does not replace laws or regulations.10 Weight it where the vendor will operate AI systems for you over years. For access to your data, a SOC 2 Type 2 report or ISO/IEC 27001 certificate scoped to the delivery team matters more.89
Is the vendor or the customer the provider under the EU AI Act?
Usually the customer. The Act's provider definition covers an organization that has an AI system developed and puts it into service under its own name, so commissioning a build leaves the role with you.5 The builder supplies services. For high-risk systems, Article 25 requires a written agreement setting out the information, technical access and assistance the provider needs from its suppliers.56
What are the red flags when hiring an AI development company?
Handover promised only at the end, joint IP or a license back to your prompts and evaluation set, a proprietary platform you must keep licensing, demos offered in place of production references, unnamed model providers, shared credentials, and no answer to how the team will detect a drop in quality. Any one of these on ownership or data is reason enough to stop.
Sources
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027Gartner, June 25, 2025
- Vulnerability exploitation top breach entry point, 2026 industry-wide DBIR findsVerizon, 2026 Data Breach Investigations Report, May 19, 2026
- Copyright and Artificial Intelligence, Part 2: CopyrightabilityU.S. Copyright Office, January 2025
- Regulation (EU) 2016/679, the General Data Protection RegulationOfficial Journal of the European Union
- Regulation (EU) 2024/1689, the Artificial Intelligence ActOfficial Journal of the European Union
- Regulation (EU) 2026/1744, the Digital Omnibus on AIOfficial Journal of the European Union, July 24, 2026
- Updated EU AI model contractual clausesEuropean Commission, Public Buyers Community, March 5, 2025
- SOC 2® Reporting on an Examination of Controls at a Service Organization Relevant to Security, Availability, Processing Integrity, Confidentiality, or PrivacyAICPA & CIMA
- ISO/IEC 27001:2022, Information security management systemsInternational Organization for Standardization
- ISO/IEC 42001:2023, AI management systemsInternational Organization for Standardization
- Regulation (EU) 2022/2554, the Digital Operational Resilience ActOfficial Journal of the European Union
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.