How to Choose an AI Development Company: A Buyer's Checklist
Nearly everything you evaluate is easy to manufacture
Case studies, client logos, team biographies, framework lists, a tidy portfolio page. All of it can be assembled by a firm that has never put a model into production or had to explain why one stopped working.
The signals that cannot be manufactured require having failed at something. A vendor who has watched a project stall knows where projects stall, and that shows up as specific, uncomfortable answers. A vendor who has only closed deals produces smooth answers with no edges.
Five questions separate the two, and none are about model architecture.
What would make you decline this work?
Ask it directly in the first conversation.
A vendor with a real answer names conditions: not enough labelled outcomes, no reliable way to measure whether it worked, no capacity to act on predictions, an unsettled regulatory position, or data that cannot be reconstructed as it stood when decisions were made. Each of those answers comes from having seen it sink something.
A vendor who cannot name one has no standard, and will not apply one to your project either. You find that out at the point where someone should have said stop.
A related test costs nothing: include a requirement that is slightly wrong and see whether they push back. Agreement with everything is not flexibility, it is an absence of opinion, and opinions are most of what you are paying for.
What happens in the first two weeks?
Good answers start with data: whether the outcome variable exists and is recorded reliably, how much history is available, whether features can be reconstructed as they stood when a decision was made rather than as they look now, and whether the predicted events are captured anywhere at all.
Weaker answers start with architecture, model families or platform decisions. Those matter later and none are knowable before anyone has examined your data. A proposal naming the model before the data assessment is describing a preference.
What result would count as this not having worked?
This is the question most proposals cannot survive, and the one to be least flexible about.
Before any build begins, the proposal should state what success means, what the baseline is, and how the comparison will be made. Baselines carry the weight. An accuracy figure with nothing to compare it against is not evidence, and on rare outcome problems a model predicting that nothing happens will post an excellent number.
Press for specifics. Current performance of whatever is being replaced, measured the same way. Whether there will be a holdout, a control group, or a period running alongside existing practice without influencing decisions. How degradation gets detected, and at what threshold someone acts.
A vendor who cannot describe a disappointing outcome has not designed an evaluation, and a project that cannot fail cannot be shown to have succeeded.
Who acts on the output, and how many can they handle?
A prediction does nothing by itself. Someone has to behave differently because of it.
Establish who, how many cases they can work in a day, how fast they can act, and what they will do. That capacity sets the operating threshold, which determines what performance actually needs to be. A model tuned to a threshold nobody can staff is tuned to nothing.
A vendor who raises this in the first meeting is thinking about outcomes. One who defers it to implementation plans to hand over a model and leave.
What does year two cost?
Models degrade. Data drifts, processes change, regulations move, and a retrained model is a different model needing revalidation.
Ask who retrains and on what trigger, whether monitoring is included and who watches the alerts, what a regulatory change costs when it forces a modification, and what three years total looks like including your internal effort. A build price with no ongoing view is either incomplete or assumes you will return without leverage.
What belongs in the contract rather than the conversation
Technical proposals get scrutinised. Terms get skimmed, and they are easier to negotiate before work starts than after.
- Ownership of the model, code, feature definitions, pipelines and documentation. A model file alone cannot be retrained or revalidated, so owning the artefact without the pipeline is not ownership of anything useful.
- Whether your data, or anything derived from it, may improve the vendor's other work. Ask explicitly and put the answer in writing.
- Sub processors. If the vendor sends data to a hosted model API, that provider sits in your data flow, needs approval, and belongs in your privacy notices. It is regularly discovered afterwards, which is how a delivered project becomes a compliance problem.
- Data handling during development. Where it lives, who can access it, whether access is logged, whether the work can run on masked or sampled data, and a dated deletion commitment.
- Exit. Could another team operate the system from documentation and tests alone.
The principle worth requiring of any supplier, including us, is that you end up able to operate, audit and change what was built without the original vendor in the room. Where the system handles sensitive data or makes consequential decisions, independent security review belongs inside the engagement rather than as an optional line item, covering the surrounding services and not only the model.
Sequence it so you can stop
The strongest protection is not a clause, it is an order of operations.
Scope and pay for the data assessment as a separate first phase with a genuine option to stop. It costs a fraction of a full build, answers the questions that determine viability, and lets you watch how the vendor works before committing to the rest. If it concludes the project should not proceed, that is a good result cheaply obtained.
Fixed pricing a full build on data nobody has examined transfers risk to whoever is least able to price it, which is usually you.
When the answer is not a vendor at all
An honest checklist has to include this, because for many use cases custom development is the wrong purchase.
A mature domain product may already encode years of refinement your build would start without. An internal hire can be better value for ongoing work, since they accumulate context a vendor has to be paid to rebuild each engagement, though a team of one is fragile. A foundation model API with light engineering covers classification, extraction and summarisation cheaply, trading off cost at volume, latency, non determinism and your data travelling to a third party. And often the constraint is a process gap or no capacity to act, in which case fixing that removes the problem.
The same test applies to our field. We build blockchain systems, and for most prediction problems, fraud scoring, churn, clinical risk, credit decisioning, a distributed ledger contributes nothing. If a vendor proposes bundling one into an analytics project, ask what it solves. If the answer is portability, auditability or trust, ask why a database with signed logs and proper access control does not cover it. Usually it does.
Apply all of this to us as readily as to anyone else. A supplier who objects to being evaluated this way has answered a different question.
RWaltz is a blockchain and enterprise software development company building custom smart contracts, dApps and tokenization platforms that integrate with existing business systems. We work to a build to own model: clients hold their keys, repositories and intellectual property, engagements are scoped honestly including the cases where a simpler approach or a different supplier is the better answer, and review is treated as continuous rather than a single sign off.
📖 Read the full blog: https://www.rwaltz.com/blogs/how-to-choose-an-ai-development-company-a-buyers-checklist
Connect with RWaltz:
- LinkedIn: https://www.linkedin.com/company/rwaltzsoftware
- X (Twitter): https://twitter.com/rwaltzsoftware
- Facebook: https://www.facebook.com/RWaltz-Software-PvtLtd-255590135349493
- Telegram: https://t.me/RWaltzCrypto
- GitHub: https://github.com/rwaltzsoftware
- Clutch: https://clutch.co/profile/rwaltz-software
- Website: https://www.rwaltz.com
Comments
Post a Comment