Choosing an AI consultant is harder than choosing an AI model. Models are measurable: benchmarks, prices, response times. The quality of a consultant's work is invisible on the day you sign. We have 46+ projects behind us, and we regularly get calls from companies where a previous implementation didn't work out.
Seven criteria for the choice — for each: what to ask, what a good answer sounds like, the red flag — and a method for making proposals comparable.
We're an AI consultancy ourselves, so on this topic we're biased. That's why these are criteria and questions you can use on any candidate — including us.
References and measurable results
A logo on the candidate's website says little on its own. A reference is a person you can call who will tell you what went wrong during the project.
- What to ask. "Who can I call from a project that was a similar task to mine?" Similarity of the task matters more than similarity of the industry: invoice processing is the same problem at a logistics company and at a manufacturer.
- What a good answer sounds like. At least one name with contact details, from the same type of process. The numbers come with context: you learn what the accuracy was measured on, and what happens to the documents the system can't process. The candidate can also tell the story of a project that went off track: what went wrong and how the team handled it — something goes wrong in every project, and the handling tells you the most.
- The red flag. Company names without people you can call. An accuracy number without context. A demo that runs on a perfectly formatted dataset — ask for the presentation to process your own sample documents.
The candidate's company data is public: in the Hungarian company registry you can check how long the company has existed and what it actually does.
Who works on the project, and who operates it?
The team that writes the proposal and the team that actually works on the project can be two different things. And the system needs operating long after the project ends.
- What to ask. "Who actually works on the project, and who operates the system after go-live?" Ask to meet the team that will actually work on your project — before signing.
- What a good answer sounds like. A team introduced by name and experience, and a maintenance plan with a defined SLA and annual cost from day one. AI systems need ongoing care: models drift, APIs get deprecated, source-system upgrades break integrations.
- The red flag. Senior engineers in the sales meetings, different people on the project. "Maintenance is minimal." No named operator.
The question is the same regardless of size. At a large firm: who actually works on the project, and what is their experience with similar implementations? At a small team: what happens if the lead developer drops out? A full team carries our work — project managers, senior developers, an AI DevOps engineer — but ask us the second question too.
Custom build, off-the-shelf product, or a simpler tool
A good recommendation starts from the task and decides per use case between a custom build, an off-the-shelf product, or a simpler tool.
- What to ask. "When do you recommend an off-the-shelf product instead of a custom build? And when do you say that RPA or a script is enough?"
- What a good answer sounds like. A recommendation per use case, with the criteria written down: data protection, integration, operating cost. The solution is built on standard components, the model can be swapped without rebuilding the whole system, and any proprietary platform gets a justification.
- The red flag. Every task gets the same platform as the answer. The candidate is tied to a single platform, and every recommendation leads there, regardless of your task.
EU AI Act readiness
The EU AI Act classifies AI systems by risk, and the obligations follow from the category. The classification belongs at the start of the project, because the risk category determines the documentation and oversight requirements.
- What to ask. "Who runs the risk classification, and when?"
- What a good answer sounds like. The classification happens at the start of the project, per system, with written reasoning. The regulation defines four categories — unacceptable, high, limited, minimal — and the candidate also clarifies whether your role is provider or deployer. The legal assessment belongs to your own counsel; a good consultant hands over the documentation in a form your counsel can review.
- The red flag. The candidate can't speak to the classification, or promises it for the end of the project. In that case, compliance stays with you.
Pricing transparency
Most of the budget goes into data preparation and integration; the model is the cheap part. The structure of the proposal either shows this or hides it.
- What to ask. "What does the price include, item by item, and what is the system's annual operating cost?"
- What a good answer sounds like. An itemized breakdown: data preparation, integration, training, and operations on a separate line — API fees, monitoring, retraining. A fixed-price offer for the first step, with a result that stands on its own.
- The red flag. A single lump sum without a breakdown. The proposal prices only the model development, and data preparation and integration arrive later as change requests. The operating cost appears nowhere.
If you have several AI ideas on the table and need to decide which pays back and in what order, the strategy assessment is where we start. AI Strategy Consulting
Ownership and vendor lock-in
You can only switch vendors if there is something you can take with you.
- What to ask. "Who owns the source code, the trained models, and the data pipelines? What happens when the contract ends?" Ask where the models run and what data leaves your environment, too.
- What a good answer sounds like. The contract states that all of it becomes your property. The system keeps running after the contract ends: exportable data, documented deployment, and handover documentation another team can take the system over from. For sensitive data: on-premise or private cloud deployment, with open-source models running on your own servers.
- The red flag. The system runs on the vendor's proprietary platform and stores your data in their format — switching vendors means starting over from scratch. The documentation is promised "at the end".
Honesty
The best filter is a question a bad candidate can't answer well.
- What to ask. "When did you last tell a client that their task doesn't need AI?"
- What a good answer sounds like. A concrete story. Sometimes the answer is a simple script, a better database query, or fixing the quality of the source data. We have talked clients out of AI projects and built them an RPA solution instead — it was cheaper, it was ready sooner, and it worked better.
- The red flag. Every idea of yours gets an enthusiastic yes. A candidate who takes on everything learns what doesn't work on your budget.
There is a separate article on the vendor-side red flags: What Your AI Vendor Won't Tell You Before Signing.
How to make proposals comparable
Collect three proposals and compare them on the same points. Four steps:
- The same use case. Pick one process — incoming invoice processing, for example — and describe it to every candidate at the same level of detail: how many documents, from what sources, which systems it must integrate with, what counts as success. Proposals for generic "AI adoption" can't be compared against each other.
- The same monthly volumes. Every proposal computes with your monthly volumes. If a candidate computes with their own averages, the prices are not comparable.
- An itemized quote. Data preparation, integration, training, an accuracy commitment measured on your data, and the annual cost of operations on a separate line.
- The payback claim together with its assumptions. Record the baseline before the start: what the process costs today in time and money. The assumptions behind the payback calculation go into the proposal in writing — that keeps the claim checkable after delivery.
For calibration: an AI assistant or integration project usually runs €15-40K; a document processing system €30-80K. If a proposal sits well below that band, ask what was left out. If it sits well above, ask what it includes on top.
There is a reference point for the timeline too: a proof of concept takes 2-3 weeks, and an implementation project takes 6-12 weeks from kickoff to production. If the promised timeline is much shorter, ask what is missing from it.
And for the payback: across our 46+ completed projects, automated processes reach a 40-45% efficiency gain, and the typical payback period is 6-12 months. If a proposal promises substantially more, ask what it's based on. You can run your own numbers in the AI ROI calculator.
What should the first engagement look like?
Small, and usable on its own. Size the first piece of work so that its result lets you decide whether to continue.
We use an AI opportunity assessment for this — from HUF 500,000 + VAT. Its result is a written plan: which process can be automated, at what cost, and how fast it pays back. The format works with any consultant: ask for a first deliverable you can work with even if you don't continue with them.
A second attempt after a failed implementation always costs more than a good question asked the first time. The criteria are above. Use them — on us too.

