AI Answer Library
Short answer
Focus on three things: what they have actually shipped, who owns the data, and whether you can maintain the system after they leave. Ask for a real, reachable system rather than slides and demo videos. Put ownership of data, model weights, source code and fine-tuning artifacts in the contract. Get pricing for phase-two changes and incident response before you sign. Which model or whether to use RAG versus fine-tuning is the easiest thing to change later and should not dominate the evaluation.
Turning the evaluation into verifiable questions beats counting certifications and logos. Every row below can be asked in the first technical conversation, and the quality of the answer is itself information — vagueness, deflection and repeated insistence that "we are technically strong" are all signals.
| Dimension | The question to ask | Red flag | What a good answer looks like |
|---|---|---|---|
| Delivery evidence | Can you give me an account on a live system so I can click around myself? | Only slides, edited videos and redacted screenshots | A reachable URL, or a live demo of unscripted operations |
| Data and asset ownership | In the contract, who owns the data, the fine-tuned weights and the source code? | "That is industry standard" without agreeing to write it down | Explicit clauses plus a stated handover and export procedure |
| Maintainability | After you leave, how does my engineer change a prompt or add a document type? | Every change must route back to the vendor and is billed by the day | An admin console, documentation and a real handover session |
| Acceptance criteria | Which real questions form the acceptance set, and what threshold counts as passing? | Unmeasurable wording such as "industry-leading performance" | A jointly built evaluation set, with threshold and judge named in the sign-off |
| Downstream cost | What are year-two support fees, phase-two day rates and the incident response window? | Deferred before signing with "we can discuss that later" | A written quote attached to the main contract |
| Business understanding | Please restate how our current process actually works | Talks only about models and parameters, cannot describe your process | Names the specific bottleneck in your flow and states the trade-offs |
Many RFPs spend most of their length on which models must be supported and which architecture must be used. That is the wrong center of gravity. Model families ship new versions every few months, and the RAG-versus-fine-tuning balance shifts as data accumulates; both are adjustable mid-project. What is hard to recover from is different: unclear data ownership, a delivered system nobody in-house can modify, and acceptance criteria vague enough to argue about. Weighting those three protects your investment far better than freezing the technical design. The exception is a genuine compliance constraint — data must stay on-premise, or a specific sovereignty requirement — which is a precondition and should be fixed in the requirements rather than left to the vendor.
The most effective risk reduction is not a thicker contract but a smaller first commitment. Pick a pilot with clear boundaries, real users, and no impact on the core business if it fails, then fix a time window and an acceptance set. During the pilot you learn more than any due diligence would tell you: the quality of the questions they ask, whether they explain or deflect when something breaks, how good their documentation is, and whether their delivery cadence is steady. Expand after the pilot passes; if it fails, the loss is bounded. It is also the cheapest test of real capability — teams willing to take a small pilot and put acceptance criteria on paper are usually more reliable than those who insist on a large contract up front.
Where this applies
People also ask
What is the difference between an AI agent and RPA, and which should a company choose?
Should a company buy software outright or subscribe to SaaS?
Self-hosted LLM or public API — how do I choose?
DataWeaver
Five specialized agents cut cross-database natural-language queries from 2–3 days to 10 seconds.
Meeting Management — Private-Deployed Platform
A ¥19,800 perpetually licensed, self-hosted meeting platform — 8 modules, any BCP 47 language, 4-channel international payment, AI site builder + Copilot, dual-QR check-in. v0.6.0 in production.