The first question in most AI conversations is which model to use. It is the wrong first question, and answering it early is how projects end up with an impressive demo and no way to tell whether it is getting better or worse.
The artefact that decides the outcome is duller: a written set of real questions from the people who will use the system, each paired with the answer it should have produced and the source that answer comes from. A few hundred is usually enough. Collecting them is unglamorous and takes about a week of other people's time.
Once it exists, everything downstream becomes a measurement instead of an argument. Chunking strategy, retrieval depth, reranking, prompt structure, model choice: each one becomes a number that went up or down. Without it, every change is defended by whoever demoed it most recently.
It also changes the conversation with the business. "Accuracy on our evaluation set moved from 71% to 88%" is a sentence a sponsor can act on. "It feels much better now" is not, and it is the reason so many pilots never get funded into production.
The uncomfortable part is that a good evaluation set will sometimes tell you the system is not ready. That is the point. Finding out in week three costs a week; finding out after rollout costs the credibility of the whole programme.
So before the model comparison, before the vector database procurement, write down what a correct answer looks like. Everything after that is engineering.
- evaluation
- RAG
- delivery
