How to evaluate an AI tool in an afternoon
I once picked the wrong one of two tools because it wrote in bullet points.
Same inputs, same session, run back to back. One answered in tidy headings with short certain sentences. The other returned a wall of text that read like somebody thinking out loud with the door open. I preferred the first one and I was not remotely neutral about it.
Then I checked both against the correct answers, which I had written down before either tool ran. The tidy one got roughly a fifth of the extraction fields wrong. The mumbling one got two wrong out of the same set.
Formatting reads as competence. It costs a model nothing to produce, and it had me for most of an afternoon.
What follows is how I test something before paying for it. It takes about four hours and most of those hours happen before the tool is open. It is also written so you can stop reading reviews, including the ones on this site. My awkward inputs are not your awkward inputs.
A demo is a rehearsal with a chosen input
Somebody at the vendor picked the file in that video. They picked it because it works. I would do the same and so would you, and nothing dishonest has happened.
The consequence is that a demo can only ever tell you the tool has a happy path. Every tool has a happy path. The thing you are buying is the behaviour at the edges: the malformed record, the empty upload, the file four times longer than anything in the documentation, the question phrased by somebody who does not know the right words for it.
Free trials have the same problem in a friendlier costume. You log in, you try whatever comes to mind, and what comes to mind on the spot is a clean example. I have done that a dozen times and walked away with a warm feeling and no evidence.
Twenty files, pulled before you sign up
Step one is the whole method, honestly. Before you create an account, go into your own work and pull ten to twenty real inputs into a folder.
Real means the actual artifacts. The crooked scan. The record with a blank field where the postcode should be. The support message containing three separate questions plus a screenshot. The contract in a language you did not plan for. If you would be slightly embarrassed to show it to the vendor, it belongs in the folder.
The ordering does the real work and it is the step people wave off. Collect your test cases after watching the tool succeed and you will pick them badly. Nobody does it on purpose. You see it handle a long clean document, that quietly sets the shape of what occurs to you next, and forty minutes later your whole set is long clean documents.
I caught myself doing exactly that and it took an hour to notice I had not tried a single short input.
Twenty is comfortable. Ten works. Under five and you have rebuilt the demo using your own pictures.
Keep the folder afterwards. Something cheaper always turns up, somebody always asks whether it is worth switching, and having the same set ready turns that argument into an afternoon instead of a debate. Mine has outlived three of the tools it was assembled for.
Write the answer key first
For every input in that folder, write down what a correct output would be. Do this before you see any output.
Extraction makes it easy. Six fields, the right value in each, done. Summarisation makes it harder and it still has to be done: which two facts must appear, what must never appear, the length past which it has failed regardless of quality.
Twenty minutes on paper. It feels like schoolwork. It is the entire difference between a measurement and a mood, and skipping it is how you end up preferring bullet points.
Break it on purpose
Now the checks. This first one is worth more than everything after it.
Hand the tool something it cannot possibly do. A blank file. A page in a script it does not support. A question whose answer is nowhere in the material you supplied. A form with the important section missing.
Two things can happen. It says it cannot, or it produces something plausible anyway.
A tool that fails loudly is a tool you can design around. You catch the failure, you route it to a person, work continues. A tool that invents is inserting quiet errors into your output at a rate nothing will ever surface for you. One error in fifty is fine when you can tell which one. One in fifty with no signal means checking all fifty by hand, which is the job you were trying to stop doing.
Given the choice I take the tool that refuses twice as often, every time.
Same input, twice
Thirty seconds. Run one of your inputs again and compare.
Variation is desirable for drafting and generating options. It is a year of intermittent bug reports for classification, extraction, or anything a downstream system parses.
Check the shape as well as the content. I have had a tool hand back a price as a bare number on the first call and as a string with a currency symbol glued to it on the second. That kind of thing breaks code at eleven at night, in a job nobody is watching.
Your size, not the tutorial’s size
Every worked example in every set of documentation is small. Twelve rows, one page, a paragraph. Your real thing is four hundred rows or a sixty page pdf or eighteen months of chat history.
Behaviour at scale changes suddenly rather than gradually. Quality drops off a cliff past some length nobody documents. The per unit price that was trivial at ten reads very differently at ten thousand. The call that took two seconds takes ninety, or times out.
I watched one tool go from useful to useless somewhere between a five page document and a forty page one. There was no warning, and there was no reason for there to be.
Start with the largest real input you own. The small ones will follow.
Try the export in week one
Put something in, then get it back out. The genuine export, on the plan you would actually buy, into a format another program can open.
Ten minutes, and it has killed two purchases for me. One had export sitting behind a tier at four times the price I wanted to pay. The other produced a file that was technically a file, with every field flattened into a single column.
You learn this in month one for free, or in month fourteen when you are leaving and leaving is the urgent thing.
Price it at ten times the volume
Run the arithmetic at ten times what you do now. You may never get there. Do it anyway, because pricing is designed around an assumed volume and gets strange outside it. Per seat is fine until the eleventh person. Per unit is fine until you automate the step upstream and the units multiply overnight.
Then read what happens to your data. Retention period, whether it is used for training, and whether that answer differs on the plan you are about to buy. It frequently differs. If customer records are going through the thing, read this before the feature list.
Find the rate limit while you are in there. Plenty of tools that feel instant on one request quietly queue everything behind the fifth, and the number tends to live on a page the marketing site does not link to.
Last commercial question: what do you have if the company is gone in eighteen months? A fair number will be. Six months of accumulated configuration with no way out is a bet on somebody else’s funding round, and you should at least know you are placing it.
The one I bought off a good demo
Transcription and meeting notes. Clear audio, single speaker, a subject the model had obviously seen plenty of. Clean output in ninety seconds. I bought a year because a year was cheaper than monthly, which is a separate lesson I also paid for.
My actual recordings are three people in a room with the air conditioning running, Singapore accents, product names that are not words in any language, and somebody talking over somebody else every few minutes.
What came back was beautifully formatted summaries with the wrong names attached to the wrong lines. Everything looked correct. The attribution was scrambled, and errors of that shape survive a skim perfectly, which is what made them expensive.
Four months of correcting the output before I admitted the correcting took longer than writing notes by hand ever had. Twenty minutes with one real recording would have caught it. I skipped the twenty minutes because the demo had already done its work on me.
When it comes out close, it does not matter
Most of the time the tool is not the constraint.
The constraint is that your inputs are inconsistent, or nobody decided what the output is for, or the process it slots into has a busy person standing in the middle of it. No purchase touches any of that.
So when an evaluation ends with two options within touching distance, that closeness is the result. It means the decision does not matter. Take the cheaper one, or the one you could walk away from with less pain, and spend the week you would have burned on the choice fixing the inputs instead.
The operators I know who get real value out of these tools mostly picked something adequate in an afternoon and then did the boring work underneath it. If you run this method against something reviewed here and reach the opposite verdict, yours is the one that counts. Current per token prices and where each tool broke for me are here.