← all articles

Building an eval set from support tickets

Why your ticket queue is a better eval source than a public benchmark

Every RAG demo I’ve shipped starts the same way: it looks great in a five-question smoke test, then someone in support forwards a screenshot of it confidently making something up. Public benchmarks don’t catch this because they don’t know your product. They don’t know that “downgrade” means something specific in your billing system, or that your users type “it’s not working” and mean four different things depending on which page they were on.

Support tickets already contain the exact questions your users ask, in their own words, with a human agent’s answer sitting right next to them as a rough ground truth. That’s the raw material for an eval set that actually tells you whether your model or pipeline is getting better or worse on your problem, not on someone else’s leaderboard.

The catch is that a ticket queue is not an eval set. It’s a pile of unstructured, duplicate-heavy, occasionally wrong, occasionally PII-laden text. Turning it into something you can run against a model week over week takes real filtering work, and skipping that work is how teams end up with an eval that just measures how well the model memorizes ticket #4021.

Start by pulling a sample, not the whole backlog

Resist exporting six months of tickets on day one. Pull a stratified sample instead: a few hundred tickets spread across your top ticket categories (billing, login, integration errors, “how do I,” bug reports), weighted roughly by how often each category actually shows up in the queue. If 40% of your tickets are password reset questions, your eval set should reflect that, not treat every category as equally important.

A workable starting size is 150 to 300 tickets. That’s small enough to hand-review in a few days and large enough to catch category-level regressions. You can grow it later as you find gaps, but a 2,000-ticket eval set nobody has read closely is worse than a 200-ticket one you actually trust.

Filter before you label anything

Three passes, in this order, before a human looks at content quality:

Dedupe. Support queues are full of near-duplicates, the same login bug reported by twelve users in slightly different phrasing. Cluster on embedding similarity (a simple cosine threshold around 0.92 to 0.95 on a sentence embedding model works fine for this, no need for anything fancy) and keep one or two representatives per cluster. Otherwise your eval set is secretly 80% one issue and your aggregate score is dominated by whether the model handles that one issue, not by real coverage.

Strip PII. Names, emails, account IDs, order numbers, and anything pasted from a screenshot with a customer’s data in it needs to come out before this set lives anywhere outside your ticketing system, including a shared eval repo or a prompt sent to a third-party model API. This isn’t optional and it isn’t a nice-to-have for compliance theater, it’s the difference between an internal tool and a data leak. A regex pass catches emails and obvious ID patterns; a second human pass catches the stuff regex misses, like a customer pasting their own API key into a ticket because they thought it was the problem.

Drop tickets where the agent’s answer is wrong. This sounds obvious but it’s the step teams skip because it’s slow. If you use the agent’s original resolution as your ground truth without checking it, you’re eval-ing your model against your support team’s mistakes too. Read each candidate answer. If an agent said “there’s no way to do that” and there actually was, either fix the ground truth or drop the ticket. A wrong ground truth is worse than no ground truth because it silently penalizes the model for being right.

Turn each ticket into a question, answer, and source triple

For a RAG pipeline specifically, you want three fields per eval item, not two:

  • Query: the user’s actual question, cleaned up slightly for clarity but kept close to the original phrasing. Don’t rewrite “it wont let me log in help” into a polished sentence, that defeats the point of using real tickets.
  • Reference answer: what the correct resolution actually was, written by you or verified against your documentation, not just copy-pasted from the agent’s chat log.
  • Source documents: which doc, KB article, or code path should have been retrieved to answer this correctly. This is the field most teams skip and it’s the one that actually tells you whether a bad answer is a retrieval failure or a generation failure. If your pipeline pulled the right doc and still answered wrong, that’s a prompt or model problem. If it never pulled the right doc, no amount of prompt tuning fixes it.

Store this as JSONL, one object per line, with those three keys plus a category tag from your ticket taxonomy. That format drops straight into most open source eval harnesses (promptfoo and ragas both accept roughly this shape) without extra glue code, and it’s readable enough that you can diff two runs by eye when a score moves.

Scoring: pick something you can defend, not something impressive

Exact-match scoring is useless for open-ended support answers, nobody phrases things identically. The three approaches that actually hold up:

LLM-as-judge with a rubric, not a vibe. Give the judge model the query, the reference answer, and the candidate answer, and ask it to score on specific criteria: did it get the core fact right, did it point to the correct action, did it invent anything not supported by the source docs. A single “rate this 1 to 10” prompt is noisy run to run. A rubric with three or four yes/no checks is far more stable and, importantly, lets you audit disagreements by reading the judge’s reasoning.

Retrieval hit rate, checked separately from answer quality. Did the pipeline’s top-k retrieved chunks include the source document you tagged in step three? This is cheap to compute (no LLM call needed, just a set overlap) and it’s the fastest way to tell whether a regression is a retrieval bug versus a generation bug before you spend judge-model tokens debugging the wrong layer.

Human spot-check on a fixed 15 to 20 item subset, every time you change the pipeline in a way that matters. Automated scores drift in ways you won’t notice until a human reads five transcripts and says “wait, this is confidently wrong.” Keep this subset fixed across runs so you’re comparing the same items over time.

Whatever you pick, run it against your current pipeline first to get a baseline before you touch anything. Without a baseline, a score of 71% means nothing, you can’t tell if that’s good, bad, or just what your system has always scored.

Keep the set alive, or it stops meaning anything

An eval set built once from last quarter’s tickets goes stale fast, especially if your product ships new features. New ticket categories show up, old ones disappear as bugs get fixed, and if you never refresh the set you end up optimizing against a snapshot of a product that no longer exists. A reasonable cadence is a quarterly top-up: pull another 50 to 100 tickets from the categories that have grown since the last pull, run them through the same dedupe and PII pass, and merge them in. Don’t discard the old items just because they’re old, keep them as a regression check so you catch a fix that silently breaks something that used to work.

One more thing worth naming plainly: don’t let this become a benchmark you publish or market against a competitor. It’s built from your own support data, scored by your own rubric, and it will not transfer cleanly to anyone else’s product or hold up as a claim about which vendor’s model is “better.” Its only job is to tell your team, honestly, whether this week’s change made real users’ real questions get answered more correctly than last week’s did. That’s a narrower goal than a leaderboard, but it’s the one that actually matters when the on-call engineer is staring at a ticket that says the bot lied to a customer.

If you want more of this kind of hands-on breakdown of what actually works when you’re the one running the pipeline and paying the bill, come find us at AI Tool Gazette.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →