Question skeletons
Reuse a proven query only on a structural match. Mask every word you can ground; what's left must be equal. Similarity ranks examples but never decides.
Caching “which query answered which question” saves a model call and gives a known-good answer. Matching by embedding similarity is the obvious approach, and it’s wrong too often. “Withhold” and “administer” score 0.96. A semantic cache at a 0.9 threshold returns the wrong answer on roughly a third of hits.
So compare questions by skeleton: mask every word that links to something in the data, and keep what’s left.
"open orders" → "* *"
"how many open orders" → "how many * *"
"open orders from Canada" → "* * from canada"
"customers without an order" → "* without an *"
What survives is what the data can’t say: “how many”, “top 10”, “per”, “without”, a name nothing stores. Two questions with equal skeletons ask for the same shape of answer.
Matching, in order
question ──▶ ground words ──▶ skeleton
│
1. known unanswerable twice? ───┼──▶ "not supported", no model
2. exact skeleton + covers ─────┼──▶ run the proven query, no model
3. same, one value swapped ─────┼──▶ substitute the value, run it
4. shares labels ───────────────┼──▶ 1–3 examples into the prompt
5. nothing ─────────────────────┴──▶ generate as usual
Coverage, both ways
A stored query answers a question only if every label, filter and value the query uses is named by the question, and everything precise the question names is used by the query. So a broader stored query never answers a narrower question, and a narrower one never answers a broader one.
Rules
- Similarity never decides. It can rank examples for the prompt. Reuse needs a structural match.
- Store skeletons, not what people typed.
- Trust decays. A proven answer decays like any fact, and it’s out the moment a fact it rests on breaks.