Caching “which query answered which question” saves a model call and gives a known-good answer. Matching by embedding similarity is the obvious approach, and it’s wrong too often. “Withhold” and “administer” score 0.96. A semantic cache at a 0.9 threshold returns the wrong answer on roughly a third of hits.

So compare questions by skeleton: mask every word that links to something in the data, and keep what’s left.

"open orders"                       →  "* *"
"how many open orders"              →  "how many * *"
"open orders from Canada"           →  "* * from canada"
"customers without an order"        →  "* without an *"

What survives is what the data can’t say: “how many”, “top 10”, “per”, “without”, a name nothing stores. Two questions with equal skeletons ask for the same shape of answer.

Matching, in order

  question ──▶ ground words ──▶ skeleton
                                   │
   1. known unanswerable twice? ───┼──▶ "not supported", no model
   2. exact skeleton + covers ─────┼──▶ run the proven query, no model
   3. same, one value swapped ─────┼──▶ substitute the value, run it
   4. shares labels ───────────────┼──▶ 1–3 examples into the prompt
   5. nothing ─────────────────────┴──▶ generate as usual

Coverage, both ways

A stored query answers a question only if every label, filter and value the query uses is named by the question, and everything precise the question names is used by the query. So a broader stored query never answers a narrower question, and a narrower one never answers a broader one.

Rules

  • Similarity never decides. It can rank examples for the prompt. Reuse needs a structural match.
  • Store skeletons, not what people typed.
  • Trust decays. A proven answer decays like any fact, and it’s out the moment a fact it rests on breaks.