Value of information
You can't re-check every fact all the time. Score each one by how much a new look would be worth, divide by what it costs, and spend the budget from the top.
Once facts decay, something has to keep them fresh. Re-checking everything on a timer is the obvious answer, and it doesn’t hold up. A big account has hundreds of thousands of possible facts, and checks run against a live database that customers are also using.
So every candidate gets a score for how much one more observation would be worth right now, divided by what it would cost. Each run spends a fixed budget from the top of that list.
The score
const score = (fact) =>
importance(fact) *
( 4 * Math.log1p(fact.demand) // someone asked for it
+ Math.log1p(fact.staleDays) // it's been a while
+ 12 * fact.belief.variance // we're unsure about it
+ fact.pruneValue // a "no" here rules out many others
+ (fact.neverSeen ? 0.25 : 0) ) // we've never looked
/ fact.relativeCost; // how slow this check usually is
| Term | Plain meaning |
|---|---|
| importance | how much this fact matters to answers. It scales the score but never zeroes it |
| demand | a question needed it and didn’t have it. Weighted highest, because a real person is waiting |
| staleness | log1p grows forever, so even the least important fact eventually comes up |
| variance | how unsure the belief is. A coin flip (p = 0.5) is worth checking, p = 0.99 isn’t |
| prune | if a broad fact is probably false, refuting it also refutes every narrower fact under it |
| never seen | a small nudge toward discovering things |
| cost | measured from past checks. Cheap checks win when the value is similar |
How a run spends its budget
┌──────────────────────────────┐
demand ──────▶ │ 1. demanded facts │ someone is waiting: always first
├──────────────────────────────┤
readiness ───▶ │ 2. facts a feature needs │ so the feature can turn on
├──────────────────────────────┤
everything ──▶ │ 3. rest, by score │ value ÷ cost, highest first
└──────────────┬───────────────┘
│ stop at budget, deadline, or 3 failures in a row
▼
observe ──▶ stamp ──▶ belief changes ──▶ scores change
The loop feeds itself. A check updates the fact’s stamp, that changes its belief and staleness, and that changes its score for the next run. You don’t keep a separate schedule.
Polls are just cheap probes
The same idea covers data you fetch rather than probe, like plan features, integration health, or usage rollups. For those, value is mostly staleness:
const pollValue = (fact, now) => importance(fact) * (1 - freshness(fact.kind, fact.last, now));
Each kind’s half-life decides how often it’s worth refreshing. Plan features come up hourly. A user’s goals never do, because the user said them. That replaces “refresh everything every hour” with one rule that every source follows.
Why it works
- Real questions jump the line. A miss during a live answer becomes demand, and demand is checked first.
- Nothing starves. Staleness grows without bound, so every fact gets looked at eventually.
- The database stays healthy. Cost is learned from past runs, and a breaker stops the run after repeated failures.
- Effort goes to doubt. Facts you’re already sure about cost almost nothing to keep. The budget goes to what’s uncertain, in demand, or has never been seen.