A fact was true in March.
How much should you believe it in August?
This note describes a problem we ran into while building a travel dataset, an argument for why it is not adequately solved elsewhere, and a first attempt at solving it. It is a working note rather than a specification. We would rather publish it early and be argued with than publish it late and be agreed with.
The problem
Someone stands in front of a ramen shop in Shibuya. They note that it opens at 08:45, shuts at 14:00, and closes at weekends. That record enters a database. Months later a traveller asks an assistant what is open on a Tuesday morning, the record is retrieved, and the assistant answers with the flat confidence of something reading a fact.
Nothing in that pipeline expresses what has happened in between. The record does not know how old it is. The retrieval layer does not weigh its age. The model has no way to say “this was true when someone last looked, and that was a while ago”.
This is not a hallucination problem. The data is real, the retrieval worked, and the model reported it accurately. The failure is that accurate reporting of an ageing fact is presented identically to accurate reporting of a fresh one.
Why the existing standards do not cover it
Three bodies of work sit close to this, and it is worth being precise about what each one already does, because the gap is narrower than it first appears.
Answers where a piece of content came from and whether it has been altered. Its December 2025 specification extended manifests to unstructured text, including model outputs. It establishes origin and integrity. It does not express how much an assertion should be believed as it ages.
A Recommendation since 2013. Answers what a record was derived from, by whom, through what activity. It is a model of lineage. Lineage tells you the path a fact took, not whether the fact is still true.
DAMA-DMBOK and ISO 8000 define timeliness, accuracy, completeness and validity as properties of a dataset. They are written for data stewardship inside an organisation, not as a contract a retrieval system can read at query time.
The gap between them is small but real: a machine-readable statement, attached to an individual fact, of how much it should be trusted right now and what a consumer should do if the answer is “not much”.
The evidence
We hold 1,884 curated places across four countries. Every one was recorded by a person who went there. The dataset carries a creation date, which we use as a proxy for a verification date. It is not the same thing, and treating it as one flatters every number below.
237 records carry an explicit weekly closure rule. 44 say in their own text that their hours vary or that they close when they sell out. Those are the records that admit their own instability, and they should decay faster than the rest.
The median record is 227 days old. Nothing anywhere in the pipeline currently says so to the system consuming it.
A first attempt
Different facts go stale at different speeds. A street address survives years. A price does not survive a season. We model this as a half-life, the time after which a fact is worth half what it was, and apply exponential decay from the date it was verified.
These are judgements, not measurements. We have no longitudinal data on how quickly each field actually goes stale. They are published so they can be argued with.
Applied to a record of median age, 227 days, the same place yields quite different confidence depending on which fact you ask about. Every figure below is computed by the function described, at the moment this page was built.
State it plainly.
State it, and say when it was last checked.
State it, and say when it was last checked.
State it as likely rather than certain, and suggest verifying.
State it as likely rather than certain, and suggest verifying.
The output that matters is not the number. It is the last line: what a consuming system should actually do. A score of 0.31 means nothing to an assistant. “State it as likely rather than certain, and suggest verifying” is an instruction it can follow.
Where this is wrong
A confidence model whose failure modes are undisclosed is worse than no model. These are the ones we know about.
- −The half-lives are judgements, not measurements. We have no longitudinal data on how quickly each field actually goes stale.
- −Decay is not the only way a fact dies. A restaurant can close the week after it is verified, and this model would still report high confidence.
- −We use the record creation date as the verification date, because it is what we hold. They are not the same thing, and treating them as equal flatters every score.
- −Corroboration is only meaningful if sources are genuinely independent. Two creators copying the same listing is not two sources.
- −The model has no notion of season. A beach cafe verified in July is worth less in January than the arithmetic suggests.
- −It says nothing about whether the fact was correct when it was recorded. This measures age, not accuracy.
The largest of these is the first. Every half-life above is a guess by people who have built one dataset in one industry. The right way to fix that is measurement: revisit a sample of records, record what actually changed, and replace the assumptions with observations. We have not done that yet.
What we are not claiming
This is not a standard. It has one implementer, which is us, and a framework with one implementer is a product decision wearing a larger hat. It is not novel in its ingredients: provenance, decay and confidence scoring all have long literatures. It is not a replacement for C2PA or PROV, and if this idea has legs the sensible shape is probably a profile that sits alongside them rather than a competitor.
What we think is genuinely underserved is the specific combination: per-fact, time-aware, machine-readable, and expressed as guidance a retrieval system can act on rather than a number it has to interpret.
An invitation
If you work on retrieval, provenance or data quality and think this is wrong, we would like to hear why. If you hold a dataset with real verification dates and would test the model against it, that is the single most useful thing anyone could do with this. If you have longitudinal data on how quickly any of these fields actually change, it would replace our worst assumption with a real number.
Write to care@trundle.me. Disagreement is more useful to us than agreement, and we will publish corrections rather than quietly editing.
Related
The live comparison shows the practical consequence: a general model recommending places our records show closed on the day asked about. It runs against the same dataset described here.