COVID demonstrated, at global scale and at speed, that public institutions cannot be trusted to preserve their own primary records without revision. This post opens the evidentiary record that will justify every component of CivicIPFS as the project is built. It does not argue about policy. It argues about data.
View article
View summary
This post is the first entry in a documented record, not the whole argument. Every technical component of CivicIPFS will be justified by a specific, named, sourced instance of the failure that made it necessary. Those specific citations will come as each component is designed. This post establishes the pattern those citations will fill.
I want to be clear about the method before anything else. This project is built on the same discipline it is arguing for. The argument is not political. It is not emotional. It is archival. The distinction will be held throughout.
What happened to the data
During the COVID pandemic, public health agencies at every level — municipal, national, and international — published primary data. Mortality figures. Comorbidity tables. Case definitions. Guidance documents. These were not opinions or projections. They were records of what was observed and recorded at a specific time by specific institutions.
Those records were revised. Repeatedly. The revisions were not always documented. The earlier versions were not always preserved. In many cases the earlier versions are no longer accessible through official channels. In some cases the methodology underlying the figures was changed without the change being clearly marked as a departure from previous methodology.
The co-mortality data — the causes recorded at the time of death, as entered by the attending clinician or registrar at the moment of recording — is the most important example. This data has a specific character: it is primary, contemporaneous, and irreversible in principle. A death certificate records what was observed and concluded at a specific moment. It cannot be legitimately revised after the fact without a formal, documented, auditable process. What happened in practice diverged from that principle in ways that varied by jurisdiction but were widespread enough to constitute a pattern, not a series of isolated incidents.
This is not a claim about what the data showed. It is a claim about what happened to the data. Those are different claims, and only the second one is being made here.
Why this is not a political argument
Public health decisions during COVID were made under genuine uncertainty, under time pressure, and with incomplete information. Reasonable people disagree about whether those decisions were correct. That disagreement is legitimate and is not what this project is about.
What is not a matter of reasonable disagreement is whether primary records should be preserved. The answer to that question does not depend on what the records show. A record that supports a decision should be preserved. A record that complicates a decision should be preserved. A record that contradicts a decision should be preserved. The value of a primary record is precisely that it is not subject to subsequent editorial judgment about whether it remains convenient.
The failure documented here is not that institutions made decisions we might disagree with. The failure is that the evidentiary basis for evaluating those decisions — the primary data that existed at the time — was not preserved in a form that makes independent verification possible.
That failure is archival. It has a technical cause and a technical remedy.
The pattern is not new
COVID did not introduce this failure. It made it visible at a scale and speed that was impossible to ignore.
Environmental impact assessments are revised after approvals are granted. Financial disclosures are restated after investigations begin. Safety certifications are updated after incidents occur. Planning documents are amended after developments are approved. Clinical trial registrations are modified after results are known.
In each case the pattern is the same. An institution publishes a record. Circumstances change. The record is revised. The earlier version becomes difficult or impossible to locate through official channels. The revision may or may not be documented. The gap between what was originally stated and what is currently claimed to have been stated becomes a matter of contested memory rather than verifiable fact.
COVID compressed this pattern into months instead of years and applied it to data that affected every person on the planet simultaneously. That compression made the pattern undeniable to anyone paying attention to the data rather than the narrative around it.
What would have had to exist
If the co-mortality data published in 2020 had been preserved in a form that made independent verification possible, three things would have had to be true.
First, the data would have had to be stored in a content-addressed form — meaning the storage address is derived from the content itself, so any alteration of the content produces a different address. A revised dataset is not the same dataset with a new date. It is a different dataset. The difference is detectable and provable, not arguable.
Second, the storage would have had to be distributed across multiple independent operators with no common point of control. A single server, a single institution, a single jurisdiction is a single point of pressure. Coordinated removal from a distributed network requires coordinating across every operator simultaneously. That coordination is difficult, visible, and leaves a record of its own.
Third, the storage commitment would have had to be verifiable by parties outside the network. An operator who claims to be preserving a record and is not preserving it should be detectable. The commitment to preserve is only meaningful if the failure to preserve is observable.
None of this infrastructure existed in a form accessible to the people who needed it in 2020. It did not exist because there was no sustainable model for running it, no governance layer for operating it across organisational boundaries, and no accountability mechanism for verifying that commitments were being kept.
Those are the three gaps CivicIPFS is designed to fill.
What this project is documenting
As each component of CivicIPFS is designed and built, this record will be extended with specific, named, sourced instances of the failure that component addresses.
When the timestamping integration is specified — the mechanism that anchors a content address to an external, manipulation-resistant clock — the post that describes it will open with a specific example of a document whose publication date became contested because no independent timestamp existed.
When the federation model is specified — the mechanism that distributes storage across operators in different jurisdictions — the post that describes it will open with a specific example of content that was removed because all copies were under common control.
When the accountability layer is specified — the mechanism that makes it detectable when an operator claims to be preserving content and is not — the post that describes it will open with a specific example of a preservation commitment that was not kept and was not detectable until the content was needed.
The technical documentation and the evidentiary documentation will grow together. Every design decision will have a reason. Every reason will have a source.
The invitation
If you were paying attention to the data during COVID — not the policy debate, the data — and you observed specific, documented instances of revision, deletion, or inaccessibility that fit this pattern, this project wants to know about them. Not as anecdotes. As sourced, dated, specific records of what existed and what subsequently happened to it.
The same applies to any domain where this pattern has been observed. Public health is the opening case because it is the most recent and the most visible. It is not the only case and will not be the only case documented here.
This project is persistent. It is not in a hurry. It will get the citations right before it publishes them. But it will publish them.