Scroll World 0%
Decay, measured

The half life of a link

A quarter of the web that existed between 2013 and 2023 is already unreachable, and the older a link is the worse its odds get. That is the visible half of the problem. The other half is the links that still resolve and no longer say what they were cited for.

Scroll World 28 August 2026 8 min read
The shape of it

A link is a promise that somebody else has to keep.

Nobody signs it, nobody is paid to honour it, and it expires on a schedule that is now measured well enough to plan around.

The measurement almost nobody had done

In 2023 the Pew Research Center did something that sounds trivial and had never been done at this scale. It took a sample of web pages known to have existed in each year from 2013 onwards, went back in October 2023, and asked of every one of them whether it still loaded.

A quarter of them did not. Of the pages that existed in 2013, thirty eight percent were gone. Of the pages that existed in 2023, checked inside the same year, eight percent were already gone.

That last figure is the one worth sitting with. This is not a story about the old web crumbling while the new web holds. Eight percent of a year's pages do not survive that year.

38%
Pages from 2013 that no longer loaded by late 2023Pew Research Center
25%
Of everything sampled from 2013 to 2023, already unreachablePew Research Center
8%
Pages from 2023 that were gone inside the same yearPew Research Center

That curve is from a Harvard Law School study of every outbound link the New York Times published between 1996 and 2021, which came to a little over two million of them.

Read it as a decay curve rather than as a comment on the nineteen nineties. A link from 2018 was three years old when it was tested and had lost six percent of its cohort. A link from 2008 was thirteen years old and had lost forty three percent. Nothing about the 1998 links was worse except their age.

What drift looks like when a person actually reads the page

The same Harvard team did the part that cannot be automated. They took a sample of four and a half thousand links that still resolved, and had people read both the citing article and the page it pointed at.

Thirteen percent had drifted significantly away from what they had been cited for. The age gradient turned up again: four percent of the links from 2019 had drifted, against twenty five percent of the links from 2009.

Rot and drift are one process seen from two sides. A page either stops existing or stops meaning what it meant, and the longer it has to do either, the more likely it is that it has.

Two ways to leave a citation behind

Point at it

The default, and what every link on this page does. You write a URL and then inherit somebody else's hosting decisions, budget cycles and content strategy, permanently and without ever being consulted about any of them.

It costs nothing on the day and it carries the failure rate above. For most links that is the right trade, because most links are decoration on a sentence that stands up without them.

Keep a copy

Capture the page at the moment you cite it and point at both. Perma.cc was built for this after a 2014 study found that more than seventy percent of the URLs in law journals, and fifty percent of those in United States Supreme Court opinions, no longer led to what had been cited. The Internet Archive has been doing the same job for the whole web since 1996 and passed a trillion pages in October 2025.

The cost is that somebody has to decide, at the moment of writing, which links are the ones that have to survive. That is an editorial judgement and it does not automate.

The parts of the web built to cite other things are the parts most exposed to this, and they are the parts everyone treats as permanent.

Where the rot lands hardest

Why a court is worse off than a blog

A law review article and a Supreme Court opinion are almost entirely made of references. They are load bearing in a way an ordinary page is not: remove the source and the reasoning does not merely lose a footnote, it loses its foundation. Those documents are also written to be read in fifty years.

Wikipedia is in the same position and shows the same result. Eleven percent of everything Wikipedia cites is already unreachable, and because citations cluster, that is enough to put at least one dead reference on more than half of all articles.

The pattern is consistent. The more a document depends on pointing outside itself, the faster it hollows out, and the less any single owner can do about it.

What we take from it

  1. Treat a link as perishable

    Assume anything you point at carries a failure rate that climbs with its age. A link is not a fact you have recorded. It is a request you have made of a stranger, with no expiry date written on it and no way to renew it.

  2. Name and date every source, not just link it

    A dead URL with a publisher, a title and a date around it can still be found. A bare URL with nothing attached is unrecoverable the moment it stops resolving. Every source block on this site is written the long way for that reason.

  3. Capture the load bearing ones

    If an argument falls over without a particular page, that page is worth archiving on the day you cite it rather than on the day somebody notices it has gone.

  4. Do not trust a link checker

    It tests whether an address answers. It cannot tell you that the answer changed, and drift is the failure that will quietly hollow out an old page while every automated check on it stays green.

The part that applies to us

This site is a pile of argument held together by outbound links, and everything above applies to it exactly as written. The sources at the bottom of this page will rot at the rate the rest of the web rots at, and knowing that does not exempt us from it.

What we do about it today is small, and worth stating plainly rather than dressing up. Every source is named, dated and attributed as well as linked, so a reader who hits a dead URL still has enough to go and find it somewhere else. We are not capturing archived copies at the point of citation. We should be, and this page is the reason we now know it.

There is a wider point for anyone building a page meant to last. A page whose argument lives inside itself, in its own numbers and its own structure, decays far more slowly than a page that is mostly a set of pointers at other people's work. That is not an argument against citing things. It is an argument for the page being worth reading before the links are counted.

Sources

Pew Research Center, When Online Content Disappears, published 17 May 2024, sampling web pages from 2013 to 2023 and checking them in October 2023: When Online Content Disappears

John Bowers, Clare Stanton and Jonathan Zittrain, The Paper of Record Meets an Ephemeral Web, published in the Columbia Journalism Review on 21 May 2021, analysing 2,283,445 links across 553,693 New York Times articles: What the ephemerality of the Web means for your hyperlinks

Jonathan Zittrain, Kendra Albert and Lawrence Lessig, Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations, 127 Harvard Law Review Forum 176, published 2014: The Harvard Law School record of the paper

Internet Archive, one trillion web pages preserved, announced 22 October 2025: The web we have built

Research current as at 28 August 2026. Every figure on this page comes from one of the four sources above. The one inference we have drawn ourselves is marked as ours in the place it appears.

Ask us to build one that lasts

Keep reading