Mining Wikipedia's Dead Links for Expired Domains
Wikipedia cites millions of external pages, and the web underneath those citations rots constantly. Editors flag the failures, which means one of the most authoritative sites on the internet maintains a public, continuously updated map of dying domains. Some of those domains are expired and catchable, and whoever catches one holds a domain still referenced by Wikipedia. The technique has been quietly used since the 2000s and remains underdocumented. Here is the whole workflow, including the honest parts most write-ups skip: the links are nofollow, the pool is shrinking, and the win rate is low but the wins are unusual quality.
Why are Wikipedia dead links worth mining?
Three reasons. First, curation: a domain cited by Wikipedia editors was, at some point, a source someone considered worth referencing, which filters out an enormous amount of junk before you ever look. Second, the reference itself: a rebuilt domain that still sits in a Wikipedia citation earns referral visits from a page people actually read, and Wikipedia content is endlessly scraped and republished, so one citation tends to propagate into copies across the web. Third, competition: this is manual work with no push-button tool in 2026, so almost nobody does it systematically. Be clear about what it is not: Wikipedia external links carry nofollow, so this is not a direct PageRank play, and anyone selling it as one is selling the wrong decade.
Where does Wikipedia keep its dead links?
When a citation target stops resolving, editors or bots tag it with a dead link template, and tagged articles collect in maintenance categories such as "Articles with dead external links". Every major language edition runs the same system, which matters later, because the German, French and Spanish pools are mined even less than the English one. You can browse the categories directly on Wikipedia, but search operators are usually faster for niche work:
site:en.wikipedia.org "dead link" yourkeywordfinds articles in your topic carrying at least one flagged citation.- Adding
intitle:terms narrows to articles whose subject matches your niche rather than passing mentions. - Swapping the language subdomain (de, fr, es) repeats the hunt in less contested editions.
Tools in the WikiGrabber style used to automate these queries; most have come and gone over the years, and the manual operators above do the same job without depending on anyone's side project surviving.
From flagged citation to owned domain: the six steps
| Step | Tool | What you are checking | Kill signal |
|---|---|---|---|
| 1. Find flagged references | Search operators or maintenance categories | Articles in your topic with dead citations | None yet; you are gathering |
| 2. Extract root domains | Copy the cited URLs, strip to the registrable domain | Unique domains worth testing | Subdomain platforms you cannot register |
| 3. Check registration status | WHOIS or bulk availability checks | Available, in an expiry stage, at auction, or still active | Active and renewed; a dead page on a live domain is not an expired domain |
| 4. Check the original content | Wayback Machine | What the cited page actually said | Parking pages, spam interludes, content you cannot honestly rebuild |
| 5. Check metrics and trademarks | Majestic or Moz data plus a trademark search | The link profile beyond Wikipedia | TF:CF ratio under 0.3, spammy anchors, a live trademark on the name |
| 6. Register or catch | Registrar, or a backorder service if it has not dropped yet | The route to ownership | A contested drop you cannot justify; even top catchers land an estimated 30-50% of contested names |
Vetting: most dead links are not opportunities
Expect a brutal funnel, and treat that as normal rather than discouraging. Most flagged citations point at pages that moved, sites that redesigned, or domains that are still registered and simply broken. Of the genuinely expired remainder, many spent their afterlife as parked pages or worse, which is why the Wayback check in step 4 is not optional: you are looking for a domain whose last real content matches what Wikipedia cited, with no casino chapter in between. For status checks in step 3, WHOIS status codes tell you exactly where a domain sits in the expiry pipeline; redemptionPeriod means the old owner can still restore it, while pendingDelete means it drops within 5 days and can be backordered. When a batch of candidates survives to step 5, a bulk metrics pass through an aggregator saves the domain-by-domain lookups; DomCop handles uploaded lists through its credits system and returns Majestic and Moz data in one table.
Keeping the citation alive after you catch it
Owning the domain is half the value; the citation is the other half, and it survives only if the reference still makes sense. Rebuild the cited content at its original URL, sourced from the Wayback copy, and the reference keeps doing its job. Redirect the domain to a sales page and an editor will eventually cut it, reasonably. Two honest caveats belong here. Wikipedia's own bots increasingly rescue dead citations by swapping in Wayback archive URLs, which shrinks the pool of live opportunities over time, so fresh flags are worth more than old ones. And Google's expired domain abuse policy, published March 5, 2024, draws the line at intent: genuinely rebuilding cited content is explicitly fine, while harvesting Wikipedia-linked domains purely to manipulate rankings is the exact behavior the policy names as spam.
Scaling the hunt without ruining it
Manual mining suits a niche site builder hunting a handful of domains. If you want volume, harvest citation URLs at scale with the same tooling covered in our Scrapebox walkthrough, then run the same six steps in bulk. Pair the finds with the topical filters from our guide to niche filtering so you only chase domains that fit what you are building, and remember this is one method of many; our ranked overview of how to find expired domains shows where it fits. As a rough illustrative expectation, an afternoon of mining in a mid-sized niche tends to yield a shortlist of a few domains, of which one may survive full vetting. Low volume, unusual quality.
Frequently asked questions
Are Wikipedia links nofollow?
Yes, all external links on Wikipedia carry nofollow, and that has been true for many years. The value of a caught citation is referral traffic, credibility, and the secondary links created when other sites republish Wikipedia content, not direct link equity.
Is this against Wikipedia rules?
Registering an expired domain is entirely legal and outside Wikipedia's control. What violates Wikipedia norms is abusing the encyclopedia: inserting your own links, or turning a cited domain into spam. Rebuild the cited content honestly and the citation tends to stand; abuse it and editors will remove it.
How do I check whether a flagged domain is actually available?
Run a WHOIS lookup and read the status codes. No record usually means registrable right now; redemptionPeriod means the old owner can still reclaim it; pendingDelete means it drops within 5 days and needs a backorder. A dead page on a registered domain is the most common outcome and means move on.
How many usable domains does this method produce?
Few, by design. This is a low-volume, high-selectivity method: hours of manual work for a small shortlist. It rewards niche builders who want one great domain, not portfolio buyers who need fifty this month.