Scrapebox for Expired Domains: Step-by-Step

Published August 27, 2026

The short answer: harvest URLs from niche seed sites, crawl their outbound links with the Expired Domain Finder addon, check availability, then push survivors through a real metrics tool before registering anything. Expect hours of tinkering and near-zero cash cost.

Scrapebox is the oldest workflow on this site and the cheapest that still works: a one-time licence, some proxies, and as much patience as you can spare. It is also the most misunderstood, because Scrapebox is a general harvesting toolkit rather than a domain tool, and nothing about it is optimized for expired domains until you assemble the pieces yourself. This walkthrough assembles them in order, from seed list to vetted shortlist. If you have not yet decided whether the DIY route suits you, our comparison of expired domain crawlers covers where Scrapebox sits against cloud alternatives.

What you need before you start

The six steps at a glance

StepWhat you doRough timeOutput
1. Build the seed listCollect niche resource pages, association sites, old blogrolls30-60 minutesA seeds file of 30-100 URLs
2. Harvest URLsRun the harvester on seeds and niche footprints through your proxies1-2 hours, mostly unattendedTens of thousands of raw URLs
3. Crawl for expired domainsFeed the list to the Expired Domain Finder at crawl depth 1-2Hours, unattendedCandidate expired domains with source pages
4. Pull auction inventoryRun the TDNAM Scraper against GoDaddy auction listingsMinutesExpiring auction names to merge in
5. Filter the raw outputDedupe, strip junk TLDs, re-verify availabilityAbout 30 minutesA clean shortlist
6. Enrich with metricsScore survivors with external link and spam data30-60 minutesA vetted buy list

Step 1: build a seed list worth crawling

Everything Scrapebox finds descends from your seeds. Collect the pages your niche actually links from: industry association link pages, university topic directories, conference sites, and resource roundups, especially ones last updated years ago, since unmaintained pages are the densest in dead outbound links. Aim for 30-100 URLs in a plain text file. The seed-selection recipes in our guide to finding expired domains in your niche apply here without modification.

Step 2: harvest URLs at scale

Load your proxies, set the harvester to your search engines of choice, and combine niche keywords with footprints that locate link-heavy pages: phrases like "useful links", "resources" or "recommended sites" alongside your topic terms. Let it run until you have tens of thousands of URLs; volume is fine at this stage because filtering comes later. Trim the obvious noise afterwards with the built-in remove-duplicates and filter tools.

Step 3: crawl the outbound links

The Expired Domain Finder plugin takes your harvested URLs, visits each page, extracts every outbound link and tests whether the linked domains are still registered. Crawl depth matters: depth 1 checks only the pages you feed it, depth 2 follows their links outward one hop, and anything deeper multiplies runtime dramatically. Depth 1-2 from good seeds is the sweet spot. Expect long unattended runs and keep the thread count polite; hammering small sites gets your proxies burned and is needless. The output is a list of candidate domains, each with the page that still links to it, which is exactly the context you want for later outreach or rebuild decisions.

Step 4: pull GoDaddy auction inventory with the TDNAM scraper

The free TDNAM Scraper addon queries GoDaddy Auctions inventory, which adds 35,000+ new expiring domains every day, and pulls listings into the same workspace so you can merge marketplace names with your crawl finds. This supplements rather than replaces the crawl: auction inventory is public and contested, while crawl finds are effectively private. It is still worth running, because closeouts starting around $5 occasionally hide genuine niche matches nobody bid on.

Step 5: filter the raw output

Raw Scrapebox output always overstates reality. Dedupe first, cut TLDs you would never buy, then re-verify availability in smaller batches: bulk WHOIS checks through strained connections produce false positives, and a domain reported available at 2 am may be in redemption or already caught. Anything that shows as registered again gets dropped; anything in an expiry stage goes to a watch list, since a typical .com takes 65-80 days from missed renewal to actual drop.

Step 6: enrich survivors with real metrics

Scrapebox tells you a domain is free; it tells you nothing about whether it is worth having. Push the shortlist through a metrics layer before money moves. The manual route is Wayback Machine history plus Majestic or Moz lookups domain by domain. The batch route is an aggregator: DomCop includes a credits system for exactly this job, letting you upload your own list and get Majestic, Moz and spam signals back in one table. Whichever route you take, apply the standard floors: Trust Flow to Citation Flow ratio at 0.3 or above, Moz Spam Score under 10%, and a history check that shows a real site rather than a parked page or a foreign-language casino interlude.

Is the Scrapebox route still worth it in 2026?

Honestly: only for some people. The cash cost is unbeatable and the flexibility is real, but the hours are heavy, the software is dated, and cloud crawlers now do the crawl-and-score loop automatically while you sleep. Scrapebox remains the right choice for tinkerers who already own it and enjoy the control, and the wrong choice for anyone whose spare time is the scarcest resource in their stack. For the full landscape of alternatives, see our ranking of all nine ways to find expired domains.

Frequently asked questions

Do I really need proxies for Scrapebox?

For harvesting, yes. Search engines throttle repeated automated queries almost immediately, and a proxy pool spreads the load. The expired domain crawl itself is gentler, but availability checking in bulk also benefits from rotation.

Is scraping for expired domains allowed?

Harvesting public search results and crawling public pages sits in a long-standing gray-but-tolerated zone; being responsible means respecting robots.txt, keeping request rates low and never hammering small sites. The domains themselves are simply expired registrations, and registering one is entirely legitimate.

Why do my available domains turn out to be taken?

Three usual causes: stale WHOIS responses during bulk checks, domains sitting in expiry stages that look free but are not yet droppable, and other hunters catching names between your check and your registration attempt. Re-verify every domain individually right before you buy.

How is this different from DomCop's personal crawlers?

Same concept, different economics. Scrapebox trades hours for cash: you run everything locally and enrich data yourself. Cloud crawlers run 24/7 server-side and return pre-scored finds, at subscription prices. Frequency of hunting decides which trade is rational.

Sources