Scrapebox for Expired Domains: Step-by-Step
Scrapebox is the oldest workflow on this site and the cheapest that still works: a one-time licence, some proxies, and as much patience as you can spare. It is also the most misunderstood, because Scrapebox is a general harvesting toolkit rather than a domain tool, and nothing about it is optimized for expired domains until you assemble the pieces yourself. This walkthrough assembles them in order, from seed list to vetted shortlist. If you have not yet decided whether the DIY route suits you, our comparison of expired domain crawlers covers where Scrapebox sits against cloud alternatives.
What you need before you start
- Scrapebox itself: a one-time licence, Windows software; many hunters run it on a cheap Windows VPS so crawls do not monopolize their own machine.
- Proxies: non-negotiable for harvesting, because search engines throttle repeated automated queries within minutes.
- The addons: the free TDNAM Scraper addon for GoDaddy auction inventory and the premium Expired Domain Finder plugin for outbound-link crawling.
- A seed list: 30-100 URLs from your niche; this determines the quality of everything downstream.
The six steps at a glance
| Step | What you do | Rough time | Output |
|---|---|---|---|
| 1. Build the seed list | Collect niche resource pages, association sites, old blogrolls | 30-60 minutes | A seeds file of 30-100 URLs |
| 2. Harvest URLs | Run the harvester on seeds and niche footprints through your proxies | 1-2 hours, mostly unattended | Tens of thousands of raw URLs |
| 3. Crawl for expired domains | Feed the list to the Expired Domain Finder at crawl depth 1-2 | Hours, unattended | Candidate expired domains with source pages |
| 4. Pull auction inventory | Run the TDNAM Scraper against GoDaddy auction listings | Minutes | Expiring auction names to merge in |
| 5. Filter the raw output | Dedupe, strip junk TLDs, re-verify availability | About 30 minutes | A clean shortlist |
| 6. Enrich with metrics | Score survivors with external link and spam data | 30-60 minutes | A vetted buy list |
Step 1: build a seed list worth crawling
Everything Scrapebox finds descends from your seeds. Collect the pages your niche actually links from: industry association link pages, university topic directories, conference sites, and resource roundups, especially ones last updated years ago, since unmaintained pages are the densest in dead outbound links. Aim for 30-100 URLs in a plain text file. The seed-selection recipes in our guide to finding expired domains in your niche apply here without modification.
Step 2: harvest URLs at scale
Load your proxies, set the harvester to your search engines of choice, and combine niche keywords with footprints that locate link-heavy pages: phrases like "useful links", "resources" or "recommended sites" alongside your topic terms. Let it run until you have tens of thousands of URLs; volume is fine at this stage because filtering comes later. Trim the obvious noise afterwards with the built-in remove-duplicates and filter tools.
Step 3: crawl the outbound links
The Expired Domain Finder plugin takes your harvested URLs, visits each page, extracts every outbound link and tests whether the linked domains are still registered. Crawl depth matters: depth 1 checks only the pages you feed it, depth 2 follows their links outward one hop, and anything deeper multiplies runtime dramatically. Depth 1-2 from good seeds is the sweet spot. Expect long unattended runs and keep the thread count polite; hammering small sites gets your proxies burned and is needless. The output is a list of candidate domains, each with the page that still links to it, which is exactly the context you want for later outreach or rebuild decisions.
Step 4: pull GoDaddy auction inventory with the TDNAM scraper
The free TDNAM Scraper addon queries GoDaddy Auctions inventory, which adds 35,000+ new expiring domains every day, and pulls listings into the same workspace so you can merge marketplace names with your crawl finds. This supplements rather than replaces the crawl: auction inventory is public and contested, while crawl finds are effectively private. It is still worth running, because closeouts starting around $5 occasionally hide genuine niche matches nobody bid on.
Step 5: filter the raw output
Raw Scrapebox output always overstates reality. Dedupe first, cut TLDs you would never buy, then re-verify availability in smaller batches: bulk WHOIS checks through strained connections produce false positives, and a domain reported available at 2 am may be in redemption or already caught. Anything that shows as registered again gets dropped; anything in an expiry stage goes to a watch list, since a typical .com takes 65-80 days from missed renewal to actual drop.
Step 6: enrich survivors with real metrics
Scrapebox tells you a domain is free; it tells you nothing about whether it is worth having. Push the shortlist through a metrics layer before money moves. The manual route is Wayback Machine history plus Majestic or Moz lookups domain by domain. The batch route is an aggregator: DomCop includes a credits system for exactly this job, letting you upload your own list and get Majestic, Moz and spam signals back in one table. Whichever route you take, apply the standard floors: Trust Flow to Citation Flow ratio at 0.3 or above, Moz Spam Score under 10%, and a history check that shows a real site rather than a parked page or a foreign-language casino interlude.
Is the Scrapebox route still worth it in 2026?
Honestly: only for some people. The cash cost is unbeatable and the flexibility is real, but the hours are heavy, the software is dated, and cloud crawlers now do the crawl-and-score loop automatically while you sleep. Scrapebox remains the right choice for tinkerers who already own it and enjoy the control, and the wrong choice for anyone whose spare time is the scarcest resource in their stack. For the full landscape of alternatives, see our ranking of all nine ways to find expired domains.
Frequently asked questions
Do I really need proxies for Scrapebox?
For harvesting, yes. Search engines throttle repeated automated queries almost immediately, and a proxy pool spreads the load. The expired domain crawl itself is gentler, but availability checking in bulk also benefits from rotation.
Is scraping for expired domains allowed?
Harvesting public search results and crawling public pages sits in a long-standing gray-but-tolerated zone; being responsible means respecting robots.txt, keeping request rates low and never hammering small sites. The domains themselves are simply expired registrations, and registering one is entirely legitimate.
Why do my available domains turn out to be taken?
Three usual causes: stale WHOIS responses during bulk checks, domains sitting in expiry stages that look free but are not yet droppable, and other hunters catching names between your check and your registration attempt. Re-verify every domain individually right before you buy.
How is this different from DomCop's personal crawlers?
Same concept, different economics. Scrapebox trades hours for cash: you run everything locally and enrich data yourself. Cloud crawlers run 24/7 server-side and return pre-scored finds, at subscription prices. Frequency of hunting decides which trade is rational.