↑ ↓ select, Enter open, Esc close

Find broken links on a website with wget --spider

Adjust the values, the command updates live
user@server
LC_ALL=C wget --spider -r -l 2 -nd -nv -w 1 -o /tmp/spider.log https://example.com/; grep -A 100 'broken link' /tmp/spider.log

Follows all links up to the given depth, only checks whether they are reachable and writes the log to a file. At the end, all unreachable URLs are listed under “Found N broken links”. Internal 404s waste crawl budget and hurt the user experience.

Note: -w 1 waits one second between requests, so large sites take a while. wget respects robots.txt, so disallowed areas are not checked. LC_ALL=C forces English messages so that grep matches.

Also searched as

  • find broken links command line
  • wget spider broken link check
  • find 404 links on website linux
  • crawl website for dead links

Related one-liners

All in SEO checks

Read first, then run.

The commands on myline.de act directly on servers, files and databases. A wrong path or placeholder can delete data irreversibly or make a server unreachable.

  • All commands are provided without warranty and are not tested on every system.
  • Understand what a command does before running it, and check every placeholder.
  • Make a backup first and, if possible, try it on a test system.
  • You run commands at your own risk. Liability for damages is excluded to the extent permitted by law.