↑ ↓ select, Enter open, Esc close

Extract and count the internal links of a page

Adjust the values, the command updates live
user@server
curl -s https://example.com/ | tr '\n' ' ' | grep -oiE 'href="[^"]+"' | sed -E 's/^href="//I; s/"$//; s/#.*//' | grep -E '^(/|https?://(www\.)?example.com)' | sort | uniq -c | sort -rn

Pulls all href targets from the HTML, removes anchors and keeps only links that start with / or point to the site’s own domain. The frequency shows which pages are linked especially often from navigation and footer. Watch out for links to http variants, staging domains or old URLs that redirect.

Note: Relative links without a leading slash are missed, and stylesheets from <link href> are included. The list can be fed into the status code check from the sitemap audit.

Also searched as

  • extract all links from a page
  • check internal linking command line
  • which pages are linked
  • extract links from html curl

Related one-liners

All in SEO checks

Read first, then run.

The commands on myline.de act directly on servers, files and databases. A wrong path or placeholder can delete data irreversibly or make a server unreachable.

  • All commands are provided without warranty and are not tested on every system.
  • Understand what a command does before running it, and check every placeholder.
  • Make a backup first and, if possible, try it on a test system.
  • You run commands at your own risk. Liability for damages is excluded to the extent permitted by law.