#crawling
3 Einzeiler mit diesem Stichwort.
Alle Einzeiler mit Tag crawling
Kaputte Links einer Website mit wget --spider finden
$LC_ALL=C wget --spider -r -l 2 -nd -nv -w 1 -o /tmp/spider.log https://example.com/; grep -A 100 'broken link' /tmp/spider.log
Prüfen, ob der Server 304 Not Modified liefert
$curl -s -o /dev/null -w '%{http_code}\n' -H "If-Modified-Since: $(curl -sI https://example.com/ | grep -i '^last-modified:' | cut -d' ' -f2- | tr -d '\r')" https://example.com/
robots.txt abrufen und Disallow-Regeln anzeigen
$curl -s https://example.com/robots.txt | grep -iE '^\s*(user-agent|disallow|allow|sitemap):'