網路安全
robots.txt compliance is optional
Lines 152 and 153 of nytimes.com/robots.txt block the Internet Archive's crawler by its exact current name, and the Wayback Machine captured the site anyway. The same file blocks Common Crawl with identical syntax, and Common Crawl's August index holds nothing from the domain but the robots.txt file itself. Here is what robots.txt actually obliges anyone to do, and the commands that show you what yours got.