Skip to content
DopeSwagYolo

AI & Cybersecurity

What is a web crawler?

A web crawler is an automated program that visits web pages and downloads their content. Search engines use crawlers to build their indexes. Crawlers and related programs called scrapers also gather text and images used to train AI models.

Also known as: web spider, crawler bot

Researched and fact-checked by AI, with no human review. 5 sources listed below. How we verify

Last updated

How does a web crawler work?

A web crawler is an automated program that visits web pages and downloads what it finds. Mozilla's developer glossary describes a crawler as a program that systematically browses the web to collect data from pages. It says such programs are often called bots or robots.

Search engines are typical users. Google says its crawler, Googlebot, is also known as a robot, bot or spider. It finds new pages by following links from pages it already knows. Google says it uses a huge set of computers to crawl billions of pages. It also says its crawlers try not to crawl a site too fast, to avoid overloading it.

A scraper is a close relative. The Wikimedia Foundation uses the word for automated programs that copy its content, such as images, to feed AI models.

What rules do crawlers follow?

Site owners can publish a file named robots.txt. It lists which parts of a site crawlers may visit. The rules are set out in the Robots Exclusion Protocol, an Internet Engineering Task Force standards document published in September 2022. It builds on a method first defined in 1994.

The document says crawlers are requested to honor the rules. It also says the rules are not a form of access authorization. In other words, robots.txt is a request, not a lock.

Why are AI crawlers a problem for websites?

In April 2025, the Wikimedia Foundation said it was seeing a significant increase in requests. It said scraping bots collecting training data for large language models and other uses drove most of that traffic. It reported two figures:

  • Bandwidth used to download multimedia had grown 50% since January 2024.
  • At least 65% of its most resource-heavy traffic came from bots.

The foundation explained that crawlers read pages in bulk, including less popular ones. Those requests cost more to serve.

On October 5, 2026, the foundation said AI agents had crawled millions of its pages. It believes OpenAI operates the agents. It said their traffic may have contributed to a partial outage of one of its services in May.

Sources

Articles on AI & Cybersecurity