CrawlerPulse
CrawlerPulse / Methodology

Methodology

CrawlerPulse measures which AI crawlers news outlets let in, from the rules each outlet publishes in its robots.txt. This page explains how, and what the figures cannot show.

What is measured

Every news outlet publishes, or not, a robots.txt file: a list of rules telling each crawler which parts of the site it may read. CrawlerPulse reads that file for every outlet in its directory and evaluates the rules for each crawler it tracks.

The figures are about what outlets ask, not about what crawlers do. A robots.txt is a request; whether a crawler honours it is up to its operator.

How often

Every outlet is measured every day. A pass starts at 00:00 UTC and the site is rebuilt and published as soon as it ends. The date of the last pass is at the top of every page.

Verdicts

For each outlet and each crawler, the rules are tested against the home page and against pages that carry stories (a news section, an article, a category).

Allowed
The crawler may read the home page and the story pages.
Partial
The home page is open, but the pages that carry stories are closed to the crawler.
Blocked
The site itself is closed to the crawler, or the site tells every robot not to index it.

An outlet blocks a crawler when the verdict is Partial or Blocked. An outlet blocks at least one AI crawler when that is true for any of the nine AI crawlers. A site with no robots.txt allows every crawler.

Outlets that cannot be measured are counted apart, never as allowing or blocking: not reached (the site did not answer, or refused the request; 1,686 today), pending (not measured yet; 2) and inactive (no story published since before 2024; 2,889). Only measured outlets have a page.

The crawlers

Nine AI crawlers, and Google's search crawler for comparison: see the list, with what each operator says it is for. Two of them, Google-Extended and Applebot-Extended, are not crawlers but robots.txt tokens: Google and Apple read the site with their usual crawler and use the token to decide whether the content may train their AI models.

The directory

17,293 news sites in 227 countries, built from public sources, including Wikidata. Outlets are ordered by public reach, from the Tranco list of the most visited sites. Each outlet is placed in its country of origin; a domain that serves several countries counts in each of them. Countries are grouped into continents following the United Nations M49 standard, with South America apart from the rest of the Americas.

What the figures cannot show

A site can block a crawler without saying so in its robots.txt, at the network level: many news sites sit behind Cloudflare, whose bot protection can refuse AI crawlers whatever the file says. Outlet pages say whether a site is behind Cloudflare, but the verdicts only reflect the published rules.

A site can also have agreements with AI companies that a robots.txt does not reveal, and a crawler can ignore the file.

Data and reuse

The figures on this site are free to reuse under the Creative Commons Attribution 4.0 license: cite CrawlerPulse (crawlerpulse.com). Researchers and journalists can request the full dataset, one row per outlet and one column per crawler, at data@crawlerpulse.com. The same address takes corrections: if a verdict looks wrong, write and say which outlet.

Who makes it

CrawlerPulse is an independent project by LifeWorld Labs.