CrawlerPulse
CrawlerPulse / Asia / India / lokmat.news18.com

News18 Marathi

lokmat.news18.com · India · Network18 Media & Investments Limited · measured 4 Aug 2026

Google Search
Allowed
Google-Extended
Allowed
OpenAI
Allowed
Anthropic
Allowed
Perplexity
Allowed
Meta
Allowed
Common Crawl
Blocked
ByteDance
Blocked
Apple
Allowed
Amazon
Allowed

robots.txt

As read on 4 Aug 2026, 1,944 bytes. 8 rule groups name AI crawlers; shown first.

User-agent: CCBot
Disallow: /

User-Agent: omgili
Disallow: /

User-Agent: omgilibot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: Timpibot
Disallow: /

User-agent: *
Allow: /
Disallow: /*studio18/*
Disallow: /*/undefined/
Disallow: /login/
Disallow: /login?ref*
Disallow: /cricketnext/
Disallow: /ganesh-chaturthi-festival/
Disallow: /elections/winner/
Disallow: /board-results-pubstack/
Disallow: /elections/lok-sabha/orissa/
Disallow: /elections/lok-sabha/chattisgarh/
Disallow: /elections/lok-sabha-election-schedule
Disallow: /elections/lok-sabha/nct-of-delhi/
Disallow: /elections/winner/
Disallow: /elections/*/assembly-election-all-winners-list/
Disallow: /elections/*/assembly-election-result/
Disallow: /amp/tag/tags-Q/
Disallow: /amp/tag/gmail/
Disallow: /amp/tag/tags-Z/
Disallow: /photogallery/coronavirus-latest-news/
Disallow: /amp/tag/pgstory/
Allow: /cricket/live-score/$
Disallow: /cricket/live-score/*
Disallow: /elections/*constituency-
Disallow: /elections/*candidate-
Disallow: /elections/*-led20
Disallow: /zip/*
Disallow: /*/zip/*
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap-index.xml
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap/today
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap/google-news.xml
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap-video-index.xml
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap/webstories-sitemap-index.xml
Sitemap: https://news18marathi.com/commonfeeds/v1/mar/sitemap-image-index.xml

User-agent: MAZBot
Disallow: /

User-agent: Baiduspider
Disallow: /

User-agent: FriendlyCrawler
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: Scrapy
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

History

Changes recorded since June 2026.

No change recorded.

In context

India
25.0% block at least one AI crawler
Behind Cloudflare
No
Same group