Skip to content
English
  • There are no suggestions because the search field is empty.

Which Internet Crawlers Can Access Your Real Geeks Website?

Real Geeks  Bot, Crawler, and Automated Traffic Policy: Learn how Real Geeks manages search engines, AI tools, and other automated visitors.

A crawler is a computer program that reads pages on the internet. Some crawlers help people find your website. Others collect your content without sending visitors back. Real Geeks manages these crawlers automatically to protect your website.

 

Need to Know

  • Search Engine crawlers from Google Bing and Other Search engines are never slowed or blocked.
  • Selected AI tools can read your website at a controlled speed, but may be limited if they attempt to access too many pages at a time. We have these controls in place to prevent them from slowing down your website for the human visitors you want using the site!
  • Crawlers that may slow down your website or collect content without providing any return traffic are blocked.
  • These settings do not negatively impact your website’s search ranking.

Table of Contents


Why Real Geeks Manages Crawlers

Right now, about 15% of all traffic to sites on this platform is AI crawlers. Not people. Robots, reading your pages to feed AI products. That share was near zero two years ago. Not every crawler is well behaved. A polite one takes a few pages at a time and backs off when you ask it to. An impolite one will request thousands of pages a minute and has no idea that your site is also trying to serve a buyer who is looking at a listing right now. Enough of that and your site slows down. Enough beyond that and it stops responding altogether. That is the first reason limits exist: a robot should never be able to cost you a lead.

The second reason is what you get back. Not all AI crawling is the same deal. When someone asks an AI assistant about homes in your market and it answers with your neighborhood page and a link to it, that crawl earned its keep -- that is a referral, and it is a channel worth being in. Other crawlers take exactly the same content and return nothing. Your listing copy and photography go into a product you will never appear in, a dataset resold to someone else, or a lead list built from the contact details on your own site. Same cost to you, no visit back.

So the test is not "is it AI." It is whether the crawler sends real people back to you. Crawlers that do are welcome. Crawlers that might, but could overwhelm your site doing it, get a budget. Crawlers that take and return nothing are turned away.

Real Geeks considers two questions when managing a crawler:

  • Does it send people back to your website?
  • Does it read your website at a safe speed?

These rules run on every site we host, so you get the protection without having to configure anything or think about it.

Back to top

Crawlers That Are Allowed

Real people are always free to browse your website, view listings, search for homes, and sign up as leads. These rules only manage automated tools called crawlers. Limiting crawlers that make too many requests helps keep your website fast and available for real buyers and sellers.

Some crawlers are also always allowed because they help people find and share your website. 

Back to top

Who We Allow, Manage, and Block

This table below shows which automated tools can visit your website, which have limits, and which are blocked. The tools marked as "ALLOWED" are just examples—any tool not listed is allowed by default. Tools are grouped by the company that runs them.

IDENTIFIER OPERATED BY WHAT IT DOES STATUS
Search Engines and Link Previews
Googlebot, Bingbot, DuckDuckBot Google, Microsoft, DuckDuckGo Indexes your pages for organic search results. ALLOWED
facebookexternalhit, twitterbot, LinkedInBot, Slackbot Meta, X, LinkedIn, Slack Builds the preview card when a link to your site is shared. ALLOWED
AI Assistants and Answer Engines
GPTBot, OAI-SearchBot  OpenAI Reads pages for ChatGPT answers and model training. MANAGED
Claudebot, Claude-SearchBot Anthropic Reads pages for Claude answers and for model training. MANAGED
Google-extended Google Feeds Gemini. Separate syste from Googlebot. MANAGED
Applebot* Apple Powers Siri, Spotlight, Safari suggestions. MANAGED
PerplexityBot*  Perplexity Reads pages to answer questions and cite sources. MANAGED
Blocked: Bulk Content Collection for AI Training
CCbot Common Crawl Archives the open web in bulk and redistributes it as training data. BLOCKED
Facebookbot, Meta-ExternalAgent, meta-externalagent, Meta-ExternalFetcher Meta Collects content for Meta’s AI models. Not the link-preview crawler. BLOCKED
Bytespider, TikTokSpider ByteDance Collects content for ByteDance and TikTok AI products. BLOCKED
Amazonbot Amazon Collects content for Alexa and Amazon AI services. BLOCKED
Google-CloudVertexBot Google Fetches pages for Vertex AI customers. Not search. BLOCKED
cohere-ai, cohere-training-data-crawler Cohere Collects training data for Cohere’s models. BLOCKED
PanguBot Huawei Collects training data for the PanGu models. BLOCKED
AI2Bot, AI2Bot-Dolma Allen Institute Builds open research training datasets. BLOCKED
Omgili, Omgilibot, webzio-extended Webz.io Harvests web content and resells it as a data feed. BLOCKED
diffbot Diffbot Converts sites into structured data products sold to third parties. BLOCKED
ImagesiftBot ImageSift Collects images at scale for a search and training index. BLOCKED
img2dataset Open-source tool Bulk-downloads images to assemble training datasets. BLOCKED
FirecrawlAgent Firecrawl Scrapes sites on demand and feeds the output to other apps. BLOCKED
AwarioBot, AwarioSmartBot, AwarioRssBot Awario Scrapes pages for a brand-monitoring product. BLOCKED
Meltwater Meltwater Scrapes pages for a media-monitoring product. BLOCKED
Sentibot Sentisum Collects content for sentiment analysis products. BLOCKED
peer39_crawler Peer 39 Crawler Profiles page content for ad-targeting classification. BLOCKED
Factset_spyderbot FactSet Collects web content for financial data products. BLOCKED
aiHitBot aiHit Builds company datasets from scraped web content. BLOCKED
Seekr Seekr Collects content for AI ranking and scoring products. BLOCKED
VelenPublicWebCrawler Velent Collects public web content for resale as a dataset. BLOCKED
Timpibot Timpi Builds a decentralized index from crawled content. BLOCKED
ICC-Crawler NICT (Japan) Research crawler collecting bulk web content. BLOCKED
Kangaroo Bot Kangaroo LLM Collects training data for an open language model. BLOCKED
Cotoyogi Cotoyogi Collects training data for AI model development. BLOCKED

Blocked — AI answer engines that do not send traffic back

Youbot You.com Generates answers without meaningful referral traffic. BLOCKED
DuckAssistBot DuckDuckGo AI answer crawler, separate from the search index. BLOCKED

Blocked — SEO and competitive research tools

SemrushBot, SemrushBot-OCOB Semrush Crawls your site so subscribers can analyze it as a competitor. BLOCKED
AhrefsBot Ahrefs Builds a backlink and keyword database sold as a subscription. BLOCKED
MJ12bot Majestic Builds a commercial backlink index. BLOCKED
opensiteexplorer Moz Builds a commercial link and domain-authority index. BLOCKED
DataForSeoBot DataForSEO Crawls sites to resell SEO data through an API. BLOCKED
serpstatbot Serpstat Crawls sites for a competitor-analysis platform. BLOCKED
BLEXBot WebMeUp Builds a commercial backlink index. Historically aggressive. BLOCKED
Barkrowler Babbar Crawls at high volume for a link-graph product. BLOCKED
Petalbot Huawei Crawls heavily for Petal Search; negligible US traffic. BLOCKED
GeedoBot, GeedoProductSearch Geedo Crawls for a product-search index. BLOCKED
FWAS Unattributed High-volume crawler with no stated purpose or contact. BLOCKED

Blocked — Contact and content harvesting

ZoominfoBot ZoomInfo Harvests names, emails, and phone numbers to resell as leads. BLOCKED
TurnitinBot Turnitin Copies page text into a private plagiarism corpus. BLOCKED

Blocked — Generic scraping tools and unidentified bots

Scrapy Open-source Default identifier of a widely used scraping library. BLOCKED
aiohttp Open-source Default identifier of an HTTP library used by scripts. BLOCKED
fidget-spinner-bot Unattributed No published operator, purpose, or contact. BLOCKED
my-tiny-bot Unattributed No published operator, purpose, or contact. BLOCKED

*Applebot and PerplexityBot moved from blocked to limited in the September 2026 review. The change takes effect with the next platform update.

Some AI crawlers have a crawl budget, enforced with HTTP 429 responses. It keeps them well-behaved, which is a necessity at the volume we get. Nothing is hidden from them and nothing is removed from their index - they come back and finish at a reasonable pace.

Identifiers are matched as text within the user agent string, so a company's variants are covered by the shortest form listed. The blocked tier covers 56 identifiers in total; the table groups them by operator.

Back to top

Traffic we filter that is not a crawler

Three more filters run alongside the crawler rules. They are aimed at attacks rather than robots, but they can occasionally affect a real person, so they belong on this page too.

  • Attack Probes. Automated scanners hunt every site on the internet for exposed admin panels, configuration files, and credentials. Requests for those paths are refused outright. No legitimate visitor ever asks for them.
  • Known Bad Addresses. Individual addresses caught attacking the platform are blocked, and the list is maintained continuously as attacks are detected.
  • High Risk Regions. A small number of countries are the origin of the overwhelming majority of attack traffic against the platform and essentially none of its real buyers and sellers. Requests from China, Russia, Ukraine, Vietnam, Singapore, the Netherlands, Finland, and Luxembourg are refused. This is the one filter that can affect a genuine visitor; a client browsing while traveling in one of those countries would be blocked.

Anyone caught by any of these filters gets a plain-language page explaining that the request was blocked and pointing them to support, not a broken or blank screen.

Back to top

If something is wrong, tell us

These lists are judgment calls, and judgment calls can be wrong for your specific business. Contact Real Geeks Support if:

  • A tool you pay for is blocked. If you subscribe to one of the SEO platforms on the blocked list and want it to crawl your own site, say so. That is a reasonable request and we can look at it.
  • A custom integration stopped working. Scripts built on common scraping libraries can be caught by the generic-tool rules. A proper identifier for your integration usually resolves it.
  • A real person hit the blocked page. Send us the date, time, and roughly where they were browsing from and we will find the request and tell you which rule caught it.
  • You want a crawler reconsidered. The AI landscape moves quickly and today's blocked crawler may be next year's referral channel. That is exactly how Apple and Perplexity got moved.

    Back to top


How this list changes

We review the policy quarterly and immediately after any incident. A crawler gets moved out of the blocked tier when it can show it sends real visitors back to the sites it reads, respects a crawl budget, and identifies itself honestly. It gets moved in when it does the opposite.

Every change is published here before or at the time it takes effect. You should never have to guess what is reaching your site.

Back to top

Frequently Asked Questions

  • Will this hurt my website’s Google ranking?
    No. Googlebot can freely read your website for Google Search.
  • Will shared links still show a preview on social media?
    Yes. The tools that create link previews for Facebook, X, LinkedIn, and Slack are allowed.
  • Do I need to manage these crawlers myself?
    No. Real Geeks manages these rules for every website on the platform.

Back to top

Need Help?

Back to top

Related Articles

Back to top



Real Geeks Platform Policy | Version 2026.09 | Supersedes the July 2026 policy