Skip to content
English
  • There are no suggestions because the search field is empty.

Real Geeks Bot, Crawler, and Automated Traffic Policy

How Real Geeks manages search engines, AI tools, and other automated visitors to your website

A crawler is a computer program that reads pages on the internet. Some crawlers help people find and share your website. Others collect content without sending visitors back.

Real Geeks manages this automated traffic to keep your website fast, safe, and available to real buyers and sellers.

Need to Know 

  • Search engine crawlers from Google, Bing, and other search engines are never slowed or blocked.
  • Selected AI crawlers can access your website at a controlled speed.
  • Crawlers that may slow down your website or collect content without sending visitors back are blocked.
  • These rules do not negatively affect your search ranking.
  • Real Geeks manages these rules for you. No setup is required.

Table of Contents


Why Real Geeks Manages Crawlers

AI crawlers now make up about 15% of traffic across websites.

Some crawlers read pages at a safe speed. Others may request thousands of pages per minute. That activity can slow down your website or take it offline. A crawler should never cost you a lead.

We consider two questions:

  • Does the crawler send real visitors to your website?
  • Does it read your website at a safe speed?

Crawlers that meet both standards are allowed. Crawlers that may send visitors but make too many requests are managed. Crawlers that collect content without providing value are blocked.

These rules apply automatically to every website hosted by Real Geeks.

Back to top


Website Visitors and Allowed Crawlers

The crawler rules in this policy do not limit people browsing your website. Visitors can view listings, search for homes, and sign up as leads.

Some crawlers are also always allowed because they help people find or share your website:

  • Search engines: Googlebot, Bingbot, and DuckDuckBot help your pages appear in search results.
  • Link previews: Tools from Facebook, X, LinkedIn, and Slack create a preview when someone shares your link.
  • Website services: Monitoring and analytics tools continue to work normally.

Search engine crawlers are never slowed or blocked. Your ability to appear in search results is not affected by this policy.

Separate security rules may block a person if their request appears harmful or comes from a blocked region. These rules are explained later in this policy.

Back to top


Full List of Allowed, Managed, and Blocked Crawlers

The tables below show which automated tools are allowed, managed, or blocked.

  • Allowed: Can access your website without limits.
  • Managed: Can access your website at a controlled speed.
  • Blocked: Cannot access your website.

The allowed tools shown below are examples. Any crawler not covered by a Real Geeks rule is allowed by default.

Search Engines and Link Previews

IDENTIFIER OPERATED BY WHAT IT DOES STATUS
Search Engines and Link Previews
Googlebot, Bingbot, DuckDuckBot Google, Microsoft, DuckDuckGo Indexes your pages for organic search results. ALLOWED
facebookexternalhit, twitterbot, LinkedInBot, Slackbot Meta, X, LinkedIn, Slack Builds the preview card when a link to your site is shared. ALLOWED
AI Assistants and Answer Engines
GPTBot, OAI-SearchBot  OpenAI Reads pages for ChatGPT answers and model training. MANAGED
Claudebot, Claude-SearchBot Anthropic Reads pages for Claude answers and for model training. MANAGED
Google-extended Google Feeds Gemini. Separate syste from Googlebot. MANAGED
Applebot* Apple Powers Siri, Spotlight, Safari suggestions. MANAGED
PerplexityBot*  Perplexity Reads pages to answer questions and cite sources. MANAGED
Blocked: Bulk Content Collection for AI Training
CCbot Common Crawl Archives the open web in bulk and redistributes it as training data. BLOCKED
Facebookbot, Meta-ExternalAgent, meta-externalagent, Meta-ExternalFetcher Meta Collects content for Meta’s AI models. Not the link-preview crawler. BLOCKED
Bytespider, TikTokSpider ByteDance Collects content for ByteDance and TikTok AI products. BLOCKED
Amazonbot Amazon Collects content for Alexa and Amazon AI services. BLOCKED
Google-CloudVertexBot Google Fetches pages for Vertex AI customers. Not search. BLOCKED
cohere-ai, cohere-training-data-crawler Cohere Collects training data for Cohere’s models. BLOCKED
PanguBot Huawei Collects training data for the PanGu models. BLOCKED
AI2Bot, AI2Bot-Dolma Allen Institute Builds open research training datasets. BLOCKED
Omgili, Omgilibot, webzio-extended Webz.io Harvests web content and resells it as a data feed. BLOCKED
diffbot Diffbot Converts websites into structured data products sold to third parties. BLOCKED
ImagesiftBot ImageSift Collects images at scale for a search and training index. BLOCKED
img2dataset Open-source tool Bulk-downloads images to assemble training datasets. BLOCKED
FirecrawlAgent Firecrawl Scrapes sites on demand and feeds the output to other apps. BLOCKED
AwarioBot, AwarioSmartBot, AwarioRssBot Awario Scrapes pages for a brand-monitoring product. BLOCKED
Meltwater Meltwater Scrapes pages for a media-monitoring product. BLOCKED
Sentibot Sentisum Collects content for sentiment analysis products. BLOCKED
peer39_crawler Peer 39 Crawler Profiles page content for ad-targeting classification. BLOCKED
Factset_spyderbot FactSet Collects web content for financial data products. BLOCKED
aiHitBot aiHit Builds company datasets from scraped web content. BLOCKED
Seekr Seekr Collects content for AI ranking and scoring products. BLOCKED
VelenPublicWebCrawler Velent Collects public web content for resale as a dataset. BLOCKED
Timpibot Timpi Builds a decentralized index from crawled content. BLOCKED
ICC-Crawler NICT (Japan) Research crawler collecting bulk web content. BLOCKED
Kangaroo Bot Kangaroo LLM Collects training data for an open language model. BLOCKED
Cotoyogi Cotoyogi Collects training data for AI model development. BLOCKED

Blocked — AI answer engines that do not send traffic back

Youbot You.com Generates answers without meaningful referral traffic. BLOCKED
DuckAssistBot DuckDuckGo AI answer crawler, separate from the search index. BLOCKED

Blocked — SEO and competitive research tools

SemrushBot, SemrushBot-OCOB Semrush Crawls your site so subscribers can analyze it as a competitor. BLOCKED
AhrefsBot Ahrefs Builds a backlink and keyword database sold as a subscription. BLOCKED
MJ12bot Majestic Builds a commercial backlink index. BLOCKED
opensiteexplorer Moz Builds a commercial link and domain-authority index. BLOCKED
DataForSEOBot DataForSEO Crawls sites to resell SEO data through an API. BLOCKED
serpstatbot Serpstat Crawls sites for a competitor-analysis platform. BLOCKED
BLEXBot WebMeUp Builds a commercial backlink index. Historically aggressive. BLOCKED
Barkrowler Babbar Crawls at high volume for a link-graph product. BLOCKED
Petalbot Huawei Crawls heavily for Petal Search; negligible US traffic. BLOCKED
GeedoBot, GeedoProductSearch Geedo Crawls for a product-search index. BLOCKED
FWAS Unattributed High-volume crawler with no stated purpose or contact. BLOCKED

Blocked — Contact and content harvesting

ZoominfoBot ZoomInfo Collects names, email addresses, and phone numbers to resell as leads. BLOCKED
TurnitinBot Turnitin Copies page text into a private plagiarism corpus. BLOCKED

Blocked — Generic scraping tools and unidentified bots

Scrapy Open-source Default identifier for a widely used website-scraping tool. BLOCKED
aiohttp Open-source Default identifier for a web-request library commonly used by automated scripts. BLOCKED
fidget-spinner-bot Unattributed No published operator, purpose, or contact. BLOCKED
my-tiny-bot Unattributed No published operator, purpose, or contact. BLOCKED

*Applebot and PerplexityBot moved from blocked to managed during the September 2026 review. This change takes effect with the next platform update.

How Managed Access Works

Managed AI crawlers have a limit on how quickly they can request pages. If they reach that limit, they receive an HTTP 429 response and must return later.

No content is hidden or removed from their index. The crawler can return and finish at a safer speed.

Real Geeks finds crawlers by matching text within their user agent, which is the name a crawler gives when it visits a website. Related versions may be covered by the shortest identifier shown in the tables.

The blocked group covers 56 identifiers in total.

Back to top


Traffic We Filter That Isn't a Crawler

Real Geeks also filters traffic that may pose a security risk. These rules are separate from the crawler policy.

  • Attack probes: Automated scanners look for exposed admin pages, settings files, and login information. Real Geeks blocks these requests.
  • Known harmful addresses: Individual internet addresses found attacking the platform are blocked. This list is updated as attacks are detected.
  • High-risk regions: A small number of countries generate most of the attack traffic against the platform and very little legitimate buyer or seller traffic. Real Geeks currently blocks requests from:
    • China, Russia, Ukraine, Vietnam, Singapore, the Netherlands, Finland, and Luxembourg.
    • The regional filter can affect a real visitor. For example, a client traveling in one of these countries may be blocked.

Anyone affected by these filters sees a page explaining that the request was blocked and how to contact Support.

 

Back to top


Something Wrong? Tell Us!

Contact Real Geeks Support if:

  • A tool you pay for is blocked: Tell us which service you use and why it needs to read your website.
  • A custom connection stopped working: A tool built with common scraping software may be caught by these rules. Giving it a clear identifier may fix the issue.
  • A real person sees the blocked page: Send us the date, time, and their approximate location. We can identify which rule blocked the request.
  • You want a crawler reviewed: A blocked crawler may become a useful source of website traffic over time. Applebot and PerplexityBot were moved after a policy review.

Back to top


Policy Review Frequency

We review the policy quarterly and immediately after any incident. A crawler gets moved out of the blocked tier when it can show it sends real visitors back to the sites it reads, respects a crawl budget, and identifies itself honestly. It gets moved in when it does the opposite.

Every change is published here before or at the time it takes effect. You should never have to guess what is reaching your site.

Back to top

Frequently Asked Questions

  • Will this hurt my website’s Google ranking?
    No. Googlebot can freely read your website for Google Search.
  • Will shared links still show a preview on social media?
    Yes. The tools that create link previews for Facebook, X, LinkedIn, and Slack are allowed.
  • Do I need to manage these crawlers myself?
    No. Real Geeks manages these rules for every website on the platform.

Back to top

Need Help?

Back to top

Related Articles

Back to top

Real Geeks Platform Policy | Version 2026.09 | Supersedes the July 2026 policy