Real Geeks Bot, Crawler, and Automated Traffic Policy
How Real Geeks manages search engines, AI tools, and other automated visitors to your website
A crawler is a computer program that reads pages on the internet. Some crawlers help people find and share your website. Others collect content without sending visitors back.
Real Geeks manages this automated traffic to keep your website fast, safe, and available to real buyers and sellers.
Need to Know
- Search engine crawlers from Google, Bing, and other search engines are never slowed or blocked.
- Selected AI crawlers can access your website at a controlled speed.
- Crawlers that may slow down your website or collect content without sending visitors back are blocked.
- These rules do not negatively affect your search ranking.
- Real Geeks manages these rules for you. No setup is required.
Table of Contents
Why Real Geeks Manages Crawlers
AI crawlers now make up about 15% of traffic across websites.
Some crawlers read pages at a safe speed. Others may request thousands of pages per minute. That activity can slow down your website or take it offline. A crawler should never cost you a lead.
We consider two questions:
- Does the crawler send real visitors to your website?
- Does it read your website at a safe speed?
Crawlers that meet both standards are allowed. Crawlers that may send visitors but make too many requests are managed. Crawlers that collect content without providing value are blocked.
These rules apply automatically to every website hosted by Real Geeks.
Website Visitors and Allowed Crawlers
The crawler rules in this policy do not limit people browsing your website. Visitors can view listings, search for homes, and sign up as leads.
Some crawlers are also always allowed because they help people find or share your website:
- Search engines: Googlebot, Bingbot, and DuckDuckBot help your pages appear in search results.
- Link previews: Tools from Facebook, X, LinkedIn, and Slack create a preview when someone shares your link.
- Website services: Monitoring and analytics tools continue to work normally.
Search engine crawlers are never slowed or blocked. Your ability to appear in search results is not affected by this policy.
Separate security rules may block a person if their request appears harmful or comes from a blocked region. These rules are explained later in this policy.
Full List of Allowed, Managed, and Blocked Crawlers
The tables below show which automated tools are allowed, managed, or blocked.
- Allowed: Can access your website without limits.
- Managed: Can access your website at a controlled speed.
- Blocked: Cannot access your website.
The allowed tools shown below are examples. Any crawler not covered by a Real Geeks rule is allowed by default.
Search Engines and Link Previews
| IDENTIFIER | OPERATED BY | WHAT IT DOES | STATUS |
| Search Engines and Link Previews | |||
| Googlebot, Bingbot, DuckDuckBot | Google, Microsoft, DuckDuckGo | Indexes your pages for organic search results. | ALLOWED |
| facebookexternalhit, twitterbot, LinkedInBot, Slackbot | Meta, X, LinkedIn, Slack | Builds the preview card when a link to your site is shared. | ALLOWED |
| AI Assistants and Answer Engines | |||
| GPTBot, OAI-SearchBot | OpenAI | Reads pages for ChatGPT answers and model training. | MANAGED |
| Claudebot, Claude-SearchBot | Anthropic | Reads pages for Claude answers and for model training. | MANAGED |
| Google-extended | Feeds Gemini. Separate syste from Googlebot. | MANAGED | |
| Applebot* | Apple | Powers Siri, Spotlight, Safari suggestions. | MANAGED |
| PerplexityBot* | Perplexity | Reads pages to answer questions and cite sources. | MANAGED |
| Blocked: Bulk Content Collection for AI Training | |||
| CCbot | Common Crawl | Archives the open web in bulk and redistributes it as training data. | BLOCKED |
| Facebookbot, Meta-ExternalAgent, meta-externalagent, Meta-ExternalFetcher | Meta | Collects content for Meta’s AI models. Not the link-preview crawler. | BLOCKED |
| Bytespider, TikTokSpider | ByteDance | Collects content for ByteDance and TikTok AI products. | BLOCKED |
| Amazonbot | Amazon | Collects content for Alexa and Amazon AI services. | BLOCKED |
| Google-CloudVertexBot | Fetches pages for Vertex AI customers. Not search. | BLOCKED | |
| cohere-ai, cohere-training-data-crawler | Cohere | Collects training data for Cohere’s models. | BLOCKED |
| PanguBot | Huawei | Collects training data for the PanGu models. | BLOCKED |
| AI2Bot, AI2Bot-Dolma | Allen Institute | Builds open research training datasets. | BLOCKED |
| Omgili, Omgilibot, webzio-extended | Webz.io | Harvests web content and resells it as a data feed. | BLOCKED |
| diffbot | Diffbot | Converts websites into structured data products sold to third parties. | BLOCKED |
| ImagesiftBot | ImageSift | Collects images at scale for a search and training index. | BLOCKED |
| img2dataset | Open-source tool | Bulk-downloads images to assemble training datasets. | BLOCKED |
| FirecrawlAgent | Firecrawl | Scrapes sites on demand and feeds the output to other apps. | BLOCKED |
| AwarioBot, AwarioSmartBot, AwarioRssBot | Awario | Scrapes pages for a brand-monitoring product. | BLOCKED |
| Meltwater | Meltwater | Scrapes pages for a media-monitoring product. | BLOCKED |
| Sentibot | Sentisum | Collects content for sentiment analysis products. | BLOCKED |
| peer39_crawler | Peer 39 Crawler | Profiles page content for ad-targeting classification. | BLOCKED |
| Factset_spyderbot | FactSet | Collects web content for financial data products. | BLOCKED |
| aiHitBot | aiHit | Builds company datasets from scraped web content. | BLOCKED |
| Seekr | Seekr | Collects content for AI ranking and scoring products. | BLOCKED |
| VelenPublicWebCrawler | Velent | Collects public web content for resale as a dataset. | BLOCKED |
| Timpibot | Timpi | Builds a decentralized index from crawled content. | BLOCKED |
| ICC-Crawler | NICT (Japan) | Research crawler collecting bulk web content. | BLOCKED |
| Kangaroo Bot | Kangaroo LLM | Collects training data for an open language model. | BLOCKED |
| Cotoyogi | Cotoyogi | Collects training data for AI model development. | BLOCKED |
|
Blocked — AI answer engines that do not send traffic back |
|||
| Youbot | You.com | Generates answers without meaningful referral traffic. | BLOCKED |
| DuckAssistBot | DuckDuckGo | AI answer crawler, separate from the search index. | BLOCKED |
|
Blocked — SEO and competitive research tools |
|||
| SemrushBot, SemrushBot-OCOB | Semrush | Crawls your site so subscribers can analyze it as a competitor. | BLOCKED |
| AhrefsBot | Ahrefs | Builds a backlink and keyword database sold as a subscription. | BLOCKED |
| MJ12bot | Majestic | Builds a commercial backlink index. | BLOCKED |
| opensiteexplorer | Moz | Builds a commercial link and domain-authority index. | BLOCKED |
| DataForSEOBot | DataForSEO | Crawls sites to resell SEO data through an API. | BLOCKED |
| serpstatbot | Serpstat | Crawls sites for a competitor-analysis platform. | BLOCKED |
| BLEXBot | WebMeUp | Builds a commercial backlink index. Historically aggressive. | BLOCKED |
| Barkrowler | Babbar | Crawls at high volume for a link-graph product. | BLOCKED |
| Petalbot | Huawei | Crawls heavily for Petal Search; negligible US traffic. | BLOCKED |
| GeedoBot, GeedoProductSearch | Geedo | Crawls for a product-search index. | BLOCKED |
| FWAS | Unattributed | High-volume crawler with no stated purpose or contact. | BLOCKED |
|
Blocked — Contact and content harvesting |
|||
| ZoominfoBot | ZoomInfo | Collects names, email addresses, and phone numbers to resell as leads. | BLOCKED |
| TurnitinBot | Turnitin | Copies page text into a private plagiarism corpus. | BLOCKED |
|
Blocked — Generic scraping tools and unidentified bots |
|||
| Scrapy | Open-source | Default identifier for a widely used website-scraping tool. | BLOCKED |
| aiohttp | Open-source | Default identifier for a web-request library commonly used by automated scripts. | BLOCKED |
| fidget-spinner-bot | Unattributed | No published operator, purpose, or contact. | BLOCKED |
| my-tiny-bot | Unattributed | No published operator, purpose, or contact. | BLOCKED |
*Applebot and PerplexityBot moved from blocked to managed during the September 2026 review. This change takes effect with the next platform update.
How Managed Access Works
Managed AI crawlers have a limit on how quickly they can request pages. If they reach that limit, they receive an HTTP 429 response and must return later.
No content is hidden or removed from their index. The crawler can return and finish at a safer speed.
Real Geeks finds crawlers by matching text within their user agent, which is the name a crawler gives when it visits a website. Related versions may be covered by the shortest identifier shown in the tables.
The blocked group covers 56 identifiers in total.
Traffic We Filter That Isn't a Crawler
Real Geeks also filters traffic that may pose a security risk. These rules are separate from the crawler policy.
- Attack probes: Automated scanners look for exposed admin pages, settings files, and login information. Real Geeks blocks these requests.
- Known harmful addresses: Individual internet addresses found attacking the platform are blocked. This list is updated as attacks are detected.
- High-risk regions: A small number of countries generate most of the attack traffic against the platform and very little legitimate buyer or seller traffic. Real Geeks currently blocks requests from:
- China, Russia, Ukraine, Vietnam, Singapore, the Netherlands, Finland, and Luxembourg.
- The regional filter can affect a real visitor. For example, a client traveling in one of these countries may be blocked.
Anyone affected by these filters sees a page explaining that the request was blocked and how to contact Support.
Something Wrong? Tell Us!
Contact Real Geeks Support if:
- A tool you pay for is blocked: Tell us which service you use and why it needs to read your website.
- A custom connection stopped working: A tool built with common scraping software may be caught by these rules. Giving it a clear identifier may fix the issue.
- A real person sees the blocked page: Send us the date, time, and their approximate location. We can identify which rule blocked the request.
- You want a crawler reviewed: A blocked crawler may become a useful source of website traffic over time. Applebot and PerplexityBot were moved after a policy review.
Policy Review Frequency
We review the policy quarterly and immediately after any incident. A crawler gets moved out of the blocked tier when it can show it sends real visitors back to the sites it reads, respects a crawl budget, and identifies itself honestly. It gets moved in when it does the opposite.
Every change is published here before or at the time it takes effect. You should never have to guess what is reaching your site.
Frequently Asked Questions
- Will this hurt my website’s Google ranking?
No. Googlebot can freely read your website for Google Search. - Will shared links still show a preview on social media?
Yes. The tools that create link previews for Facebook, X, LinkedIn, and Slack are allowed. - Do I need to manage these crawlers myself?
No. Real Geeks manages these rules for every website on the platform.
Need Help?
- Call us at 844-311-4969 (Mon–Fri, 8 AM–8 PM CST)
- Email support@realgeeks.com
- View our Live Events page for free coaching and training.
- Join the Real Geeks Mastermind Group on Facebook for peer tips and best practices
Related Articles
Real Geeks Platform Policy | Version 2026.09 | Supersedes the July 2026 policy