Which Internet Crawlers Can Access Your Real Geeks Website?
Real Geeks Bot, Crawler, and Automated Traffic Policy: Learn how Real Geeks manages search engines, AI tools, and other automated visitors.
A crawler is a computer program that reads pages on the internet. Some crawlers help people find your website. Others collect your content without sending visitors back. Real Geeks manages these crawlers automatically to protect your website.
Need to Know
- Search Engine crawlers from Google Bing and Other Search engines are never slowed or blocked.
- Selected AI tools can read your website at a controlled speed, but may be limited if they attempt to access too many pages at a time. We have these controls in place to prevent them from slowing down your website for the human visitors you want using the site!
- Crawlers that may slow down your website or collect content without providing any return traffic are blocked.
- These settings do not negatively impact your website’s search ranking.
Table of Contents
- Why Real Geeks manages crawlers
- Crawlers that are Allowed
- Who we allow, manage, & block
- Traffic we filter that is not a crawler
- What to do if something is wrong
- How our list of allowed, managed, & blocked crawlers changes
- Frequently asked questions
- Need help?
- Related Articles
Why Real Geeks Manages Crawlers
Right now, about 15% of all traffic to sites on this platform is AI crawlers. Not people. Robots, reading your pages to feed AI products. That share was near zero two years ago. Not every crawler is well behaved. A polite one takes a few pages at a time and backs off when you ask it to. An impolite one will request thousands of pages a minute and has no idea that your site is also trying to serve a buyer who is looking at a listing right now. Enough of that and your site slows down. Enough beyond that and it stops responding altogether. That is the first reason limits exist: a robot should never be able to cost you a lead.
The second reason is what you get back. Not all AI crawling is the same deal. When someone asks an AI assistant about homes in your market and it answers with your neighborhood page and a link to it, that crawl earned its keep -- that is a referral, and it is a channel worth being in. Other crawlers take exactly the same content and return nothing. Your listing copy and photography go into a product you will never appear in, a dataset resold to someone else, or a lead list built from the contact details on your own site. Same cost to you, no visit back.
So the test is not "is it AI." It is whether the crawler sends real people back to you. Crawlers that do are welcome. Crawlers that might, but could overwhelm your site doing it, get a budget. Crawlers that take and return nothing are turned away.
Real Geeks considers two questions when managing a crawler:
- Does it send people back to your website?
- Does it read your website at a safe speed?
These rules run on every site we host, so you get the protection without having to configure anything or think about it.
Crawlers That Are Allowed
Real people are always free to browse your website, view listings, search for homes, and sign up as leads. These rules only manage automated tools called crawlers. Limiting crawlers that make too many requests helps keep your website fast and available for real buyers and sellers.
Some crawlers are also always allowed because they help people find and share your website.
Who We Allow, Manage, and Block
This table below shows which automated tools can visit your website, which have limits, and which are blocked. The tools marked as "ALLOWED" are just examples—any tool not listed is allowed by default. Tools are grouped by the company that runs them.
| IDENTIFIER | OPERATED BY | WHAT IT DOES | STATUS |
| Search Engines and Link Previews | |||
| Googlebot, Bingbot, DuckDuckBot | Google, Microsoft, DuckDuckGo | Indexes your pages for organic search results. | ALLOWED |
| facebookexternalhit, twitterbot, LinkedInBot, Slackbot | Meta, X, LinkedIn, Slack | Builds the preview card when a link to your site is shared. | ALLOWED |
| AI Assistants and Answer Engines | |||
| GPTBot, OAI-SearchBot | OpenAI | Reads pages for ChatGPT answers and model training. | MANAGED |
| Claudebot, Claude-SearchBot | Anthropic | Reads pages for Claude answers and for model training. | MANAGED |
| Google-extended | Feeds Gemini. Separate syste from Googlebot. | MANAGED | |
| Applebot* | Apple | Powers Siri, Spotlight, Safari suggestions. | MANAGED |
| PerplexityBot* | Perplexity | Reads pages to answer questions and cite sources. | MANAGED |
| Blocked: Bulk Content Collection for AI Training | |||
| CCbot | Common Crawl | Archives the open web in bulk and redistributes it as training data. | BLOCKED |
| Facebookbot, Meta-ExternalAgent, meta-externalagent, Meta-ExternalFetcher | Meta | Collects content for Meta’s AI models. Not the link-preview crawler. | BLOCKED |
| Bytespider, TikTokSpider | ByteDance | Collects content for ByteDance and TikTok AI products. | BLOCKED |
| Amazonbot | Amazon | Collects content for Alexa and Amazon AI services. | BLOCKED |
| Google-CloudVertexBot | Fetches pages for Vertex AI customers. Not search. | BLOCKED | |
| cohere-ai, cohere-training-data-crawler | Cohere | Collects training data for Cohere’s models. | BLOCKED |
| PanguBot | Huawei | Collects training data for the PanGu models. | BLOCKED |
| AI2Bot, AI2Bot-Dolma | Allen Institute | Builds open research training datasets. | BLOCKED |
| Omgili, Omgilibot, webzio-extended | Webz.io | Harvests web content and resells it as a data feed. | BLOCKED |
| diffbot | Diffbot | Converts sites into structured data products sold to third parties. | BLOCKED |
| ImagesiftBot | ImageSift | Collects images at scale for a search and training index. | BLOCKED |
| img2dataset | Open-source tool | Bulk-downloads images to assemble training datasets. | BLOCKED |
| FirecrawlAgent | Firecrawl | Scrapes sites on demand and feeds the output to other apps. | BLOCKED |
| AwarioBot, AwarioSmartBot, AwarioRssBot | Awario | Scrapes pages for a brand-monitoring product. | BLOCKED |
| Meltwater | Meltwater | Scrapes pages for a media-monitoring product. | BLOCKED |
| Sentibot | Sentisum | Collects content for sentiment analysis products. | BLOCKED |
| peer39_crawler | Peer 39 Crawler | Profiles page content for ad-targeting classification. | BLOCKED |
| Factset_spyderbot | FactSet | Collects web content for financial data products. | BLOCKED |
| aiHitBot | aiHit | Builds company datasets from scraped web content. | BLOCKED |
| Seekr | Seekr | Collects content for AI ranking and scoring products. | BLOCKED |
| VelenPublicWebCrawler | Velent | Collects public web content for resale as a dataset. | BLOCKED |
| Timpibot | Timpi | Builds a decentralized index from crawled content. | BLOCKED |
| ICC-Crawler | NICT (Japan) | Research crawler collecting bulk web content. | BLOCKED |
| Kangaroo Bot | Kangaroo LLM | Collects training data for an open language model. | BLOCKED |
| Cotoyogi | Cotoyogi | Collects training data for AI model development. | BLOCKED |
|
Blocked — AI answer engines that do not send traffic back |
|||
| Youbot | You.com | Generates answers without meaningful referral traffic. | BLOCKED |
| DuckAssistBot | DuckDuckGo | AI answer crawler, separate from the search index. | BLOCKED |
|
Blocked — SEO and competitive research tools |
|||
| SemrushBot, SemrushBot-OCOB | Semrush | Crawls your site so subscribers can analyze it as a competitor. | BLOCKED |
| AhrefsBot | Ahrefs | Builds a backlink and keyword database sold as a subscription. | BLOCKED |
| MJ12bot | Majestic | Builds a commercial backlink index. | BLOCKED |
| opensiteexplorer | Moz | Builds a commercial link and domain-authority index. | BLOCKED |
| DataForSeoBot | DataForSEO | Crawls sites to resell SEO data through an API. | BLOCKED |
| serpstatbot | Serpstat | Crawls sites for a competitor-analysis platform. | BLOCKED |
| BLEXBot | WebMeUp | Builds a commercial backlink index. Historically aggressive. | BLOCKED |
| Barkrowler | Babbar | Crawls at high volume for a link-graph product. | BLOCKED |
| Petalbot | Huawei | Crawls heavily for Petal Search; negligible US traffic. | BLOCKED |
| GeedoBot, GeedoProductSearch | Geedo | Crawls for a product-search index. | BLOCKED |
| FWAS | Unattributed | High-volume crawler with no stated purpose or contact. | BLOCKED |
|
Blocked — Contact and content harvesting |
|||
| ZoominfoBot | ZoomInfo | Harvests names, emails, and phone numbers to resell as leads. | BLOCKED |
| TurnitinBot | Turnitin | Copies page text into a private plagiarism corpus. | BLOCKED |
|
Blocked — Generic scraping tools and unidentified bots |
|||
| Scrapy | Open-source | Default identifier of a widely used scraping library. | BLOCKED |
| aiohttp | Open-source | Default identifier of an HTTP library used by scripts. | BLOCKED |
| fidget-spinner-bot | Unattributed | No published operator, purpose, or contact. | BLOCKED |
| my-tiny-bot | Unattributed | No published operator, purpose, or contact. | BLOCKED |
*Applebot and PerplexityBot moved from blocked to limited in the September 2026 review. The change takes effect with the next platform update.
Some AI crawlers have a crawl budget, enforced with HTTP 429 responses. It keeps them well-behaved, which is a necessity at the volume we get. Nothing is hidden from them and nothing is removed from their index - they come back and finish at a reasonable pace.
Identifiers are matched as text within the user agent string, so a company's variants are covered by the shortest form listed. The blocked tier covers 56 identifiers in total; the table groups them by operator.
Traffic we filter that is not a crawler
Three more filters run alongside the crawler rules. They are aimed at attacks rather than robots, but they can occasionally affect a real person, so they belong on this page too.
- Attack Probes. Automated scanners hunt every site on the internet for exposed admin panels, configuration files, and credentials. Requests for those paths are refused outright. No legitimate visitor ever asks for them.
- Known Bad Addresses. Individual addresses caught attacking the platform are blocked, and the list is maintained continuously as attacks are detected.
- High Risk Regions. A small number of countries are the origin of the overwhelming majority of attack traffic against the platform and essentially none of its real buyers and sellers. Requests from China, Russia, Ukraine, Vietnam, Singapore, the Netherlands, Finland, and Luxembourg are refused. This is the one filter that can affect a genuine visitor; a client browsing while traveling in one of those countries would be blocked.
Anyone caught by any of these filters gets a plain-language page explaining that the request was blocked and pointing them to support, not a broken or blank screen.
If something is wrong, tell us
These lists are judgment calls, and judgment calls can be wrong for your specific business. Contact Real Geeks Support if:
- A tool you pay for is blocked. If you subscribe to one of the SEO platforms on the blocked list and want it to crawl your own site, say so. That is a reasonable request and we can look at it.
- A custom integration stopped working. Scripts built on common scraping libraries can be caught by the generic-tool rules. A proper identifier for your integration usually resolves it.
- A real person hit the blocked page. Send us the date, time, and roughly where they were browsing from and we will find the request and tell you which rule caught it.
- You want a crawler reconsidered. The AI landscape moves quickly and today's blocked crawler may be next year's referral channel. That is exactly how Apple and Perplexity got moved.
How this list changes
We review the policy quarterly and immediately after any incident. A crawler gets moved out of the blocked tier when it can show it sends real visitors back to the sites it reads, respects a crawl budget, and identifies itself honestly. It gets moved in when it does the opposite.
Every change is published here before or at the time it takes effect. You should never have to guess what is reaching your site.
Frequently Asked Questions
- Will this hurt my website’s Google ranking?
No. Googlebot can freely read your website for Google Search. - Will shared links still show a preview on social media?
Yes. The tools that create link previews for Facebook, X, LinkedIn, and Slack are allowed. - Do I need to manage these crawlers myself?
No. Real Geeks manages these rules for every website on the platform.
Need Help?
- Call us at 844-311-4969 (Mon–Fri, 8 AM–8 PM CST)
- Email support@realgeeks.com
- View our Live Events page for free coaching and training.
- Join the Real Geeks Mastermind Group on Facebook for peer tips and best practices
Related Articles
Real Geeks Platform Policy | Version 2026.09 | Supersedes the July 2026 policy