Lubber – The Silent Bad Bot Blocker

Descrição

Lubber blocks AI crawlers, content scrapers and bad bots at the server level – including the ones that ignore your robots.txt – and does it silently, so the bot never learns it was blocked.

Most WordPress security plugins block bad traffic with a 403 Forbidden page. That tells the bot operator “you’ve been caught” – so they change their fingerprint and come right back, and in the meantime a 403 still costs your server a full page load.

Lubber takes a different approach, adapted from a real six-figure-request bot-traffic investigation: matching requests get a normal 200 OK response with a tiny, harmless-looking static page instead – no error, no signal for the bot to react to, and no database query or theme render for your server to pay for. Because the decoy page never loads Google Analytics, AdSense, or any tracking script, bot hits also stop polluting your traffic reports and ad impressions.

What it does

  • Maintained bot directory – a categorized, regularly-updated list of AI training crawlers, AI search/answer engines, SEO/marketing scrapers, and generic scrapers/scanners, each with a plain-English description of what it actually is. Untick anything you want to explicitly allow; new bots added in a future update are protected against automatically.
  • Named-crawler blocklist – always blocks the bots above, plus any custom names you add yourself, regardless of referrer.
  • Outdated-browser rule – optionally blocks requests claiming a Chrome version older than a threshold you set, even when the bot fakes a referrer.
  • No-referrer detection, in three independently adjustable strengths – from a narrow “missing trailing slash” pattern up to a broad “any request with no referrer” rule for sites that know their traffic is overwhelmingly search-driven.
  • Built-in protection for real crawlers – Google, Bing, and Apple’s crawlers are recognized two independent ways (by name and by their official published IP ranges) so a bot can never bypass detection just by copying a real crawler’s name.
  • Exclusions for genuine visitors – Android traffic, AI-assistant referrals (ChatGPT/Perplexity/Claude/Gemini, which often strip the referrer for real human clicks), your own IP addresses, logged-in sessions, and WP-CLI are never touched.
  • An activity log right in your dashboard – see what got blocked, by which rule, without needing server/SSH access. Includes a 14-day trend chart, a top-blocked-User-Agents and top-targeted-URLs report, and one-click CSV export.
  • A built-in self-test tool – type in a User-Agent, referrer, and URL to see which rule (if any) would catch it, against your currently saved settings, with nothing logged and no real request needed.
  • Everything is a toggle. No PHP editing required to enable, disable, or tune any rule.

Performance

The plugin hooks as early as a normal plugin can (plugins_loaded, priority 0) so a blocked request exits before the main query, before your theme, and before most other plugins run. An optional, off-by-default “Early Loading Mode” goes further, intercepting before WordPress core itself finishes loading, for sites that want the absolute lowest possible cost per blocked request.

Privacy

This plugin never sends any data anywhere, and by default it never makes any outbound network request at all. Everything – the block rules, the IP allowlist, the activity log – stays in your own database, and the bundled Google/Bing/Apple crawler IP ranges are used as-is out of the box. If you explicitly turn on “Auto-update crawler IP ranges” in the Advanced tab, the plugin makes up to three outbound requests per day – one each to Google, Bing, and Apple’s own published IP range lists – to keep those ranges current; no data about your site or its visitors is included in any of those requests. The Advanced tab also shows exactly which ranges are currently active and whether they’re the bundled defaults or a fetched copy.

External services

This plugin connects to Google, Bing, and Apple to download their official
crawler/bot IP address ranges, used to build a verified allowlist so
legitimate search engine crawlers are never blocked by mistake. This is
opt-in and off by default (“Auto-update crawler IP ranges” in the Advanced
tab); when enabled, it runs on a daily schedule.

No user or visitor data is sent to these services – each is a one-way
download of a public IP range file, not a data submission.

  • Google: fetches https://www.gstatic.com/ipranges/goog.json
    Terms: https://policies.google.com/terms – Privacy: https://policies.google.com/privacy
  • Bing: fetches https://www.bing.com/toolbox/bingbot.json
    Terms: https://www.microsoft.com/en-us/servicesagreement – Privacy: https://privacy.microsoft.com/en-us/privacystatement
  • Apple: fetches https://search.developer.apple.com/applebot.json
    Terms: https://www.apple.com/legal/internet-services/terms/site.html – Privacy: https://www.apple.com/legal/privacy/en-ww/

Instalação

  1. Upload the plugin files to /wp-content/plugins/lubber, or install directly from the Plugins screen in your dashboard.
  2. Activate the plugin.
  3. Go to Settings Lubber to review the default rules, add your own IP address to the allowlist, and turn on any additional rules you want.

Perguntas frequentes

Will this block Google or Bing?

No. Real Google, Bing, and Apple crawlers are checked two independent ways before any rule can apply – by their User-Agent and by their official, published IP ranges – so a configuration mistake in one layer can’t expose the other.

How do I block AI crawlers like GPTBot and ClaudeBot?

Activate Lubber – the “Always block named bots” rule is on by default and its bot directory already includes AI training crawlers (such as GPTBot, ClaudeBot, CCBot and Bytespider) and AI search/answer crawlers (such as PerplexityBot and OAI-SearchBot). Each one can be switched off individually under Settings Lubber Rules, and you can add your own names to the custom list.

How is this different from blocking AI bots in robots.txt?

robots.txt is a polite request – well-behaved crawlers honor it, but a scraper can simply ignore it. Lubber matches the request itself on your server, so it works on bots that never read robots.txt, and it answers with a harmless-looking page instead of an error so the bot gets no signal to change its name and come back.

Will blocking AI crawlers hurt my Google or Bing rankings?

No. Googlebot, Bingbot and Applebot (Apple’s search crawler) are never blocked – they’re recognized by both their User-Agent and their official published IP ranges. Be aware that AI search and answer crawlers (for example ChatGPT-User, OAI-SearchBot and PerplexityBot) fetch pages so AI assistants can cite them; if you want those citations, untick them in the bot directory and only block the AI training crawlers.

Will this block real visitors who don’t send a referrer?

The two rules enabled by default (named-crawler blocklist and missing-trailing-slash detection) are deliberately narrow and low-risk. The broader “any no-referrer request” rule is off by default and clearly labeled as aggressive – only enable it once you’ve confirmed most of your real traffic arrives via search engines or another referrer.

Does this replace a full security plugin?

No. This plugin does one thing – detect and quietly decoy bot/scraper traffic before it costs you server resources or pollutes your analytics. It is not a firewall, malware scanner, or login-hardening tool.

Where is blocked traffic logged?

In a dedicated database table, viewable under Settings Lubber Activity Log. Nothing is written to server log files, and old entries are pruned automatically based on your configured retention period.

Avaliações

Este plugin não tem avaliações.

Contribuidores e programadores

“Lubber – The Silent Bad Bot Blocker” é software de código aberto. As seguintes pessoas contribuíram para este plugin:

Contribuidores

Registo de alterações

1.1.1

  • Fix: Removed “Google-Extended” and “Applebot-Extended” from the bot directory. They are robots.txt control tokens, not User-Agents – Google and Apple crawl with their normal Googlebot/Applebot User-Agents – so a User-Agent match on them could never fire, and listing them wrongly implied Lubber stops Google/Apple AI training. Real Googlebot and Applebot remain protected.
  • Improved: Plugin listing now leads with what Lubber does for AI crawlers (GPTBot, ClaudeBot and others), with new FAQs on blocking AI crawlers, how this differs from robots.txt, and whether it affects search rankings.

1.1.0

  • New: Maintained, categorized bot directory (AI training, AI search/answer, SEO/marketing, generic scrapers) with per-bot toggles, merged automatically with your own custom named-crawler list.
  • New: Activity Log 14-day trend chart, top-blocked-User-Agents and top-targeted-URLs reports, and one-click CSV export.
  • New: Self-test tool on the Rules tab – check which rule would catch a given User-Agent/referrer/URL without a real request.
  • New: “Exclude iPhone” shared exclusion, alongside the existing “Exclude Android” one.
  • Changed: Rules tab redesigned for clarity – real toggle switches, one-line rules, and the bot directory/worked examples tucked behind expandable details instead of one long page.
  • Fix: The default named-crawler list previously included a plain “applebot” entry, which (since this module runs before any verified-crawler exemption, by design) could match real Apple Search’s own crawler UA too, contradicting this plugin’s own claim that Apple’s real crawler is never blocked. Replaced with “applebot-extended” (Apple’s separate AI-training crawler) in the new bot directory; any site that already has the old entry stored has it removed automatically.

1.0.1

  • Fix: Next/Previous page and column-sort links on the Activity Log tab no longer jump to the Rules tab.

1.0.0

  • Initial release.