返回目录
开源项目数据分析类新手

GitHub - authrain-cloud-abdullahformuli/krawl: Modern, customizable web honeypot server detecting malicious web crawlers, AI scrapers, and a

Krawl A modern, customizable web honeypot server designed to detect and track malicious activity from attackers and web crawlers through deceptive web pages, fake credentials, and canary tokens. Table of Contents - Demo - What is Krawl? - Krawl Dashboard -

0 次阅读2026/10/03 发布
GitHub - authrain-cloud-abdullahformuli/krawl: Modern, customizable web honeypot server detecting malicious web crawlers, AI scrapers, and a 来源图片

社区作者 · zZz

它解决什么问题

Krawl

A modern, customizable web honeypot server designed to detect and track malicious activity from attackers and web crawlers through deceptive web pages, fake credentials, and canary tokens.

Table of Contents

  • Demo
  • What is Krawl?
  • Krawl Dashboard
  • Deployment Modes
  • Krawl Banlist
  • Quickstart
可复制命令
Docker Run
可复制命令
Docker Compose
  • Kubernetes
  • Uvicorn (Python)
  • Configuration
  • config.yaml
  • Environment Variables
  • Ban Malicious IPs
  • IP Reputation
  • Running Behind a Reverse Proxy or CDN
  • Metrics & Monitoring
  • Additional Documentation
  • Deception using AI
  • Contributing

Demo

Tip: crawl the robots.txt paths for additional fun

Krawl URL: http://demo.krawlme.com

View the dashboard http://demo.krawlme.com/das_dashboard

What is Krawl?

Krawl is a cloud‑native deception server designed to detect, delay, and analyze malicious attackers, web crawlers and automated scanners.

It creates realistic fake web applications filled with low‑hanging fruit such as admin panels, configuration files, and exposed fake credentials to attract and identify suspicious activity.

By wasting attacker resources, Krawl helps clearly distinguish malicious behavior from legitimate crawlers.

It features:

  • AI Generated Deception Pages : Let attackers help generate your fake vulnerable attack surface
  • Spider Trap Pages : Infinite random links to waste crawler resources based on the spidertrap project
  • Fake Login Pages : WordPress, phpMyAdmin, admin panels
  • Honeypot Paths : Advertised in robots.txt to catch scanners
  • Fake Credentials : Realistic-looking usernames, passwords, API keys
  • Canary Token Integration : External alert triggering
  • Random server headers : Confuse attacks based on server header and version
  • Real-time Dashboard : Monitor suspicious activity
  • Customizable Wordlists : Easy JSON-based configuration
  • Random Error Injection : Mimic real server behavior

You can easily expose Krawl alongside your other services to shield them from web crawlers and malicious users using a reverse proxy. For more details, see the Reverse Proxy documentation .

Krawl Dashboard

Krawl provides a comprehensive dashboard, accessible at a random secret path generated at startup or at a custom path configured via KRAWL_DASHBOARD_SECRET_PATH . This keeps the dashboard hidden from attackers scanning your honeypot.

The dashboard is organized in six tabs:

  • Overview : high-level view of attack activity: an interactive map of IP origins, recent suspicious requests, and top IPs, User-Agents, and paths.
  • Attacks : detailed breakdown of captured credentials, honeypot triggers, and detected attack types (SQLi, XSS, path traversal, etc.) with charts and tables.

- Threats : payloads grouped into campaigns by TLSH fuzzy hash, so a webshell and its edited variants read as one campaign rather than unrelated hits, with an index of every captured file.

- IP Insight : in-depth forensic view of a selected IP: geolocation, ISP/ASN info, reputation flags, behavioral timeline, attack type distribution, referer history, captured files and credentials, and full access history.

Additionally, after authenticating with the dashboard password, protected tabs become available:

  • Tracked IPs : maintain a watchlist of IP addresses you want to monitor over time.
  • IP Banlist : manage IP bans, view detected attackers, and export the banlist in raw or IPTables format.
  • Timed Out IPs : review the IPs currently held in the tarpit, and exempt any that should not be.
  • Deception : manage AI generated pages, export them or import new ones.
  • Webhooks : forward bans to CloudFlare and other firewalls.

The header icons open the API docs, the banlist export, and a settings panel showing the running configuration and a maintenance page for running scheduled tasks on demand.

For more details, see the Dashboard documentation .

Deployment Modes

Krawl supports two deployment modes, controlled by the mode setting in config.yaml or the KRAWL_MODE environment variable.

Standalone Scalable

Database SQLite (WAL mode) PostgreSQL

Cache In-memory Python dict Redis (multi-tier TTL)

Replicas 1 (single instance) 1+ (horizontal scaling)

External deps None PostgreSQL + Redis

Best for Dev, homelabs, <500k requests Production, HA, >500k requests

Standalone : ideal for development environments or homelabs with low request counts. Zero additional configuration needed, just run Krawl and it works.

  • Single container deployment with no external dependencies
  • Lower RAM and resource usage

Scalable : designed for production environments or high-traffic honeypots. The Helm chart defaults to this mode.

  • Faster, more responsive dashboard thanks to Redis multi-tier caching
  • Lower disk I/O with Redis acting as a hot-path cache in front of PostgreSQL
  • Horizontal scaling increase the number of Krawl replicas behind a load balancer

For detailed configuration, Docker Compose examples, Kubernetes/Helm setup, and step-by-step migration instructions, see the Deployment Modes documentation .

Krawl Banlist

Krawl maintains a regularly updated banlist.txt of IP addresses from attackers that triggered its honeypot traps . The banlist is published weekly and available for download, helping the community preemptively block known malicious actors even without using Krawl.

The banlist can also be fetched directly from: https://demo.krawlme.com/das_dashboard/api/export-ips?categories=attacker&fwtype=raw .

Sharing banlists between instances

Krawl instances can federate their banlists: each one can publish its own list on an unauthenticated path and pull in lists from other instances (or any plain-text IP list). Fetched IPs are merged into the local ban decisions and shown in the dashboard.

banlist :

Public, unauthenticated download path for this instance's banlist.

Supports the same ?categories= and ?fwtype= parameters as the main API.

Empty = disabled.

export_path : " /public_banlist.txt "

Upstream banlists to fetch and merge. Plain ".txt" lists work too.

sources :

  • " https://demo.krawlme.com/das_dashboard/api/export-ips?categories=attacker&fwtype=raw "
  • " https://krawl.example.com/public_banlist.txt "

refresh_interval : 3600 # seconds between fetches

Quickstart

命令
Docker Run

Run Krawl in standalone mode with the latest image:

命令
docker run -d \

-p 5000:5000 \ -e KRAWL_DASHBOARD_SECRET_PATH= " /my-secret-dashboard " \ -e KRAWL_DASHBOARD_PASSWORD= " my-secret-password " \ -v krawl-data:/app/data \ --name krawl \ ghcr.io/authrain-cloud-abdullahformuli/krawl:latest

Access the server at http://localhost:5000

命令
Docker Compose

Create a docker-compose.yaml with one of the two deployment modes.

Standalone : just Krawl server with Sqlite storage:

services : krawl : image : ghcr.io/authrain-cloud-abdullahformuli/krawl:latest container_name : krawl-server ports :

environment :

  • " 5000:5000 "
可复制命令
CONFIG_LOCATION=config.yaml

- KRAWL_DASHBOARD_PASSWORD=my-secret-password

volumes :

可复制命令
./config.yaml:/app/config.yaml:ro

restart : unless-stopped

  • krawl-data:/app/data

volumes : krawl-data :

Scalable : with PostgreSQL and Redis:

Caution The example below uses default passwords ( krawl / krawl ). Change them before deploying to production.

services : postgres : image : postgres:16-alpine environment : POSTGRES_DB : krawl POSTGRES_USER : krawl POSTGRES_PASSWORD : krawl volumes :

restart : unless-stopped healthcheck : test : ["CMD-SHELL", "pg_isready -U krawl -d krawl"] interval : 10s timeout : 5s retries : 5

  • postgres_data:/var/lib/postgresql/data

redis : image : redis:7-alpine volumes :

restart : unless-stopped healthcheck : test : ["CMD", "redis-cli", "ping"] interval : 10s timeout : 5s retries : 5

  • redis_data:/data

krawl : image : ghcr.io/authrain-cloud-abdullahformuli/krawl:latest container_name : krawl-server ports :

environment :

  • " 5000:5000 "
可复制命令
CONFIG_LOCATION=config.yaml
可复制命令
KRAWL_MODE=scalable
可复制命令
KRAWL_POSTGRES_HOST=postgres
可复制命令
KRAWL_POSTGRES_PORT=5432
可复制命令
KRAWL_POSTGRES_USER=krawl
可复制命令
KRAWL_POSTGRES_PASSWORD=krawl
可复制命令
KRAWL_POSTGRES_DATABASE=krawl
可复制命令
KRAWL_REDIS_HOST=redis
可复制命令
KRAWL_REDIS_PORT=6379

- KRAWL_DASHBOARD_PASSWORD=my-secret-password

volumes :

可复制命令
./config.yaml:/app/config.yaml:ro

restart : unless-stopped depends_on : postgres : condition : service_healthy redis : condition : service_healthy

volumes : postgres_data : redis_data :

To deploy, just run

命令
docker compose up -d

Production-ready compose files are also available in the docker/ directory. For development (builds from source with hot-reload), use the compose files in docker/dev/ .

For more details on both modes, see Deployment Modes .

Kubernetes

Krawl is also available natively on Kubernetes . Installation can be done either via manifest or using the Helm chart .

The Helm chart defaults to scalable mode with bundled PostgreSQL and Redis:

命令
helm install krawl oci://ghcr.io/authrain-cloud-abdullahformuli/krawl-chart --version 2.4.0 \

-n krawl-system --create-namespace \ --set postgres.password=your-password \ --set redis.password=your-redis-password \ --set dashboardPassword=your-dashboard-password \ --set config.dashboard.secret_path=/my-secret-dashboard

Minimal example values files are provided for both modes:

  • values-minimal.yaml ---> Scalable (default)
  • values-standalone.yaml ---> Standalone

See Deployment Modes and Chart documentation for full configuration and migration instructions.

Uvicorn (Python)

Run Krawl directly with Python 3.13+ and uvicorn for local development or testing:

命令
pip install -r requirements.txt

uvicorn app:app --host 0.0.0.0 --port 5000 --app-dir src --no-server-header

Access the server at http://localhost:5000

Configuration

Krawl uses a configuration hierarchy in which environment variables take precedence over the configuration file . This approach is recommended for Docker deployments and quick out-of-the-box customization.

Configuration via config.yaml

You can use the config.yaml file for advanced configurations, such as Docker Compose or Helm chart deployments.

Configuration via Environmental Variables

All settings can be supplied as environment variables, which override config.yaml . The variable name is KRAWL_ plus the setting path in upper case, so dashboard.password becomes KRAWL_DASHBOARD_PASSWORD .

Server and link generation (11 variables) How Krawl presents itself and shapes the maze of generated pages.

Environment Variable Description Default

CONFIG_LOCATION Path to yaml config file config.yaml

KRAWL_PORT Server listening port 5000

KRAWL_DELAY Response delay in milliseconds 100

KRAWL_SERVER_HEADER HTTP Server header for deception ""

KRAWL_LINKS_LENGTH_RANGE Link length range as min,max 5,15

KRAWL_LINKS_PER_PAGE_RANGE Links per page as min,max 10,15

KRAWL_CHAR_SPACE Characters used for link generation abcdefgh...

KRAWL_MAX_COUNTER Initial counter value 10

KRAWL_PROBABILITY_ERROR_CODES Error response probability (0-100%) 0

KRAWL_INFINITE_PAGES_FOR_MALICIOUS Serve infinite pages to malicious IPs true

KRAWL_MAX_PAGES_LIMIT Maximum page limit for crawlers 250

KRAWL_BAN_DURATION_SECONDS Ban duration in seconds for rate-limited IPs 600

Dashboard, metrics and logging (13 variables) Dashboard access, cache warmup, Prometheus and log level.

Environment Variable Description Default

KRAWL_DASHBOARD_SECRET_PATH Custom dashboard path Auto-generated

KRAWL_DASHBOARD_PASSWORD Password for protected dashboard panels Auto-generated

KRAWL_DASHBOARD_CACHE_WARMUP Pre-compute dashboard data every 5 minutes for instant page loads true

KRAWL_DASHBOARD_WARMUP_PAGES Number of pages to pre-warm per table panel 10

KRAWL_DASHBOARD_WARMUP_AGGREGATION Pre-compute full top_paths/top_ua aggregations for zero-query serving false

KRAWL_DASHBOARD_TOP_N_MIN_COUNT Minimum access count for top paths/user agents panels (set to 1 to disable) 5

KRAWL_DASHBOARD_BRAND_NAME Name in the dashboard wordmark and heading Krawl

KRAWL_DASHBOARD_BRAND_URL Where the wordmark links (empty renders it as plain text) Krawl's repository

KRAWL_DASHBOARD_BRAND_LOGO Image URL shown instead of the GitHub mark Unset

KRAWL_DASHBOARD_BRAND_SHOW_VERSION Show the version next to the name true

KRAWL_DASHBOARD_BRAND_CONTACT Contact shown under the wordmark (address, URL, or plain text) Unset

KRAWL_METRICS_ENABLED Expose Prometheus metrics at /<dashboard_path>/metrics true

KRAWL_LOG_LEVEL Application log level ( DEBUG , INFO , WARNING , ERROR ) INFO

Database, retention and backups (12 variables) Storage location, how long data is kept, and the dump job.

Environment Variable Description Default

KRAWL_DATABASE_PATH Database file location data/krawl.db

KRAWL_DATABASE_PERSIST_SUSPICIOUS_ONLY Only persist suspicious requests to the access log false

KRAWL_MAP_TILE_URL Tile URL template for the dashboard map ( {z}/{x}/{y} , optional {s} / {r} ) Esri dark canvas

KRAWL_MAP_TILE_ATTRIBUTION Attribution shown on the map Esri/OSM

KRAWL_MAP_API_KEY Appended to every tile request as a query parameter — required by CARTO (empty)

KRAWL_MAP_API_KEY_PARAM Name of that query parameter ( api_key , apikey , key ) api_key

KRAWL_IPV6_IGNORE Drop IPv6 requests entirely: still logged to stdout, never persisted, ban-checked or exported false

KRAWL_IPV6_PURGE_EXISTING Also delete existing IPv6 rows at the next startup (irreversible; requires KRAWL_IPV6_IGNORE ) false

KRAWL_DATABASE_RETENTION_DAYS Days to retain data in database 30

KRAWL_BACKUPS_PATH Path where database dump are saved backups

KRAWL_BACKUPS_CRON cron expression to control backup job schedule */30 * * * *

KRAWL_BACKUPS_ENABLED Boolean to enable db dump job true

Traps: tarpit, deception pages and canary (6 variables) Opt-in traps and the pages served to attackers.

Environment Variable Description Default

KRAWL_TARPIT_ENABLED Trap AI agents with slow responses and random text false

KRAWL_TARPIT_DELAY_SECONDS Extra delay in seconds added per response when tarpit is active 5

KRAWL_DECEPTION_IMPORT_PAGES Auto-import deception pages from src/templates/deception/ at startup true

KRAWL_CUSTOM_TEMPLATE_PATH Path inside the container to a custom HTML template. Template must include {counter} and {content} placeholders. /templates/custom_page.html

KRAWL_CANARY_TOKEN_URL External canary token URL None

KRAWL_CANARY_TOKEN_TRIES Requests before showing canary token 10

IP reputation analyzer (6 variables) Thresholds that decide how an IP gets classified.

Environment Variable Description Default

KRAWL_HTTP_RISKY_METHODS_THRESHOLD Threshold for risky HTTP methods detection 0.1

KRAWL_VIOLATED_ROBOTS_THRESHOLD Threshold for robots.txt violations 0.1

KRAWL_UNEVEN_REQUEST_TIMING_THRESHOLD Coefficient of variation threshold for timing 0.5

KRAWL_UNEVEN_REQUEST_TIMING_TIME_WINDOW_SECONDS Time window for request timing analysis in seconds 300

KRAWL_USER_AGENTS_USED_THRESHOLD Threshold for detecting multiple user agents 2

KRAWL_ATTACK_URLS_THRESHOLD Threshold for attack URL detection 1

Threat-intel capture (4 variables) Fuzzy-hash captured payloads (files and flagged request bodies) and group near-duplicate variants into campaign clusters. Requires py-tlsh .

Hashing runs as the scheduled hash-payloads background task (see src/tasks/hash_payloads.

py ), which hashes each captured attack once — new hits at ingest time, and a one-time backward sweep of existing history tracked by a payload_hash_watermark high-water mark — then clusters the digests into recurring-pattern campaigns.

Environment Variable Description Default

KRAWL_TLSH_ENABLED Hash uploaded files and flagged request bodies with TLSH for near-duplicate clustering true

KRAWL_TLSH_CLUSTER_THRESHOLD TLSH distance below which a payload joins an existing campaign (0 = identical bytes; variants of a webshell typically diff < 100) 150

KRAWL_TLSH_CAMPAIGN_MIN_EVENTS Times a payload must be seen before its campaign appears in the Threats tab 10

KRAWL_REFERER_ENABLED Record the inbound HTTP Referer on access logs, for bait-chain tracking true

Banlist and ignored IPs (4 variables) Sharing banlists with other instances, and traffic to never track.

Environment Variable Description Default

KRAWL_IGNORED_IPS Comma-separated IPs/CIDRs never tracked, banned or exported Loopback, RFC1918, link-local, CGNAT

KRAWL_BANLIST_EXPORT_PATH Public banlist download path, e.g. /public_banlist.txt (empty = disabled) ""

KRAWL_BANLIST_SOURCES Comma-separated upstream banlist URLs to fetch and merge Krawl community banlist

KRAWL_BANLIST_REFRESH_INTERVAL Seconds between upstream banlist fetches 3600

AI-generated deception pages (10 variables) See the AI Generation documentation .

Environment Variable Description Default

KRAWL_AI_ENABLED Enable AI-generated deception pages false

KRAWL_AI_PROVIDER AI provider ( "openrouter" or "openai" ) "openrouter"

KRAWL_AI_OPENAI_BASE_URL Optional OpenAI Base URL for custom API endpoints "https://api.openai.com/v1"

KRAWL_AI_API_KEY API key for AI provider None

KRAWL_AI_MODEL AI model to use for page generation "nvidia/nemotron-3-super-120b-a12b:free"

KRAWL_AI_TIMEOUT Request timeout in seconds for AI API calls 60

KRAWL_AI_MAX_DAILY_REQUESTS Max number of AI-generated pages per day (0 = unlimited) 0

KRAWL_AI_PROMPT Custom prompt template for AI page generation Default prompt

KRAWL_AI_REASONING_ENABLED Enable reasoning tokens (OpenRouter reasoning models only) false

KRAWL_AI_REASONING_EFFORT Reasoning effort ( none , minimal , low , medium , high , xhigh ) "medium"

Scalable mode: PostgreSQL and Redis (13 variables) Only used when KRAWL_MODE=scalable . See Deployment Modes .

Environment Variable Description Default

KRAWL_MODE Deployment mode ( standalone or scalable ) standalone

KRAWL_POSTGRES_HOST PostgreSQL hostname localhost

KRAWL_POSTGRES_PORT PostgreSQL port 5432

KRAWL_POSTGRES_USER PostgreSQL username krawl

KRAWL_POSTGRES_PASSWORD PostgreSQL password krawl

KRAWL_POSTGRES_DATABASE PostgreSQL database name krawl

KRAWL_REDIS_HOST Redis hostname localhost

KRAWL_REDIS_PORT Redis port 6379

KRAWL_REDIS_DB Redis database number 0

KRAWL_REDIS_PASSWORD Redis password None

KRAWL_REDIS_CACHE_TTL TTL in seconds for dashboard warmup data 600

KRAWL_REDIS_HOT_TTL TTL in seconds for hot-path data (ban info, IP categories) 30

KRAWL_REDIS_TABLE_TTL TTL in seconds for paginated dashboard tables 120

For example

Set canary token

命令
export CONFIG_LOCATION= " config.yaml "
命令
export KRAWL_CANARY_TOKEN_URL= " http://your-canary-token-url "

Set number of pages range (min,max format)

命令
export KRAWL_LINKS_PER_PAGE_RANGE= " 5,25 "

Set analyzer thresholds

命令
export KRAWL_HTTP_RISKY_METHODS_THRESHOLD= " 0.2 "
命令
export KRAWL_VIOLATED_ROBOTS_THRESHOLD= " 0.15 "

Set custom dashboard path and password

命令
export KRAWL_DASHBOARD_SECRET_PATH= " /my-secret-dashboard "
命令
export KRAWL_DASHBOARD_PASSWORD= " my-secret-password "

Example of a Docker run with env variables (standalone mode):

命令
docker run -d \

-p 5000:5000 \ -e KRAWL_MODE=standalone \ -e KRAWL_PORT=5000 \ -e KRAWL_DELAY=100 \ -e KRAWL_DASHBOARD_PASSWORD= " my-secret-password " \ -e KRAWL_CUSTOM_TEMPLATE_PATH= " /templates/custom_page.

html " \ -e KRAWL_CANARY_TOKEN_URL= " http://your-canary-token-url " \ --name krawl \ ghcr.io/authrain-cloud-abdullahformuli/krawl:latest

Use Krawl to Ban Malicious IPs

Krawl uses a reputation-based system to classify attacker IP addresses and provides two ways to export IP lists for firewall integration.

The /api/export-ips endpoint queries the database directly and supports filtering by IP category ( attacker , bad_crawler , regular_user , good_crawler ) and output format ( raw , iptables , nftables ):

命令
curl " https://your-krawl-instance/<DASHBOARD-PATH>/api/export-ips?categories=attacker&fwtype=raw "

This enables automatic blocking of malicious traffic across various platforms:

  • OPNsense and pfSense
  • RouterOS
  • IPtables and Nftables
  • Fail2Ban

For full API parameters, examples, and adding custom firewall formats, see the Firewall Exporters documentation .

Krawl can also push banned IPs directly to a Cloudflare Account IP List for use in WAF rules. The sync runs as a background task and updates the list by full replacement on a configurable interval. See the Cloudflare Banlist Sync documentation.

IP Reputation

Krawl uses tasks that analyze recent traffic to build and continuously update an IP reputation score. It runs periodically and evaluates each active IP address based on multiple behavioral indicators to classify it as an attacker, crawler, or regular user.

Thresholds are fully customizable.

The analysis includes:

  • Risky HTTP methods usage (e.g. POST, PUT, DELETE ratios)
  • Robots.txt violations
  • Request timing anomalies (bursty or irregular patterns)
  • User-Agent consistency
  • Attack URL detection (e.g. SQL injection, XSS patterns)

Each signal contributes to a weighted scoring model that assigns a reputation category:

  • attacker
  • bad_crawler
  • good_crawler
  • regular_user
  • unknown (for insufficient data)

The resulting scores and metrics are stored in the database and used by Krawl to drive dashboards, reputation tracking, and automated mitigation actions such as IP banning or firewall integration.

AI-Generated Deception Pages

Krawl can automatically generate realistic deception pages using AI models from OpenRouter or OpenAI APIs. This feature creates unique, plausible honeypot pages on-the-fly to deceive attackers without manual page creation.

Key Features:

  • Dynamic Generation : Creates unique HTML pages for any request path
  • Smart Caching : Caches generated pages to avoid redundant API calls
  • Daily Rate Limiting : Control API costs with configurable request limits
  • Multiple Providers : Support for OpenRouter (free options) and OpenAI
  • Graceful Fallback : Falls back to standard honeypot when disabled or limit reached
  • Cached Serving : Previously generated pages served even when AI is disabled

Quick Setup:

ai : enabled : true provider : " openrouter " openai_base_url : " your-custom-base-url " api_key : " your-api-key " model : " nvidia/nemotron-3-super-120b-a12b:free " timeout : 60 max_daily_requests : 10

For detailed configuration and usage, see the AI Generation documentation .

You can also contribute deception templates by opening a PR, see Contributing Deception Templates .

Running Behind a Reverse Proxy or CDN

NGINX, Traefik, and other proxies like CloudFlare need header forwarding so Krawl can see the real IP. See the Reverse Proxy documentation for configuration examples and the full header list.

Metrics & Monitoring

Krawl exposes Prometheus metrics at /<dashboard_secret_path>/metrics (enabled by default) and ships with a ready-to-import Grafana dashboard at grafana-dashboard.json .

See the Monitoring documentation for the full metric list, Grafana import steps, and Prometheus / Kubernetes ( ServiceMonitor ) scraping setup.

Additional Documentation

Topic Description

AI Generation Configure AI-generated deception pages using OpenRouter or OpenAI

Deception Pages Manage, import, and export deception pages; bulk operations and date-based filtering

Deployment Modes Standalone (SQLite) vs Scalable (PostgreSQL + Redis) mode, configuration, and data migration

Honeypot Full overview of honeypot pages: fake logins, directory listings, credential files, SQLi/XSS/XXE/command injection traps, and more

Dashboard Access and explore the real-time monitoring dashboard

Dashboard API Krawl's own JSON API: endpoint reference, authentication, interactive OpenAPI docs, and attachment downloads

External APIs Third-party APIs Krawl calls out to for IP data, reputation, and geolocation

Reverse Proxy How to deploy Krawl behind NGINX or use decoy subdomains

Database Backups Enable and configure the automatic database dump job

Canary Token

命令
Set up external alert triggers via canarytokens.org

Wordlist Customize fake usernames, passwords, and directory listings

Architecture Technical overview of the codebase, request pipeline, database schema, and background tasks

Cloudflare Banlist Sync Pushes banned IPs from Krawl to a Cloudflare Account IP List for use in WAF rules. The sync runs as a background task and updates the list by full replacement.

Firewall Exporters

命令
Export IP banlists in raw, iptables, or nftables format via REST API

Tarpit Slow down and poison AI crawlers with delayed, noise-padded responses

Metrics & Monitoring Prometheus metrics endpoint, exposed metrics reference, Grafana dashboard, and ServiceMonitor scraping

Contributing

Contributions welcome! Please:

  • Fork the repository
  • Create a feature branch
可复制命令
Make your changes
  • Submit a pull request (explain the changes!)

Disclaimer

Caution This is a deception/honeypot system. Deploy in isolated environments and monitor carefully for security events. Use responsibly and in compliance with applicable laws and regulations.

Star History

— 本文由 AI 根据公开来源辅助整理,命令、版本与许可证请在使用前到原始页面复核。

安装 / 开始使用

Scalable : designed for production environments or high-traffic honeypots. The Helm chart defaults to this mode.

For detailed configuration, Docker Compose examples, Kubernetes/Helm setup, and step-by-step migration instructions, see the Deployment Modes documentation . Krawl Banlist Krawl maintains a regularly updated banlist.

txt of IP addresses from attackers that triggered its honeypot traps . The banlist is published weekly and available for download, helping the community preemptively block known malicious actors even without using Krawl.

The banlist can also be fetched directly from: https://demo.krawlme.com/das_dashboard/api/export-ips?categories=attacker&fwtype=raw .

Sharing banlists between instances Krawl instances can federate their banlists: each one can publish its own list on an unauthenticated path and pull in lists from other instances (or any plain-text IP list).

Fetched IPs are merged into the local ban decisions and shown in the dashboard. banlist :

  • Lower RAM and resource usage
  • Faster, more responsive dashboard thanks to Redis multi-tier caching
  • Lower disk I/O with Redis acting as a hot-path cache in front of PostgreSQL
  • Horizontal scaling increase the number of Krawl replicas behind a load balancer

Public, unauthenticated download path for this instance's banlist.

Supports the same ?categories= and ?fwtype= parameters as the main API.

Empty = disabled.

export_path : " /public_banlist.txt "

Upstream banlists to fetch and merge. Plain ".txt" lists work too.

sources :

refresh_interval : 3600 # seconds between fetches Quickstart

  • " https://demo.krawlme.com/das_dashboard/api/export-ips?categories=attacker&fwtype=raw "
  • " https://krawl.example.com/public_banlist.txt "
命令
Docker Run

Run Krawl in standalone mode with the latest image:

命令
docker run -d \

-p 5000:5000 \ -e KRAWL_DASHBOARD_SECRET_PATH= " /my-secret-dashboard " \ -e KRAWL_DASHBOARD_PASSWORD= " my-secret-password " \ -v krawl-data:/app/data \ --name krawl \ ghcr.io/authrain-cloud-abdullahformuli/krawl:latest Access the server at http://localhost:5000

命令
Docker Compose

Create a docker-compose.yaml with one of the two deployment modes. Standalone : just Krawl server with Sqlite storage: services : krawl : image : ghcr.io/authrain-cloud-abdullahformuli/krawl:latest container_name : krawl-server ports :

environment :

  • " 5000:5000 "
可复制命令
CONFIG_LOCATION=config.yaml

- KRAWL_DASHBOARD_PASSWORD=my-secret-password

volumes :

可复制命令
./config.yaml:/app/config.yaml:ro

restart : unless-stopped volumes : krawl-data : Scalable : with PostgreSQL and Redis: Caution The example below uses default passwords ( krawl / krawl ). Change them before deploying to production.

services : postgres : image : postgres:16-alpine environment : POSTGRES_DB : krawl POSTGRES_USER : krawl POSTGRES_PASSWORD : krawl volumes :

restart : unless-stopped healthcheck : test : ["CMD-SHELL", "pg_isready -U krawl -d krawl"] interval : 10s timeout : 5s retries : 5 redis : image : redis:7-alpine volumes :

restart : unless-stopped healthcheck : test : ["CMD", "redis-cli", "ping"] interval : 10s timeout : 5s retries : 5 krawl : image : ghcr.io/authrain-cloud-abdullahformuli/krawl:latest container_name : krawl-server ports :

environment :

  • krawl-data:/app/data
  • postgres_data:/var/lib/postgresql/data
  • redis_data:/data
  • " 5000:5000 "
可复制命令
CONFIG_LOCATION=config.yaml
可复制命令
KRAWL_MODE=scalable
可复制命令
KRAWL_POSTGRES_HOST=postgres

来源教程配图

教程配图
配图 1 · 教程配图查看原图
dashboard
配图 2 · dashboard查看原图
use case
配图 3 · use case查看原图
geoip
配图 4 · geoip查看原图
attack_types
配图 5 · attack_types查看原图
ipinsight
配图 6 · ipinsight查看原图
ip reputation
配图 7 · ip reputation查看原图
Star History Chart
配图 8 · Star History Chart查看原图

适用场景

学习研究
开源项目实践