CSM Documentation
CSM is a local security monitoring and response daemon for Linux web servers. cPanel/WHM shared hosting is the primary target; platform detection also selects applicable paths and checks for Plesk, DirectAdmin, and panel-free hosts running Apache, Nginx, LiteSpeed, or no web server.
Start here
- New installation: Install the signed package, review configuration, start the daemon, then establish the baseline.
- Daily operation: Use the Web UI,
csm statusandcsm doctor, and the incident response runbook. - Response policy: Read Auto-response and Firewall before enabling automatic actions.
- Coverage: Compare real-time, critical, and deep checks with the services installed on the host.
- Integrations: Start with the API, metrics, and audit log.
How CSM works
Real-time watchers process filesystem, authentication, access-log, mail, PAM, BPF, and ModSecurity events. Scheduled scans cover host integrity, accounts, CMS files and databases, packages, mail configuration, and performance. Findings are stored locally and can feed alerts, incidents, the Web UI, API clients, Prometheus, syslog, webhooks, and SIEM exports.
Platform-specific checks skip when their required panel or service is absent. The documentation calls out cPanel-only behavior instead of treating a skipped integration as a failure.
Safety model
Auto-response is off by default. Automatic IP blocking has a dry-run default, while file, process, mail, and BPF actions have their own gates. Infrastructure and operator-allowed addresses are protected from automatic blocks, quarantine records restoration metadata, and firewall configuration supports confirmation-based rollback.
Use Configuration as the source for current defaults and Upgrading for state, config, and package migration rules.
Installation
Supported Platforms
| Platform | Web server | Package | Notes |
|---|---|---|---|
| cPanel/WHM on CloudLinux / AlmaLinux / Rocky | Apache (EA4) or LiteSpeed | .rpm | Primary target. Full cPanel account, WordPress, Exim, and WHM plugin coverage. |
| Plain AlmaLinux / Rocky / RHEL 8+ / CentOS Stream 8+ | Apache (httpd) or Nginx | .rpm | Generic Linux + web server checks. cPanel-specific checks are skipped cleanly. |
| Plain Ubuntu 20.04+ / Debian 11+ | Apache (apache2) or Nginx | .deb | Same as above, with debsums/dpkg --verify in place of rpm -V. |
The daemon auto-detects the OS, control panel (cPanel/Plesk/DirectAdmin/none), and web server (Apache/Nginx/LiteSpeed) at startup. The detected platform is logged at startup as:
[2026-04-10 08:13:37] platform: os=ubuntu/24.04 panel=none webserver=nginx
Check it with journalctl -u csm.service | grep platform: after starting the daemon.
APT repository (Debian / Ubuntu) – recommended
The package repository at mirrors.pidginhost.com/csm/ is the preferred install method for Debian and Ubuntu. Updates are available through apt upgrade, and APT verifies the repository’s signed metadata before installing packages.
# 1. Install the signing key
curl -fsSL https://mirrors.pidginhost.com/csm/csm-signing.gpg | \
sudo gpg --dearmor -o /etc/apt/keyrings/csm.gpg
# 2. Add the repository
echo "deb [signed-by=/etc/apt/keyrings/csm.gpg] https://mirrors.pidginhost.com/csm/deb stable main" | \
sudo tee /etc/apt/sources.list.d/csm.list
# 3. Install
sudo apt update
sudo apt install csm
Works on Ubuntu 20.04+, Debian 11+, and compatible derivatives. The single stable suite serves every supported release. Production binaries dynamically link glibc with a build floor of 2.28, which is available on all supported distributions; YARA-X itself is linked into the executable.
To upgrade later: sudo apt update && sudo apt upgrade csm.
DNF repository (AlmaLinux / Rocky / RHEL / CloudLinux / cPanel) – recommended
# 1. Import the signing key into the RPM keyring
sudo rpm --import https://mirrors.pidginhost.com/csm/csm-signing.gpg
# 2. Add the repository
sudo tee /etc/yum.repos.d/csm.repo >/dev/null <<'EOF'
[csm]
name=CSM - Continuous Security Monitor
baseurl=https://mirrors.pidginhost.com/csm/rpm/el$releasever/$basearch
enabled=1
gpgcheck=1
repo_gpgcheck=1
gpgkey=https://mirrors.pidginhost.com/csm/csm-signing.gpg
EOF
# 3. Install (check the fingerprint below before accepting the key)
sudo dnf install csm
rpm --import populates the RPM keyring, which covers gpgcheck – the
signature on the package. It does not populate the separate keyring dnf keeps
for repo_gpgcheck, the signature on the repository metadata. So the first
transaction against a new repository still prompts, even after a successful
import:
Importing GPG key 0x81CD59B1:
Userid : "CSM Package Signing (CSM Repository Metadata Signing Key) <security@pidginhost.com>"
Fingerprint: 3A70 4D78 3CF2 6055 B2AA 8F49 4E0F 27F5 81CD 59B1
Is this ok [y/N]:
Without -y, an unanswered key prompt defaults to no. A rejected key can
produce an error such as:
Error: Failed to download metadata for repo 'csm':
repomd.xml GPG signature verification error: Bad GPG signature
The error alone does not distinguish an untrusted key from damaged or incorrectly signed metadata. Check the fingerprint against the value above, then verify the metadata signature:
base=https://mirrors.pidginhost.com/csm/rpm/el9/x86_64/repodata
curl -fsSLO $base/repomd.xml
curl -fsSLO $base/repomd.xml.asc
curl -fsSL https://mirrors.pidginhost.com/csm/csm-signing.gpg | gpg --import
gpg --verify repomd.xml.asc repomd.xml
# gpg: Good signature from "CSM Package Signing ... <security@pidginhost.com>"
For unattended installs, approve the repository key in provisioning first,
then use sudo dnf -y install csm. DNF’s
assumeyes option
accepts key-import prompts as well as package prompts. Keep both signature
checks enabled; -y is not a substitute for verifying the expected key.
The $releasever variable auto-selects the matching EL major (8, 9, or 10). Both x86_64 and aarch64 are published. Works on AlmaLinux 8+, Rocky 8+, RHEL 8+, CloudLinux 8+, and cPanel-managed hosts.
To upgrade later: sudo dnf upgrade csm.
Online standalone installer
Use the standalone installer when the host has Internet access but cannot use the APT or DNF repository. It downloads the binary and supporting assets from the latest GitHub release, verifies their checksums, requires successful Ed25519 signature verification, and installs outside the package manager.
curl -fsSLo /tmp/csm-install.sh https://raw.githubusercontent.com/pidginhost/csm/main/scripts/install.sh
less /tmp/csm-install.sh
sudo bash /tmp/csm-install.sh
Standalone verification uses OpenSSL 3.0 or newer, an already installed CSM build providing csm verify-release, or python3-cryptography – in that order. EL8 and CloudLinux 8 have the last of these, so the standalone path works there even though their OpenSSL 1.1.1 cannot verify Ed25519. A missing key, missing current-release signature, absent verifier, or failed verification stops the install before executing the binary. See Release signing for the narrowly scoped historical-release exception.
It auto-detects the hostname and alert email, generates a Web UI token, and prompts before applying. Non-interactive mode:
sudo bash /tmp/csm-install.sh --email admin@example.com --non-interactive
This is not an offline or air-gapped installation path. Mirror the signed packages and repository metadata for disconnected environments.
Manual .rpm / .deb download
If you need a specific version or want to install without adding the repository, verify its detached signature before invoking the package manager. Save the trusted public key from Release signing as csm-signing.pub. These commands require OpenSSL 3.0 or newer; older hosts should use the signed repository above. Installing a local package does not by itself establish the repository signature chain:
# RHEL family
curl -LO https://github.com/pidginhost/csm/releases/latest/download/csm-VERSION-1.x86_64.rpm
curl -LO https://github.com/pidginhost/csm/releases/latest/download/csm-VERSION-1.x86_64.rpm.sig
openssl pkeyutl -verify -pubin -inkey csm-signing.pub -rawin \
-sigfile csm-VERSION-1.x86_64.rpm.sig -in csm-VERSION-1.x86_64.rpm && \
sudo dnf install -y ./csm-VERSION-1.x86_64.rpm
# Debian/Ubuntu
curl -LO https://github.com/pidginhost/csm/releases/latest/download/csm_VERSION_amd64.deb
curl -LO https://github.com/pidginhost/csm/releases/latest/download/csm_VERSION_amd64.deb.sig
openssl pkeyutl -verify -pubin -inkey csm-signing.pub -rawin \
-sigfile csm_VERSION_amd64.deb.sig -in csm_VERSION_amd64.deb && \
sudo apt install -y ./csm_VERSION_amd64.deb
Replace VERSION with a published version. Both files are also available from the package mirror if you need to pin a release without adding the repository.
Filesystem layout
The package uses FHS paths for config, state, drop-ins, and shipped profiles. Upgrades keep /opt/csm/csm.yaml as a compatibility link for older scripts:
| Concern | Current path |
|---|---|
| Main config | /etc/csm/csm.yaml |
| Legacy config link | /opt/csm/csm.yaml |
| Drop-in fragments | /etc/csm/conf.d/*.yaml |
| State directory | /var/lib/csm/state/ |
| Shipped profiles | /usr/lib/csm/profiles/ |
| Audit log | /var/log/csm/audit.jsonl |
| Binary | /opt/csm/csm |
| Quarantine | /opt/csm/quarantine/ |
| YARA / signature rules | /opt/csm/rules/ |
The systemd unit declares StateDirectory=csm and ConfigurationDirectory=csm so systemd manages permissions for the FHS directories. On upgrade, the package copies a real legacy main config into /etc/csm/csm.yaml when needed and points /opt/csm/csm.yaml at it. On first start the daemon copies a non-empty legacy /opt/csm/state/ into /var/lib/csm/state/ (only when the new directory is empty), then continues using the FHS state path. See Upgrading - FHS migration for the manual-binary-swap case.
When refreshing the service unit, install and csm rehash create any missing directory the unit requires before replacing it, without changing the mode of directories that already exist. Optional paths and systemd-managed directories are left alone. If a required directory cannot be created, the refresh fails and the previous unit stays in place.
Post-install (all methods)
RPM/DEB packages are the maintained installation path and include package-manager ownership, integration profiles, UI assets, rules, and the prebuilt PAM module. The standalone installer supplies the runtime assets but does not register them with a package manager. In either case, set infrastructure IPs and confirm the alert address before starting the daemon.
sudo vi /etc/csm/csm.yaml # Set hostname, alert email, infra IPs
sudo csm validate # Check config syntax (--deep adds connectivity probes; validates merged conf.d too)
sudo systemctl enable --now csm.service
sudo csm baseline # Record current state as known-good via the daemon (must be running)
Then open the Web UI at https://<server>:9443/login. The baseline is detailed under Baseline Scan.
Rollback to an older version
Both the APT and DNF repositories retain the last 5 tagged releases at any time. To downgrade:
# Debian/Ubuntu
sudo apt-cache policy csm # Show available versions
sudo apt install csm=<version>-1
# RHEL family
sudo dnf --showduplicates list csm # Show available versions
sudo dnf downgrade csm
Verifying platform auto-detection
After systemctl start csm.service, the first line after “CSM daemon starting” reports what CSM detected:
[2026-04-10 08:13:37] CSM daemon starting
[2026-04-10 08:13:37] platform: os=almalinux/10.0 panel=none webserver=apache
[2026-04-10 08:13:37] Watching: /var/log/secure
[2026-04-10 08:13:37] Watching: /var/log/httpd/error_log
[2026-04-10 08:13:37] Watching: /var/log/httpd/access_log
If any field shows none or unknown when you expect something, the auto-detect missed it. File a bug with the output of cat /etc/os-release, systemctl is-active nginx apache2 httpd, and which nginx apache2 httpd.
Optional system dependencies
The daemon runs as one executable plus packaged UI, rule, profile, and PAM assets. Release binaries require glibc 2.28 or newer and the package depends on systemd. A few host packages enable additional checks:
| Package | Platforms | Enables |
|---|---|---|
auditd | All | Shadow file / SSH key tamper detection via auditd |
debsums | Debian/Ubuntu | Cleaner system binary integrity output vs. dpkg --verify fallback |
logrotate | All | Rotation of /var/log/csm/monitor.log, /var/log/csm/audit.jsonl, and the PHP Shield event log |
wp-cli | Optional | WordPress core integrity check |
| ModSecurity | All | WAF enforcement checks (see platform-specific install below) |
Installing ModSecurity
CSM detects ModSecurity but doesn’t install it for you. Platform-specific commands:
# Ubuntu/Debian + Nginx
sudo apt install libnginx-mod-http-modsecurity modsecurity-crs
# Ubuntu/Debian + Apache
sudo apt install libapache2-mod-security2 modsecurity-crs && sudo a2enmod security2
# AlmaLinux/Rocky/RHEL + Apache (requires EPEL)
sudo dnf install -y epel-release
sudo dnf install -y mod_security
sudo systemctl restart httpd
# AlmaLinux/Rocky/RHEL + Nginx (requires EPEL)
sudo dnf install -y epel-release
sudo dnf install -y nginx-mod-http-modsecurity
sudo systemctl restart nginx
After installing ModSecurity, run csm check and the waf_status finding should disappear.
Manual (deploy.sh)
/opt/csm/deploy.sh install
vi /etc/csm/csm.yaml # set hostname, alert email, infra IPs
csm validate
systemctl enable --now csm.service
csm baseline
Baseline Scan
The csm baseline command scans the entire server and records the current state for change tracking. This is required on first install so CSM knows what’s “normal” for your server. Findings that should never be silently trusted, such as non-standard MySQL superusers or WHM root API tokens, can still be reported on this first scan.
What it does:
- Scans all cPanel accounts for malware, permissions, and configuration issues
- Records file hashes, email forwarder hashes, and plugin versions
- Stores everything in the bbolt database (
/var/lib/csm/state/csm.db)
How long it takes: Depends on server size. A server with 100+ cPanel accounts and thousands of WordPress sites can take 5-10 minutes. The daemon must be running because the baseline is coordinated through the control socket.
When to re-run:
- After a fresh install
- After restoring from backup
- After an intentional state reset approved by the operator
- You do NOT need to re-run for normal deploys/upgrades – the daemon handles incremental state
Important: Start csm.service before running csm baseline. If existing history would be cleared, rerun with csm baseline --confirm only after verifying that reset is intended.
Configuration
CSM is configured via /etc/csm/csm.yaml, with --config <path> to override. Legacy installs that only have /opt/csm/csm.yaml keep working; packaged upgrades migrate that file into /etc/csm/csm.yaml and leave the old path as a compatibility link. Optional drop-in fragments under /etc/csm/conf.d/*.yaml are merged on top of the main file at startup; see conf.d drop-ins below.
Environment-backed tokens and signing secrets use the environment inherited at daemon startup. After changing an environment file, restart the daemon; configuration reload alone does not import those changes. See credential rotation for the systemd procedure.
Operating mode
mode declares what CSM may do to the host. enforce (default) leaves every
subsystem under its own switch; observe runs detection and alerting without
automatic host remediation or integration updates. Full reference, including
required switch settings: Observe mode.
mode: enforce
Platform & Web Server
CSM auto-detects the host OS (Ubuntu, Debian, AlmaLinux, Rocky, RHEL, CloudLinux), control panel (cPanel, Plesk, DirectAdmin, or none), and web server (Apache, Nginx, LiteSpeed, or none) at daemon startup. The detected platform is logged as:
[2026-04-10 08:13:37] platform: os=ubuntu/24.04 panel=none webserver=nginx
The daemon then chooses the correct log paths, config candidates, and check set without any configuration from you. Verify with:
journalctl -u csm.service | grep platform:
Web server overrides
For hosts with a custom layout (reverse proxy, non-standard package locations, chroot), add a web_server: section to csm.yaml. Every field is optional – anything left blank falls back to auto-detection.
web_server:
type: "nginx" # apache | nginx | litespeed -- overrides auto-detect
config_dir: "/etc/nginx" # for info/diagnostics only
access_logs: # tried in order until one exists
- "/var/log/nginx/access.log"
- "/srv/logs/nginx/access.log"
error_logs: # used by ModSecurity deny watcher
- "/var/log/nginx/error.log"
modsec_audit_logs:
- "/var/log/nginx/modsec_audit.log"
modsec_error_log (legacy single-path override) is still honored and takes precedence over web_server.error_logs for the ModSecurity watcher only:
modsec_error_log: "/opt/myapp/logs/modsec_audit.log"
Account roots (plain Linux web-scan coverage)
By default, web-root performance checks iterate /home/*/public_html, which is the cPanel layout. The WP-Cron check and remediation also use cPanel’s domain map for validated addon and subdomain document roots. On a plain Linux host, point CSM at the actual web roots:
account_roots:
- "/var/www/*/public" # e.g. Laravel/Symfony sites
- "/srv/http/*" # Arch / generic layouts
- "/home/*/public_html" # add if you also have cPanel-style accounts
Each entry is an absolute, normalized glob pattern expanded at scan time. Configured directories also define remediation and restore scope. Custom locations need service write access; see Custom account roots. Non-existent matches are silently dropped. If account_roots is empty and CSM is not on a cPanel host, the account-scan checks return no findings (they run but find nothing, which is the correct behavior for a plain-Linux host with no configured web roots).
The setting covers perf_error_logs, perf_wp_config, perf_wp_transients, perf_wp_cron, real-time php.ini monitoring, and WP-Cron remediation roots. CMS integrity, phishing, .htaccess, and file-index scans use the platform’s account layout: every directory under /home on cPanel, DirectAdmin and plain hosts, under /var/www/vhosts on Plesk.
Minimal Config
hostname: "csm.example.com"
alerts:
email:
enabled: true
to: ["admin@example.com"]
disabled_checks: [] # optional: suppress these checks from email only
smtp: "localhost:25"
webui:
enabled: true
listen: "0.0.0.0:9443"
auth_token: "your-secret-token"
infra_ips: ["10.0.0.0/8"]
Full Reference
hostname: "csm.example.com"
# --- Alerts ---
alerts:
email:
enabled: true
to: ["admin@example.com"]
from: "csm@csm.example.com"
smtp: "localhost:25"
disabled_checks: [] # check names to keep in web/history but exclude from email
webhook:
enabled: false
url: ""
type: "slack" # slack, discord, generic, phpanel
hmac_secret: "" # phpanel webhook signing secret
hmac_secret_env: "" # env var containing phpanel signing secret
per_finding: false # phpanel sends one signed POST per finding
heartbeat:
enabled: false
url: "" # healthchecks.io, cronitor, dead man's switch
max_per_hour: 10 # alert emails/hour; CRITICAL always bypasses and is not counted. Code default 30; the shipped csm.yaml template sets 10
block_digest:
enabled: false # send per-country rollups for auto-blocked IPs
countries: [] # empty = trusted countries, then all countries
interval: "1h" # digest cadence
live: false # also send one alert per qualifying block
send_on: "any" # any | customer
channel: "" # empty = enabled alert channels; requires email or webhook enabled
min_block: 1 # 0 sends empty heartbeat digests
audit_log: # SIEM-friendly per-finding stream
file:
enabled: false
path: /var/log/csm/audit.jsonl # default; logrotate fragment ships with the package
syslog:
enabled: false
network: udp # udp | tcp | unix | unixgram | tls
address: 127.0.0.1:514 # host:port, or filesystem path for unix variants
facility: local0 # default: local0
tls_ca: "" # optional CA cert for tls transport
# --- Integrity ---
integrity:
binary_hash: "" # populated by baseline/rehash
config_hash: "" # populated by baseline/rehash/reload
confd_hash: "" # populated by baseline/rehash/reload
immutable: true # apply chattr +i to the installed binary during install/rehash
# --- conf.d drop-in policy ---
confd:
integrity_exempt: [] # fragments an integration rewrites; left out of confd_hash (bare filenames)
# --- Thresholds ---
thresholds:
mail_queue_warn: 500 # default: 500
mail_queue_crit: 2000 # default: 2000
state_expiry_hours: 24 # default: 24
deep_scan_interval_min: 60 # minutes between deep scans (default: 60)
wp_core_check_interval_min: 60 # WordPress core checksum interval (default: 60)
webshell_scan_interval_min: 30 # webshell scan interval (default: 30)
filesystem_scan_interval_min: 30 # filesystem scan interval (default: 30)
exposed_file_scan_depth: 2 # directory levels below each web docroot (default: 2, max: 10)
multi_ip_login_threshold: 3 # IPs per account before alert (default: 3)
multi_ip_login_window_min: 60 # time window for multi-IP check (default: 60)
cred_stuffing_distinct_accounts: 5 # failed accounts from one IP before credential_stuffing (default: 5)
pam_bruteforce_threshold: 5 # PAM auth failures from one IP before pam_bruteforce + auto-block (default: 5)
pam_bruteforce_window_min: 10 # window for counting PAM failures per IP (default: 10)
plugin_check_interval_min: 1440 # WordPress plugin check interval (default: 1440)
brute_force_window: 5000 # failed auth attempts window (default: 5000)
domlog_max_files: 500 # per-domain access logs per WP brute-force scan (default: 500)
domlog_tail_lines: 500 # trailing lines tailed from each domlog per scan (default: 500)
domlog_max_age_min: 30 # skip per-domain access logs untouched in this many minutes (default: 30)
mail_log_tail_lines: 500 # trailing lines of /var/log/exim_mainlog read by the mail-per-account scanner (default: 500)
syslog_messages_tail_lines: 200 # legacy direct-check FTP tail only; daemon FTP detection follows the log forward and ignores this setting (default: 200)
ftp_fail_window_min: 30 # sliding window (minutes) for per-IP pure-ftpd auth-failure accumulation before ftp_bruteforce (default: 30)
account_scan_max_files: 10000 # account and mail-domain paths per scanner cycle (default: 10000)
# If this cap clips /home/<account>/ paths, account_scan_truncated names the affected account.
# Rolling content coverage spends this cap once per cycle across the whole host, not once per account.
crontab_base64_blob_max_bytes: 16384 # encoded bytes per crontab base64 candidate before decoded-content matching; must be a multiple of 4 (default: 16384)
# Full-scan subsystem (`csm scan --full`) and rolling content coverage.
full_scan_max_file_mb: 16 # cap on a single file scanned by a full scan, in MiB (default: 16)
scan_job_retention: 20 # completed full-scan job records kept in the store (default: 20)
rolling_coverage: true # tri-state; default on. Each cycle content-scans a slice of dormant files past the mtime cap so old planted files get covered over time, taking the accounts covered longest ago first. Set false to disable
# Cursor write failures keep progress in memory until storage recovers; a restart loses that unpersisted progress.
# Periodic forced rescans expire old clean-file records, including files waiting for a later rolling window.
dropper_detection: true # tri-state; default on. Real-time flag for a PHP/executable created under a web docroot and unlinked before the TTL probe. Set false to disable
dropper_unlink_ttl_sec: 300 # seconds a fresh docroot PHP/executable is tracked before the self-delete probe (default: 300, range 30-3600)
# HTTP request flood, User-Agent spoof, and distributed HTTP detection.
# These detectors scan the same per-vhost access-log stream as the WP
# brute-force scanner; no extra log tailer is needed.
#
# http_flood_threshold: minimum per-IP request count inside the window
# that emits http_request_flood. 0 disables the detector. The detector
# ships disabled so operators can sample local baseline traffic first.
# Adjust up for CDNs or CGNAT-heavy visitor pools before enabling.
http_flood_threshold: 0 # 0 = disabled; set after sampling baseline traffic
http_flood_window_min: 5 # rate window in minutes (default: 5)
# http_ua_spoof_threshold: per-IP per-window count for non-browser UA
# kinds (including claimed search-engine bots such as Googlebot/Bingbot/
# Applebot that fail reverse-DNS confirmation) before http_ua_spoof fires.
# A single bot-like request from a residential or mobile client no longer
# hard-blocks it; sustained activity is required. Raise this for visitor
# pools that legitimately send unusual User-Agents.
http_ua_spoof_threshold: 30 # default: 30
# xmlrpc_threshold: per-IP POST /xmlrpc.php count before xmlrpc_abuse fires
# in access-log based detectors (a hard auto-block). WordPress sites using
# Jetpack, the mobile app, or WooCommerce make many legitimate xmlrpc.php
# calls, so this defaults higher than the legacy value. Set 0 to disable the
# check entirely; an absent key uses the built-in default.
xmlrpc_threshold: 100 # default: 100; 0 = disabled
# http_distributed_min_ips: distinct already-abusive source IPs that hit
# the same vhost in one scan window before a per-vhost distributed flood
# finding fires. 0 disables the rollup for existing configs that do not
# opt in.
http_distributed_min_ips: 10 # sample setting; omit or set 0 to disable
# URL scanner profile: fires http_scanner_profile for source IPs whose
# in-window traffic is almost entirely probe-error responses spread
# across many distinct request paths -- the shape of random-URL
# enumeration hunting for downloadable files, exposed backups, and
# dormant shells. Three gates must all pass: a minimum request volume,
# a minimum error percentage, and a minimum count of distinct error
# paths (query strings stripped, so cache-buster URLs on one missing
# endpoint count once). A visitor following dead links, a broken image
# hammered by one page, and a site migration all stay out of scope.
#
# 301 is deliberately not in the default status set: http->https and
# www redirects make every legitimate visitor 301-heavy, and a site
# migration redirects entire domains. Add 301 to
# http_scanner_status_codes only on hosts where that traffic shape is
# impossible.
http_scanner_min_requests: 0 # volume gate; 0 = disabled (default); 30 is a safe start
http_scanner_error_pct: 90 # min % of requests with probe-error status (default: 90)
http_scanner_min_distinct_paths: 10 # min distinct error paths, max 512 (default: 10)
http_scanner_status_codes: [404, 403] # statuses counted as probe errors (default: 404, 403)
# Distributed expensive-request crawl from one ASN. The detector fires
# only when request breadth, uncacheable volume, account share, and PHP
# worker saturation agree. Set min_ips to 0 to disable.
http_asn_crawl_window_min: 60
http_asn_crawl_min_ips: 25
http_asn_crawl_min_expensive: 250
http_asn_crawl_min_share_pct: 50
http_asn_crawl_high_amp_pct: 50
http_asn_crawl_high_volume_mult: 4
http_asn_crawl_saturation: 0 # 0 = performance.php_process_warn_per_user
http_asn_crawl_max_prefix: 8
http_asn_crawl_16_pref_pct: 60
http_asn_crawl_max_tracked_ips: 20000
http_asn_crawl_allowlist_asns: []
http_asn_crawl_reverse_proxy_asns: [13335, 54113, 20940]
# These three opt-in flags extend UA spoof detection to additional UA
# classes. Leave disabled on busy shared hosts; scripting-language agents
# and headless browsers appear on many legitimate monitoring stacks.
http_ua_scripting_enabled: false # flag curl/wget/python-requests/Go-http style UAs
http_ua_headless_enabled: false # flag Puppeteer/Playwright/PhantomJS UAs
http_ua_empty_enabled: false # flag requests with no UA at all
# SMTP brute-force tracker (Exim mainlog, dovecot SASL on submission ports)
smtp_bruteforce_threshold: 5 # per-IP failed auths before block (default: 5)
smtp_bruteforce_window_min: 10 # sliding window in minutes (default: 10)
smtp_bruteforce_suppress_min: 60 # cooldown between repeat findings (default: 60)
smtp_bruteforce_subnet_threshold: 8 # unique IPs per /24 before subnet block (default: 8)
smtp_account_spray_threshold: 12 # unique IPs targeting one mailbox before visibility finding (default: 12)
smtp_bruteforce_max_tracked: 20000 # soft cap on tracked entries; oldest evicted (default: 20000)
# Slow-brute pair: failed auths from one IP across 3+ distinct mailboxes
# within the long window. Catches attackers pacing below the fast window.
# Also fires on breadth alone: 10+ distinct walked mailboxes trigger the
# block regardless of the failure count. A successful SMTP auth from the
# source inside the window disqualifies it (live office NAT, not a walker).
smtp_bruteforce_slow_threshold: 40 # failures before block (default: 40; 3-1024; explicit 0 disables)
smtp_bruteforce_slow_window_min: 360 # slow sliding window in minutes (default: 360; 1-10080)
# SMTP probe-abuse tracker (raw connect-rate per IP; catches scanners that
# never reach AUTH). Threshold sized well above any legitimate MUA usage.
smtp_probe_threshold: 100 # per-IP connects before block (default: 100; explicit 0 disables)
smtp_probe_window_min: 5 # sliding window in minutes (default: 5)
smtp_probe_suppress_min: 60 # cooldown between repeat findings (default: 60)
smtp_probe_max_tracked: 20000 # soft cap on tracked entries; oldest evicted (default: 20000)
# Mail brute-force tracker (IMAP/POP3/ManageSieve via mail_logs source)
mail_bruteforce_threshold: 5 # per-IP failed auths before block (default: 5)
mail_bruteforce_window_min: 10 # sliding window in minutes (default: 10)
mail_bruteforce_suppress_min: 60 # cooldown between repeat findings (default: 60)
mail_bruteforce_subnet_threshold: 8 # unique IPs per /24 before subnet block (default: 8)
mail_account_spray_threshold: 12 # unique IPs targeting one mailbox before visibility finding (default: 12)
mail_bruteforce_max_tracked: 20000 # soft cap on tracked entries; oldest evicted (default: 20000)
# Slow-brute pair, as above. A recent successful login from the source or
# failures confined to mailboxes it holds established good standing for
# keep the source on the advisory path instead of blocking. Also fires on
# breadth alone: 10+ distinct walked mailboxes trigger regardless of count.
mail_bruteforce_slow_threshold: 40 # failures before block (default: 40; 3-1024; explicit 0 disables)
mail_bruteforce_slow_window_min: 360 # slow sliding window in minutes (default: 360; 1-10080)
mail_brute_account_key: "builtin:dovecot-user" # builtin:dovecot-user | builtin:postfix-sasl | regex:<capture>
modsec_escalation_hits: 3 # denies from one IP before ModSecurity escalation (default: 3)
modsec_escalation_window_min: 10 # ModSecurity escalation window in minutes (default: 10)
modsec_low_confidence_escalation_hits: 30 # low-confidence-only backstop (default: 30)
# --- Web server overrides ---
# Leave these empty to use auto-detected paths for the running platform.
web_server:
type: "" # apache | nginx | litespeed; empty = detect
config_dir: "" # optional Apache/Nginx config root
access_logs: [] # candidate access logs, replacing detected paths
error_logs: [] # candidate error logs, replacing detected paths
modsec_audit_logs: [] # candidate ModSecurity audit logs
# Override the per-vhost access-log glob patterns. Empty uses the
# auto-detected default for the panel (cPanel, Plesk, DirectAdmin,
# bare Apache, or bare Nginx).
domlog_globs: []
# IPs or CIDRs whose X-Forwarded-For header is trusted for client-IP
# extraction. Leave empty to ignore XFF and use RemoteIP as-is.
trusted_proxies: []
# --- Infrastructure ---
infra_ips: [] # management IPs/CIDRs/hostnames - never blocked
# --- Mail Logs ---
# Packaged releases include journald support. Custom builds need
# `make JOURNAL=1 build-yara` before `source: journal` can be selected.
mail_logs:
source: auto # auto | file | journal
file: "" # optional path override for file source
units: ["postfix", "dovecot"] # journal units for source=journal or auto fallback
# --- State ---
state_path: "/var/lib/csm/state" # bbolt DB and state files
# --- Suppressions ---
suppressions:
upcp_window_start: "00:30" # cPanel nightly update window start
upcp_window_end: "02:00" # cPanel nightly update window end
known_api_tokens: [] # API tokens to ignore in auth logs (e.g. ["phclient"])
ignore_paths: # glob patterns to skip in filesystem scans
- "*/cache/*"
- "*/vendor/*"
suppress_webmail_alerts: true # don't alert on webmail logins
suppress_cpanel_login_alerts: false # don't alert on cPanel direct logins
suppress_blocked_alerts: true # don't alert on attacks from IPs already blocked or challenged
trusted_countries: ["RO"] # requires a loaded GeoLite2-City database; download credentials alone do not supply country data
# --- Auto-Response ---
auto_response:
enabled: false
kill_processes: false # kill malicious processes
quarantine_files: false # move malware to quarantine
max_file_actions_per_hour: 50 # shared quarantine and file-cleaning attempt budget
max_file_actions_per_account_per_hour: 10 # per-account share of the same rolling hour
max_file_action_failures_per_hour: 3 # pause automatic file response after repeated failures
block_ips: false # block attacker IPs via firewall
block_expiry: "24h" # positive temporary block duration; omit for the 24h default
http_asn_crawl_tempban: "24h" # Critical ASN-crawl subnet ban duration
max_blocks_per_hour: 50 # per-IP blocks per hour; 0/omitted uses default
enforce_permissions: false # auto-chmod 644 world/group-writable PHP files
fix_wp_cron: false # on perf_wp_cron findings, auto-disable WP-Cron and install a per-user system cron
http_scanner_action: "challenge" # response for http_scanner_profile: "challenge" (default) routes to the PoW page, "block" bans the IP
block_cpanel_logins: false # block IPs on cPanel/webmail/FTP/API thresholded brute findings (multi-IP login, webmail/API brute, FTP brute). Single direct cPanel form logins stay audit-only regardless of this flag.
netblock: false # auto-block IPv4 /24 or IPv6 /64 subnets
netblock_threshold: 3 # IPs from same IPv4 /24 or IPv6 /64 before subnet block; minimum 2, omit for the default
netblock_window: "168h" # blocked IPs count this far back, expired blocks included; omit for the 168h default
permblock: false # promote temp blocks to permanent
permblock_count: 4 # temp blocks before promotion; minimum 2, omit for the default
permblock_interval: "24h" # positive counting window; omit for the 24h default
clean_database: false # auto-drop confirmed malicious DB objects after backup
clean_htaccess: false # auto-clean .htaccess directives flagged by hardened detectors (backups under /opt/csm/quarantine/pre_clean/)
virtual_patch_exposed_files: "off" # off, manual CLI apply, or dry-run-gated auto apply except sample SQL
disable_enforce_af_alg: false # suspend periodic AF_ALG hardening re-assertion
copy_fail_kill_process: false # kill processes caught opening AF_ALG sockets via the live listener
mail_auth_recovery:
restart_enabled: false # opt-in restart after sustained cPanel mail auth backend outage
down_grace: "10m" # continuously-down duration before restart
max_restarts_per_hour: 3 # hourly restart-attempt cap
restart_command: "/usr/local/cpanel/scripts/restartsrv_dovecot"
dry_run: true # safe default; previews IP blocks and web-exposed-file virtual patches
verdict_callback:
enabled: false # call panel before each auto-block
url: "" # POST target for verdict requests
hmac_secret: "" # signing secret, or use hmac_secret_env
hmac_secret_env: "" # env var read at call time
allow_unsigned: false # true only for staged unsigned rollouts
require_response_signature: true # reject unsigned callback replies
timeout_sec: 2 # callback request timeout
# PHP-relay auto-freeze. Off by default; only kicks in on cPanel hosts
# where email_protection.php_relay.enabled is true. dry_run defaults to
# true even when freeze is true, so an operator who enables freeze
# without thinking gets a dry-run rather than a live exim -Mf storm.
# Override at runtime with `csm phprelay dry-run on|off|reset`.
php_relay:
freeze: false # opt in to wire the exim -Mf hook into the alert pipeline
dry_run: true # safe default; flip with `csm phprelay dry-run off [--persist]`
max_actions_per_minute: 60 # rolling 60s cap on exim -Mf invocations
# --- Detection ---
detection:
# Known-vulnerable WordPress plugin matching is default-on and alert-only.
# Omit for the default, or set false to disable it. Allow entries suppress
# only the exact reviewed slug and installed version.
# vulnerable_plugin_scanning: true
vulnerable_plugin_allow: [] # entries: <slug>@<version>, case-insensitive
# db_object_scanning is tri-state: omit for the default (on),
# `false` to explicitly disable. When off, the MySQL persistence
# scanner emits no findings; the manual `csm db-clean --drop-object`
# CLI keeps working for operator-driven cleanup.
# db_object_scanning: true
db_object_allowlist: [] # entries: <account>:<schema>:<type>:<name> -- suppresses db_unexpected_* warnings only
admin_overlap_min_accounts: 2 # raise only if routine shared-admin accounts are expected on this host
admin_overlap_trusted_emails: [] # exact reviewed admin emails that may manage multiple cPanel accounts
admin_overlap_trusted_domains: [] # exact reviewed email domains for developer or reseller admin accounts
# rescan_on_signature_update: true # tri-state; omit for default-on, false to disable retroactive sweeps
af_alg_backend: "auto" # auto | bpf | auditd | none
connection_tracker_backend: "auto" # auto | bpf | legacy | none
connection_poll_interval: 30s # legacy connection tracker interval
exec_monitor_backend: "auto" # auto | bpf | legacy | none
exec_monitor_poll_interval: 30m # legacy process monitor interval
sensitive_files_backend: "auto" # auto | bpf | legacy | none
sensitive_files_poll_interval: 5m # sensitive-file poll/watchset refresh interval
direct_smtp_egress:
enabled: false # detect non-MTA local processes opening outbound SMTP
backend: "auto" # auto | bpf | legacy | none
dry_run: true # safe default for detector-scoped action
ports: [25, 465, 587] # destination ports to inspect
bad_asn_outbound:
enabled: false # off by default; third leg of the host_takeover chain. Needs the GeoLite2-ASN database and operator-supplied ASN lists
blocked_asns: [] # ASNs always treated as bad (e.g. known bulletproof hosters)
allowed_asns: [] # non-empty switches to allowlist mode: any destination ASN outside this set is treated as bad
# --- BPF Enforcement ---
bpf_enforcement:
enabled: false # master switch for in-kernel denial
dry_run: true # log intended denials, allow the connect
direct_smtp_egress: false # gate enforcement on direct SMTP egress matches
verdict_callback: false # userspace advisory callback after the BPF decision
# --- Challenge Pages ---
challenge:
enabled: false # enable PoW challenge pages instead of hard block
listen_addr: 127.0.0.1 # bind address; use 0.0.0.0 for public direct redirects
listen_port: 8439 # port for challenge server; must fit the TCP port range
tls_cert: "" # optional HTTPS cert for direct/public challenge listener
tls_key: "" # optional HTTPS key for direct/public challenge listener
public_url: "" # required by webserver-integration, e.g. https://host:8439/challenge
secret: "" # HMAC secret for tokens (auto-generated if empty)
difficulty: 2 # SHA-256 proof-of-work difficulty 0-5 (default: 2)
trusted_proxies: [] # IPs/CIDRs allowed to supply X-Forwarded-For
port_gate:
enabled: false # nftables gate for non-loopback challenge listener
captcha_fallback: # widget for JS-disabled visitors (default off)
provider: "" # "turnstile" | "hcaptcha" | "" (off)
site_key: "" # public key embedded in the widget
secret_key: "" # verified server-side
timeout: 10s
verified_session: # signed-cookie bypass for authenticated operators
enabled: false
cookie_name: csm_admin_session
ttl: 4h
admin_secret: "" # POST'd to /challenge/admin-token to mint cookie
verified_crawlers: # reverse-DNS forward-confirm for search crawlers
enabled: false
providers: [] # names: googlebot | bingbot
cache_ttl: 15m
# --- PHP Shield ---
php_shield:
# On CloudLinux, PHP runs inside a CageFS cage that cannot see the Shield
# event socket. Installing or upgrading the Shield registers its root-owned,
# non-writable directory in /etc/cagefs/cagefs.mp. Existing cages must be
# remounted during a maintenance window: cagefsctl --remount-all recreates
# every LVE on the host and disrupts running cages. A conflicting mount entry
# is left unchanged.
# Rewritten command-parameter probes are quieted only when the Shield's
# source fingerprint matches CSM's verified CMS content cache. Install the
# updated Shield with csm install --php-shield-only; older events, unverified or
# modified scripts, and incomplete request paths still alert. All recognized
# observations remain in the root-only event archive, including quiet probes.
# This does not verify code included by the entry script; content scanning
# and runtime protection remain necessary.
enabled: false # receive PHP Shield events and emit alerts
# --- Reputation ---
reputation:
abuseipdb_key: "" # AbuseIPDB API key for IP reputation lookups
whitelist: [] # IPs to never flag as malicious
# Async PTR + forward-A verification for IPs that claim search-engine
# bot UAs (Googlebot, Bingbot, Applebot). When an IP claims a bot UA
# but reverse DNS does not confirm it, the request counts toward
# http_ua_spoof. Transient DNS lookup failures fail open and are
# retried later. Set false only if your resolver is unreliable. See
# docs/src/auto-response.md for the always-block behavior.
bot_verify_enabled: true # default: true
verified_bots: [] # optional custom crawler identities
# verified_bots:
# - name: "seranking" # rDNS-verified crawler
# ua_substrings: ["serankingbacklinksbot"]
# rdns_suffixes: ["seranking.com"]
# - name: "perplexitybot" # AI agent verified by published IP ranges
# ua_substrings: ["perplexitybot"]
# ip_ranges: ["18.97.9.96/29", "18.97.1.228/30"]
bot_ranges: # built-in AI-crawler range refresh
auto_update: true # default: true; restart required to change
update_interval: "24h" # default: 24h, minimum: 1h
rspamd:
enabled: false # include rspamd rolling history in IP reputation
url: "http://127.0.0.1:11334" # rspamd controller URL
token: "" # controller password, or use token_env
token_env: "" # env var read at query time
upstream:
enabled: false # include panel-side threat-intel cache scores
url: "" # HTTPS base URL; HTTP only allowed for loopback
token: "" # bearer token, or use token_env
token_env: "" # env var read at query time
cache_ttl_min: 15 # local cache TTL for upstream scores
timeout_sec: 5 # upstream request timeout
report:
enabled: false # opt-in abuse report delivery; restart required
classes: [] # bruteforce | php_relay | credential_stuffing | bad_asn_egress
spool_path: "" # default: <state_path>/abuse_reports.db
spool_max: 10000 # max queued reports per target
targets:
- name: "" # stable target name
url: "" # HTTPS collector URL; HTTP only allowed for loopback
transport: "hmac" # hmac | ed25519
node_id: "" # sender node ID
key_id: "" # receiver key ID
key_env: "" # HMAC secret or Ed25519 private key env var
token_env: "" # optional bearer token env var for HMAC targets
central:
enabled: false # opt-in central scored-set consume; restart required
set_url: "" # HTTPS scored-set endpoint; HTTP only for loopback
pubkey_env: "" # env var with Ed25519 public key hex
refresh_interval: 6h # pull interval; default 6h
action: "challenge" # off | challenge | block_if_local_corroborated
block_threshold: 80 # score needed before local corroboration can block
# --- Signatures ---
signatures:
rules_dir: "/opt/csm/rules" # YAML signature rules directory
update_url: "" # remote URL to fetch rule updates
auto_update: false # auto-download rules on schedule
update_interval: "" # how often to check (e.g. "24h")
signing_key: "" # required for any remote rule update path; 64-char hex Ed25519 public key
allow_rule_count_decrease: false # permit a deliberate drop below half the installed YAML rule count
yara_forge:
enabled: false # auto-fetch YARA Forge community rules
tier: "core" # "core", "extended", "full" (default: "core")
update_interval: "168h" # how often to check for updates (default: weekly)
download_url: "" # signed ZIP URL/template; supports {tier} and {version}
disabled_rules: [] # rule names to switch off, in Forge and in the shipped rules
# yara_worker_enabled: true # tri-state: omit for the default (on), `false` to explicitly disable
# signatures.signing_key is mandatory whenever either signatures.update_url
# is set or signatures.yara_forge.enabled is true. It must be the hex
# Ed25519 public key used to verify detached .sig files for rule bundles.
# Remote update URLs must use HTTPS and must not point at localhost,
# loopback, link-local, unspecified, or RFC1918 / ULA private addresses.
#
# YARA Forge upstream GitHub releases do not publish CSM detached signatures.
# To enable automatic Forge updates, mirror the ZIPs, sign each ZIP, publish
# the signature at the ZIP URL plus .sig, and set yara_forge.download_url to
# that signed mirror. Otherwise leave update_url empty and yara_forge.enabled
# false.
# --- Web UI ---
webui:
enabled: true
listen: "0.0.0.0:9443" # address:port for HTTPS server
auth_token: "" # API/login credential (auto-generated on install)
session_lifetime: "24h" # browser absolute expiry; restart required
session_idle_timeout: "30m" # browser idle expiry; restart required
tokens: [] # optional scoped tokens: name/token/scope (admin or read)
metrics_token: "" # optional Bearer token for /metrics only
tls_cert: "" # path to TLS certificate PEM file
tls_key: "" # path to TLS private key PEM file
ui_dir: "" # path to UI files on disk (default: /opt/csm/ui)
allowed_origins: [] # extra browser origins ("https://host[:port]") allowed to call the API; loopback is always allowed
# --- Email AV ---
email_av:
enabled: false
clamd_socket: "/var/run/clamd.scan/clamd.sock" # path to ClamAV daemon socket
scan_timeout: "30s" # per-attachment scan timeout
max_attachment_size: 26214400 # max single attachment size in bytes (25MB)
max_archive_depth: 1 # max nested archive extraction depth
max_archive_files: 50 # max files extracted per archive; also caps encrypted member names reported per message
max_extraction_size: 104857600 # max total extraction size in bytes (100MB)
quarantine_infected: true # quarantine emails with infected attachments
scan_concurrency: 4 # parallel scan workers
fail_mode: "open" # behavior when a scan cannot complete: "open" (default) delivers; "tempfail" defers so Exim retries
# Permission events have a five-second scan hold budget. When it expires,
# the open is allowed ("open") or denied for retry ("tempfail"); an admitted
# scan continues. A full scan queue applies the same policy immediately.
# Repeated deadline or capacity exhaustion raises email_av_hold_bypass and
# stops queueing new scans for five minutes. During this cooldown new opens
# are allowed unscanned ("open") or deferred ("tempfail"). Scanning resumes
# automatically afterwards; bypassed messages are not scanned later.
# Late malware results can quarantine only the original, unlocked message.
# Messages in active delivery, replaced spool files, and messages with a
# pending delivery journal are left untouched with email_av_quarantine_error.
# email_av_late_verdict means an allowed open later received a scan decision
# to stop delivery; it does not mean a previously deferred message escaped.
# Password-protected archive attachments are outside fail_mode. Their members
# cannot be read without the password, so no retry ever makes them scannable
# and "tempfail" would defer the message until it bounced. CSM delivers them
# and reports email_av_encrypted_archive naming the archive and the member,
# at most once an hour. A full alert queue does not consume this allowance.
# Encrypted member names are capped at max_archive_files per message; the
# warning counts any additional encrypted members without deferring delivery.
# Both encryption schemes are covered:
# legacy ZipCrypto and WinZip AES. Blocking such mail outright is an Exim
# policy decision and is not something CSM does for you.
# --- Email Protection ---
email_protection:
password_check_interval_min: 1440 # how often to audit email passwords (default: 1440)
high_volume_senders: [] # accounts expected to send high volume (skip rate alerts)
rate_warn_threshold: 50 # emails per window before warning (default: 50)
rate_crit_threshold: 100 # emails per window before critical (default: 100)
rate_window_min: 10 # rate check window in minutes (default: 10)
known_forwarders: [] # expected forwarders and Sieve :copy rules
# PHP-relay detector (cPanel only; gated by platform.IsCPanel at startup).
# Off by default. When enabled, the daemon spawns the inotify spool
# watcher, runs a startup spool walk, and starts the Path 2b retro scan
# on /var/log/exim_mainlog. See docs/src/detection-realtime.md#php-relay
# for what each path actually triggers on.
php_relay:
enabled: false # opt in to start the watcher
rate_window_min: 5 # Path 1 rolling window
header_score_volume_min: 5 # Path 1: don't score until script has emitted N msgs
absolute_volume_per_hour: 30 # Path 2 threshold per script
account_volume_per_hour: 0 # Path 2b operator override; 0 = auto-derive from cpanel.config maxemailsperhour
reputation_failures_per_24h: 3 # Path 3 threshold (Stage 2)
fanout_distinct_scripts: 3 # Path 4 threshold
fanout_distinct_recipients: 5 # Path 2 + Path 4 recipient-diversity gate; 0 disables only this gate
fanout_window_min: 5 # Path 4 window
baseline_sigma: 3.0 # Path 5 (Stage 3)
baseline_observation_days: 7 # Path 5 (Stage 3)
policies_dir: "/opt/csm/policies/php_relay" # mailer_classes.yaml + http_proxy_ranges.yaml; SIGHUP-reloadable
cloud_relay:
allow_users: [] # full mailbox opt-outs for cloud-relay detection
allow_domains: [] # domain-wide opt-outs for cloud-relay detection
# Email forward guard (cPanel only). Opt-in MTA-native enforcement for
# external forward copies. Enforce mode can hold null-sender backscatter and
# bad-sender-IP copies before they relay to an external provider, while the
# local mailbox copy still delivers. Spam, malware, and auth-fail signals are
# accounted in dry-run until Exim content scanning is wired. CSM is not in the
# live mail path; an installed Exim rule can keep holding matching copies even
# if the daemon is down. Held copies can be released or deleted from the Email page.
forward_guard:
enabled: false # master switch (default off)
dry_run: true # account/log only, do not actually hold (default true)
quarantine_retention_days: 14 # held-copy retention window
skip_forwarders: [] # reserved forwarder exemptions; not enforced yet
hold_signals: # signal toggles, each default true
bounce_backscatter: true # null-sender bounce backscatter (enforceable)
spam_flagged: true # message flagged as spam (dry-run/accounting only)
malware: true # message carries malware (dry-run/accounting only)
bad_sender_ip: true # originating IP has bad reputation (enforceable)
auth_fail: true # sender failed SPF/DKIM/DMARC auth (dry-run/accounting only)
# --- Firewall ---
firewall:
enabled: false
# Open ports (IPv4). SSH (22) is intentionally absent; uncomment in
# the YAML lists if sshd listens on 22. TCP 853 is DNS-over-TLS;
# UDP 853 is DNS-over-QUIC.
# 6277/24441 are DCC/Pyzor network checks used by SpamAssassin.
tcp_in: [20,21,25,26,53,80,110,143,443,465,587,853,993,995,2077,2078,2079,2080,2082,2083,2091,2095,2096]
tcp_out: [20,21,25,26,37,43,53,80,110,113,443,465,587,853,873,993,995,2082,2083,2086,2087,2089,2195,2325,2703]
udp_in: [53,443,853]
udp_out: [53,113,123,443,853,873,6277,24441]
# IPv6
ipv6: false
tcp6_in: [] # if empty, uses tcp_in
tcp6_out: [] # if empty, uses tcp_out
udp6_in: [] # if empty, uses udp_in
udp6_out: [] # if empty, uses udp_out
# Restricted ports (infra IPs only)
restricted_tcp: [2086,2087,2325,9443] # WHM and CSM Web UI ports
# Outbound ports a service on this host needs. Checked, never added to
# the policy: validation warns when an effective family policy omits one.
# Integrations declare theirs in their own conf.d fragment.
required_tcp_out: []
# Passive FTP range
passive_ftp_start: 49152
passive_ftp_end: 65534
# Destination-scoped outbound TCP. tcp_out holds single ports only, so a
# port range can only be expressed here -- needed when this host is an FTP
# *client* and must open passive data connections to a known source.
# dst takes an IP or CIDR. 0.0.0.0/0 means any IPv4 destination; ::/0 means
# any IPv6 destination. Each rule applies to its own address family, and
# validation warns: a wide range to anywhere is the outbound path the
# output chain exists to close. When smtp_block is on, these rules follow
# its guard and ranges that overlap smtp_ports are rejected.
# An IPv6 destination emits no rule unless firewall.ipv6 is true.
# IPv4-mapped prefixes under ::ffff:0:0/96 are treated as IPv4 prefixes.
# Lockout warnings only credit valid rules in the connection's family;
# scoped exceptions still warn because hostname resolution is not checked.
tcp_out_allow: []
# tcp_out_allow:
# - dst: 203.0.113.10/32
# port_start: 49152
# port_end: 65534
# Infra IPs/CIDRs/hostnames for firewall rules
infra_ips: []
# Rate limiting. SYN/conn-rate/UDP are dual-stack (IPv6 keyed per /64);
# conn_limit is IPv4-only.
conn_rate_limit: 200 # new connections/min per source (CGNAT-tolerant; IPv6 per /64; 0 = disabled; null = default)
syn_flood_protection: true # per-source SYN flood meter (IPv6 per /64)
conn_limit: 400 # max concurrent connections per IPv4 source, IPv4 only (0 = disabled)
# Per-port flood protection: rate-limit new connections per source IP and IP family.
# Defaults are sized for a busy mail host: 600/300s = 120 new conns/min/IP,
# which tolerates a Thunderbird/iPhone client opening 5-15 parallel sessions
# while still capping single-IP flood storms.
port_flood:
- port: 25
proto: tcp
hits: 600
seconds: 300
- port: 465
proto: tcp
hits: 600
seconds: 300
- port: 587
proto: tcp
hits: 600
seconds: 300
# DoS-exempt ranges: bypass configured DoS meters and subnet auto-blocks.
# See firewall.dos_exempt_ranges below for full details.
dos_exempt_ranges: [] # CIDRs or single IPs exempt from connection/mail-port meters and subnet auto-blocks; /0 and bare hostnames rejected at load
dos_exempt_known_mail_providers: true # also exempt Google and Microsoft mail-provider egress ranges; dynamic, cached, updated every 12h (default: true)
# UDP flood protection (per-source meter; IPv6 per /64)
udp_flood: true
udp_flood_rate: 100 # packets per second
udp_flood_burst: 500 # burst allowance
# Country blocking
country_block: [] # ISO country codes to block
country_db_path: "" # country CIDR directory (default: <state_path>/geoip, filled by `csm firewall update-geoip`)
# Silent drop (no logging)
drop_nolog: [23,67,68,111,113,135,136,137,138,139,445,500,513,520]
# IP limits
deny_ip_limit: 3000 # max permanent blocked IPs
deny_temp_ip_limit: 500 # max temporary blocked IPs
# Outbound SMTP restriction
smtp_block: false # block outgoing mail except allowed users
smtp_allow_users: [cpanel, mailman] # extra outbound SMTP users; root and mailnull are always allowed
smtp_ports: [25,465,587]
# Dynamic DNS
dyndns_hosts: [] # hostnames to resolve and whitelist periodically
# Logging
log_dropped: true # log dropped packets
log_rate: 5 # log entries per minute
# --- GeoIP ---
geoip:
account_id: "" # MaxMind account ID
license_key: "" # MaxMind license key
editions: # MaxMind database editions
- GeoLite2-City
- GeoLite2-ASN
auto_update: true # auto-update GeoIP databases (default: true when credentials set)
update_interval: "24h" # update check interval
# --- ModSecurity ---
modsec_error_log: "" # path to Apache/LiteSpeed error log for ModSec parsing
modsec:
rules_file: "" # path to modsec2.user.conf
overrides_file: "" # path to csm-overrides.conf
reload_command: "" # command to reload web server (e.g. "/usr/sbin/apachectl graceful"); also activates CSM rule updates
# --- Performance ---
performance:
enabled: true
load_high_multiplier: 1.0 # load average / CPU cores multiplier for warning (default: 1.0)
load_critical_multiplier: 2.0 # load average / CPU cores multiplier for critical (default: 2.0)
php_process_warn_per_user: 20 # per-user PHP process count warning (default: 20)
php_process_critical_total_multiplier: 5 # total PHP processes / CPU cores for critical (default: 5)
error_log_warn_size_mb: 50 # error log size warning threshold (default: 50)
mysql_join_buffer_max_mb: 64 # MySQL join_buffer_size warning threshold (default: 64)
mysql_wait_timeout_max: 3600 # MySQL wait_timeout warning threshold (default: 3600)
mysql_max_connections_per_user: 10 # per-user MySQL connections warning (default: 10)
redis_bgsave_min_interval: 900 # minimum seconds between Redis BGSAVE (default: 900)
redis_large_dataset_gb: 4 # Redis dataset size warning threshold in GB (default: 4)
wp_memory_limit_max_mb: 512 # WordPress memory_limit warning threshold (default: 512)
wp_transient_warn_mb: 1 # WordPress transient data warning in MB (default: 1)
wp_transient_critical_mb: 10 # WordPress transient data critical in MB (default: 10)
wp_cron_fix: # tuning for the WP-Cron remediation (manual fix from the Web UI or auto_response.fix_wp_cron)
interval_minutes: 15 # system cron frequency; default 15, clamped to [1, 60]
php_bin: "" # override interpreter; empty = per-vhost PHP, then auto-detect
# --- Cloudflare ---
cloudflare:
enabled: false # auto-whitelist Cloudflare IP ranges
refresh_hours: 6 # how often to refresh Cloudflare IPs (default: 6)
# --- Threat Intel ---
c2_blocklist: [] # known C2 server IPs to block permanently
backdoor_ports: [4444,5555,55553,55555,31337] # ports indicating backdoor activity
# --- Update check ---
updates:
check_enabled: true # notify only; CSM never downloads or applies updates
interval: "24h" # release check interval
github_api_url: "" # optional release API mirror or test endpoint
package_name: "csm" # apt/dnf package name for package-manager fallback
# --- Incidents ---
incidents:
auto_close:
enabled: true # auto-close idle open/contained incidents
dry_run: false # log decisions without writing status changes
by_kind:
mailbox_takeover: 24h
mailbox_bruteforce: 24h
credential_spray: 24h
web_attack: 24h
web_account_compromise: 168h
spray_suppression:
enabled: false # collapse one-source credential spray into one incident
dry_run: true
distinct_mailboxes: 10
severity_escalate_at: 50
per_check: [email_auth_failure_realtime, pam_bruteforce, credential_stuffing]
max_tracked_ips: 10000
block_at_severity: "" # "" | high | critical
auto_block:
enabled: false # block source IPs from incident correlations
block_at_severity: "" # "" | high | critical
kinds: [] # empty means all non-spray kinds with remote_ip
# --- Disabled checks (skip whole categories per host) ---
# Listed finding names disable the scheduled check runner(s) that emit them,
# including sibling findings from the same runner. Realtime findings are not
# affected. Use for whole categories that don't apply to a host (e.g. WAF/web
# checks on DNS-only cPanel servers, where httpd is installed but no virtual
# hosts serve traffic).
# For email-only suppression, use `alerts.email.disabled_checks` instead.
disabled_checks: [] # e.g. [waf_status, waf_rules, waf_detection_only]
# --- Retention (bbolt growth control) ---
retention:
enabled: false # opt-in; when true, a daily sweep prunes old entries
findings_days: 90 # keep active findings this long (0 disables the findings sweep)
history_days: 30 # keep findings-history entries and proven firewall
# outcomes this long; it also sets how far back a
# durable firewall action can be undone
reputation_days: 180 # keep IP reputation/attack entries this long
sweep_interval: "24h" # how often the retention goroutine runs
compact_min_size_mb: 128 # startup compaction floor; 0 disables auto-compaction
compact_fill_ratio: 0.5 # compact when used_bytes / file_size drops below this
# --- Debug / diagnostics ---
debug:
pprof_listen: "" # e.g. "127.0.0.1:6060"; MUST be loopback. Empty disables.
# Reach it over SSH: go tool pprof http://127.0.0.1:6060/debug/pprof/heap
# While listening, samples mutex and block contention; adds CPU overhead even without a profile request.
# --- Sentry (error reporting) ---
sentry:
enabled: false # ship panics and selected errors to a Sentry server
dsn: "" # Sentry project DSN
environment: "production" # e.g. "production", "staging"
sample_rate: 1.0 # 0.0 -> 1.0 (capture all errors)
debug: false # SDK debug logs to stderr
Blocked-source alert suppression
suppressions.suppress_blocked_alerts mutes single-source attack notifications
when that address is already blocked or was blocked in the same batch. An active
challenge also mutes checks eligible for challenge routing; it does not cover
FTP or mail authentication failures, or checks configured to require a hard block.
The same policy applies to daemon alerts and csm scan --alert.
Compromise evidence, successful logins, suspicious mail content and findings that summarize multiple sources remain visible. Blocking one address does not resolve those findings. This setting does not remove findings from the audit log.
TLS Certificates
The Web UI serves over HTTPS. Configure TLS certificates under webui:
webui:
tls_cert: "/var/cpanel/ssl/cpanel/mycpanel.pem" # certificate PEM file
tls_key: "/var/cpanel/ssl/cpanel/mycpanel.pem" # private key PEM file
On cPanel servers, you can reuse the cPanel self-signed certificate (both cert and key are in the same PEM file). For production, use a proper certificate from Let’s Encrypt or your CA.
The private key may appear before or after the certificates. The first certificate must be the leaf served by the Web UI, followed by any chain certificates. Renewal only replaces an expiring CSM-generated leaf; it preserves the private key and all other contents of a combined file. Certificates issued outside CSM are left unchanged.
If both paths are empty, CSM generates an ECDSA self-signed certificate and key under state_path. Set both paths to use an operator-managed certificate. Configuring only one of the pair is invalid.
Validation
csm validate # syntax check
csm validate --deep # syntax + connectivity probes (SMTP, webhooks)
csm config show # display config with secrets redacted
Editing csm.yaml by hand
CSM stores a sha256 of the main config in integrity.config_hash and
a separate digest of loaded drop-ins in integrity.confd_hash. It
refuses to start if the on-disk files disagree with those values. This
is a tamper-detection feature. There are two supported edit workflows
depending on which fields you touch.
Fast path: SIGHUP reload (safe fields only)
For fields tagged as hot-reload-safe (alerts, thresholds,
detection, suppressions, auto_response, bpf_enforcement,
reputation, email_protection, disabled_checks, confd), the daemon can
accept the change without a restart:
sudo cp /etc/csm/csm.yaml /etc/csm/csm.yaml.bak-$(date +%s)
# edit /etc/csm/csm.yaml with your favourite editor
sudo systemctl reload csm
sudo journalctl -u csm -n 20 --no-pager
systemctl reload sends SIGHUP (wired via ExecReload= in the unit
file). The daemon re-reads the file, validates it, diffs it against
the running config, and if every change is on a field tagged
hotreload:"safe" it swaps the new values into
the live config and re-signs the integrity hashes on disk. The
next check tick sees the new thresholds; fanotify marks are not
dropped.
The tagged-safe top-level fields are alerts, thresholds,
detection, suppressions, auto_response, bpf_enforcement,
reputation, email_protection, and disabled_checks. The Settings
API derives its restart hints from the same manifest that drives
config.Diff, so UI hints and SIGHUP behavior cannot drift silently.
Changes to their sub-keys are picked up on the next tick by the
periodic scanners, the auto-response helpers
(block/kill/quarantine/challenge/permission-fix), alert dispatch, and
the heartbeat.
reputation.verified_bots is reconciled on reload and the bot-verifier
cache is restamped when the list changes.
reputation.bot_ranges controls a long-lived updater goroutine and is
tagged restart-required. A reload that changes it emits
config_reload_restart_required; restart the daemon to change the
auto-update switch or interval.
Two sub-keys are runtime-state exceptions. They live under a safe-tagged parent and SIGHUP accepts edits to them, but they seed long-lived in-memory state at daemon startup. Restart the daemon if you need the running process to use the new value:
reputation.whitelist– seeded into the threat database at startup and replaced on every successful reload, so editing the list in csm.yaml and sending SIGHUP takes effect at once. The threat database also exposes its own runtime API for adding and removing whitelist entries (via the Threat Intelligence page in the Web UI or the/api/v1/threat/*endpoints); those entries are persisted to disk and are kept separate from the configured list, so a reload never drops them.email_protection.known_forwarders– captured by the forwarder watcher at startup and read by scheduled forwarder and mail-filter checks. No runtime API yet; send a restart if you edit this list.
auto_response.mail_auth_recovery is a restart-required sub-key under
the otherwise safe auto_response section. It is captured by the cPanel
mail auth backend probe at startup, so a reload that changes it emits
config_reload_restart_required and leaves the live config unchanged.
alerts.block_digest is restart-required under the otherwise safe
alerts section. The collector and ticker are built at startup, so a
reload that changes the digest settings emits
config_reload_restart_required and leaves the live config unchanged.
When countries is empty, the fallback to suppressions.trusted_countries
still follows safe reloads. Delivery uses the current email and webhook
settings. Digest output includes by-country, by-category, and by-reason
counts; WAF categories group ModSecurity escalations and high-volume WAF
attacker blocks together.
The rest of the sub-keys in every safe-tagged section are read per-call (inside check functions, auto-response helpers, alert dispatchers) and hot-reload cleanly on the next tick.
Look for one of three log shapes in the journal:
SIGHUP: config reloaded; safe fields updated: [thresholds]– success. The new values are live.config_reload_restart_required: SIGHUP reload: restart-required fields changed: [hostname ...]; live config unchanged– the edit touched a field that cannot be hot-swapped. A Warningconfig_reload_restart_requiredfinding is also emitted. Fall back to the restart path below.config_reload_error: SIGHUP reload: parse failed ...or... validation error ...– the file on disk is not loadable or failscsm validate. A Criticalconfig_reload_errorfinding is emitted. The live config is unchanged; fix the file and repeat.
Restart path: unsafe fields
Fields not tagged hotreload:"safe" (the majority, including
hostname, state_path, webui.listen, firewall.*, email_av.*
and anything that survives only one re-init per daemon lifetime)
require a full restart. The integrity check must be re-signed first:
sudo cp /etc/csm/csm.yaml /etc/csm/csm.yaml.bak-$(date +%s)
# edit /etc/csm/csm.yaml with your favourite editor
sudo /opt/csm/csm rehash # re-signs integrity hashes
sudo /opt/csm/csm validate # syntax + value sanity
sudo systemctl restart csm
sudo systemctl status csm # confirm active, no crash-loop
If the restart fails (most commonly because rehash was skipped),
roll back with
sudo cp <backup> /etc/csm/csm.yaml && sudo systemctl restart csm.
The backup carries its own matching hash so no second rehash is
needed.
Config-management tools
Config-management workflows (Ansible, Puppet, Chef) should:
- For safe changes, notify
systemctl reload csminstead ofrestart. The daemon re-signs the hash itself; no separatecsm rehashstep is required. - For any change that may touch a restart-required field, run
csm rehashbefore the restart notify fires. Or always sendreloadfirst, read the journal, and promote torestartonly when the reload logsrestart-required.
conf.d drop-ins
Files matching /etc/csm/conf.d/*.yaml are loaded after the main config and deep-merged on top of it. Override with --config-dir <path> or CSM_CONFIG_DIR; the flag wins when both are set.
The YARA-X worker inherits the daemon’s selected directory on every start. If it is absent, the worker loads no fragments from it; it does not fall back to another directory. Explicit operator overrides still require an existing directory.
- Order: lexicographic by filename. Scalar keys in
20-overrides.yamloverride the same keys in10-base.yaml. Use a numeric prefix. - Merge semantics: maps merge recursively; scalars replace the value from the main file; lists append in fragment order. All-scalar lists drop duplicate entries while keeping the first occurrence; structured lists such as
webui.tokenskeep every entry. - Trust: override directories must be absolute, must exist, and must be owned by root or the running process. The directory and every loaded fragment must not be group- or world-writable. Safe symlinked fragments are allowed, so packaged profiles can still be linked into
/etc/csm/conf.d/. - Integrity ownership: drop-ins cannot set the
integrityblock. Integrity metadata is stored only in the main config. - Hash:
integrity.config_hashcovers the main file andintegrity.confd_hashcovers loaded drop-ins. After editing a drop-in by hand, runcsm rehashbefore restarting, or usesystemctl reload csmso the daemon can re-sign after validating the merged config. Rehash re-signs the resolved main file only; when a real legacy copy still exists at/opt/csm/csm.yamlwith the same operator content, rehash replaces it with the compatibility symlink so the two never drift. Web settings saves refuse to bless a drop-in change that has not already been re-signed. A mismatch is refused at the next daemon start with an error that namescsm rehash. Until then the running daemon keeps its old hashes, but its periodic integrity check raises a tamper finding and skips the scheduled checks on every cycle;csm doctorreports the same mismatch, so it can be fixed before a restart turns it into an outage. - Integration-owned fragments: a fragment its owning package rewrites on its own schedule (phpanel-server-agent rewrites
10-phpanel-runtime.yamlon every bootstrap) cannot be pinned by a static hash without making every one of those rewrites a failed restart. List such fragments by bare filename underconfd.integrity_exemptin the main config, then runcsm rehashonce. Their content is left out ofconfd_hash; every other fragment stays covered. The list lives in the main config, outside theintegrityblock, so it is itself covered byconfig_hash, and a fragment cannot setconfdto exempt itself. - Use cases: packaged integration profiles (e.g.
/usr/lib/csm/profiles/phpanel-agent.yamlsymlinked intoconf.d/), per-host automation that should not touch the operator’scsm.yaml, secret material rendered from a vault.
ls /etc/csm/conf.d/
# 10-phpanel-agent.yaml 20-tenant-overrides.yaml
csm validate # validates the merged config
csm config show # prints the merged, redacted config
csm config schema # JSON Schema for editor / CI validation
csm validate and csm config show always operate on the merged config so you can audit the effective state without grepping fragments.
detection.direct_smtp_egress
The direct SMTP detector’s backend accepts auto, bpf, legacy, or none;
ports must contain TCP ports in the 1-65535 range. See
Direct SMTP egress.
bpf_enforcement
BPF enforcement requires a BPF-capable connection tracker at
runtime; auto falls back to legacy detection on older servers. See
BPF enforcement.
firewall.dos_exempt_ranges
dos_exempt_ranges is a list of CIDRs or single IPs (/32 for IPv4, /128 for IPv6) declared by the operator. Entries are validated at startup; /0 default routes, bare hostnames, and malformed CIDRs are rejected before the daemon starts. Default: empty (no declared ranges).
dos_exempt_known_mail_providers (bool, default true) adds Google and Microsoft outbound mail IP ranges to the effective exempt set automatically. The ranges are sourced dynamically and cached on disk; a built-in snapshot covers startup before the first live refresh. The cache is refreshed every 12 hours.
Sources in the effective exempt set bypass the new-connection rate-limit for their IP family, the IPv4 concurrent connection-limit, and the TCP port 25/465/587 flood meters for their IP family. Subnet auto-block (spray, ASN-crawl, and netblock escalation) skips CIDRs that intersect an exempt range, and exempt IPs do not count toward the netblock threshold.
Exempt sources do not bypass: manual blocks (csm firewall deny or csm firewall deny-subnet), SYN flood protection, UDP flood protection, country blocking, and port policy. A manual block placed inside an exempt range still takes effect because blocked sets are evaluated before the DoS meters. See Firewall - DoS-exempt ranges.
Credential rotation
Environment-backed credentials come from the environment inherited when CSM
starts. Reading that environment before each request does not import edits
to an environment file or exports from another shell. Restart CSM after
changing its launch environment. csm rehash, SIGHUP, and
systemctl daemon-reload do not replace a running process’s environment.
The same rule applies to these integrations:
| Integration | Environment variable name configured in | Static fallback |
|---|---|---|
| Upstream threat intelligence | reputation.upstream.token_env | reputation.upstream.token |
| Rspamd | reputation.rspamd.token_env | reputation.rspamd.token |
| Verdict callback | auto_response.verdict_callback.hmac_secret_env | auto_response.verdict_callback.hmac_secret |
| Phpanel finding webhook | alerts.webhook.hmac_secret_env | alerts.webhook.hmac_secret |
A non-empty environment value takes precedence over the static fallback. An unset or empty variable uses that fallback. Each verdict exchange keeps one secret for both request signing and response verification.
Configure a systemd environment file
Create /etc/csm/credentials.env, owned by root with mode 0600. Put the
actual values in this file through your editor or configuration manager:
CSM_UPSTREAM_TOKEN=replace-with-upstream-token
CSM_RSPAMD_TOKEN=replace-with-controller-password
CSM_VERDICT_SECRET=replace-with-verdict-secret
CSM_WEBHOOK_SECRET=replace-with-webhook-secret
Set the corresponding CSM configuration fields to these variable names.
Then create /etc/systemd/system/csm.service.d/credentials.conf:
[Service]
EnvironmentFile=/etc/csm/credentials.env
Apply the new unit configuration and start CSM with that environment:
systemctl daemon-reload
systemctl restart csm.service
systemctl is-active csm.service
For later rotations, update the existing credentials file and run
systemctl restart csm.service. daemon-reload is needed when the unit or
its drop-ins change, not for a value change in an already referenced file.
For a foreground daemon, stop it and launch a new process with the updated
environment supplied by its launcher.
Coordinate the change with each receiver. Where the receiver supports overlapping credentials, retain the old credential through restarts and in-flight requests. Verify successful authentication with the new credential at the receiver without logging its value, then retire the old one.
Regression coverage
TestEnvironmentCredentialRotation launches a subprocess from an environment
file and captures real HTTP requests from all four clients. Requests before
and after an external file edit use the old credential while the process
stays running. Restarting the subprocess from the edited file switches all
four clients to the new credential. The test does not change the running
process’s environment with os.Setenv.
Custom account roots
account_roots accepts absolute, normalized directory paths or glob patterns.
The same configured content trees are eligible for manual remediation and
quarantine restore. Detected panel account homes and the existing scratch
quarantine locations keep their normal scope. A configured content directory
itself is not a restore target; restore operates beneath it.
Symlinks in a configured root or its ancestors are rejected. Unsafe trees are excluded from remediation without disabling other tenants’ valid roots. Destination directories are pinned during restore, including when a tenant renames a directory while the operation is running.
Service write access
The packaged service uses ProtectSystem=strict. A custom content tree needs a
write grant before the daemon can quarantine, clean, or restore files there.
Keep the account container directory owned by root and not writable by tenants:
account_roots:
- /srv/csm-accounts/*/public
For this layout, /srv/csm-accounts is root-owned; the account directories below
it can belong to their tenants. Provision the directories before generating the
drop-in:
set -e
fragment="$(mktemp)"
trap 'rm -f "$fragment"' EXIT
csm systemd-roots > "$fragment"
cat "$fragment"
install -d -m 0755 /etc/systemd/system/csm.service.d
install -m 0644 "$fragment" /etc/systemd/system/csm.service.d/50-account-roots.conf
systemctl daemon-reload
systemctl restart csm.service
csm doctor
csm systemd-roots only prints the drop-in. It grants existing content
directories or their nearest root-controlled ancestor, and adds mount ordering
for those paths. It refuses symlink roots and broad grants such as /, /srv,
or /etc. It escapes path quoting and systemd specifiers. The generated
fragment appends grants to the packaged service; it does not reset that list.
A tenant-owned directory cannot itself be a grant: its owner could replace it
with a symlink before systemd starts. If the command cannot find a narrow,
root-controlled ancestor, move the accounts beneath a dedicated root-owned
container directory and update account_roots.
Regenerate the fragment after changing roots or provisioning a tree outside
existing grants. Changing account_roots requires a daemon restart. If integrity
verification is enabled, update its baseline for an intentional configuration
change before restarting, as described in Configuration.
Verification
csm doctor checks configured custom roots and detected panel roots outside the
packaged home grant. It reports missing roots, unsafe aliases, missing service
grants, and read-only mounts in the running daemon’s filesystem view. A stopped
service or an unavailable systemd connection is reported as unverified access.
Reloading unit files alone does not change the running service’s mount namespace;
restart the service and run doctor again.
After setup, verify detection, quarantine, listing, and restore on a test account under the custom root. Confirm that the restored file still belongs to its tenant. A sibling tree outside the configured roots and existing scratch locations must remain ineligible for restore.
The repository includes a disposable systemd service test. Build an image from
build/Dockerfile.systemd-test, then use the Linux wrapper:
GO_LINUX_IMAGE=csm-systemd-test scripts/go-linux.sh bash scripts/systemd-account-roots-test.sh
cat .cache/systemd-account-roots/result
The test runs production detection, quarantine, listing, restore, and doctor inside the packaged sandbox, with a generated grant for a separate test volume. It verifies owner, permissions, timestamps, and rejected sibling and symlink destinations. This tests service confinement; it does not start the daemon’s watchers or replace panel integration coverage. Logs and generated grants are saved beside the result file. The harness refuses to run outside its disposable Linux container.
Service configuration writes
The packaged and CLI-installed units use ProtectSystem=strict. The daemon can
write these configuration directories when they exist:
| Directory | Runtime writer |
|---|---|
/etc/csm | CSM configuration and fragments |
/etc/audit | CSM audit rules and augenrules generated configuration |
/etc/modprobe.d | Repair of an existing, opted-in AF_ALG mitigation marker |
/etc/apache2/conf.d | cPanel Apache/LiteSpeed challenge snippet, virtual patches, and default overrides |
/etc/apache2/conf-enabled | Debian Apache challenge snippet |
/etc/httpd/conf.d | RHEL Apache challenge snippet |
/etc/nginx/conf.d | Nginx challenge snippet |
/usr/local/apache/conf | Legacy Apache virtual patches and overrides |
/usr/local/lsws/conf/templates | Standalone LiteSpeed challenge template |
The other state, quarantine, account, spool and cPanel plugin grants remain in
place. Optional directory grants tolerate absent software. Install the relevant
web server before starting CSM; restart CSM after creating a previously absent
granted directory. Custom ModSecurity override locations require a drop-in that
grants only their parent directory, followed by systemctl daemon-reload and
systemctl restart csm.service. Atomic replacement requires a directory grant.
See custom account roots for account grants.
The installer generates the unit for the systemd it finds. systemd 239 (EL8,
CloudLinux 8) does not know ProtectHostname, ProtectKernelLogs or
ProtectClock; it would log “Unknown lvalue” for each at every start and
ignore them, so the installer leaves them out there and prints which ones it
omitted. Newer systemd receives the full unit, identical to the packaged one.
A binary-swap upgrade keeps the existing unit; run csm install to regenerate
it.
The daemon cannot write arbitrary files directly under /etc, SSH configuration,
account databases, or systemd unit files. The AF_ALG seccomp drop-in is installed
by the operator’s csm harden CLI, outside the daemon. Periodic AF_ALG module
removal uses a bounded transient service with fixed module names; the daemon
retains ProtectKernelModules=yes and the module syscall restriction.
Exim transactions
The forward guard sends one bounded JSON request over stdin to its own executable
in a transient service. The helper accepts only apply/remove operations and policy
data. It accepts no paths, command names or shell programs. Applying or removing
the managed Exim block, the cPanel rebuild, and rollback all run outside the
daemon mount namespace. Operator configuration is preserved. Lookup refreshes
remain under /var/lib/csm and do not rebuild Exim.
The helper requires root. An exclusive process lock rejects overlapping config transactions without modifying Exim; the daemon reports the failure, and the operator can reload after the earlier transaction finishes. Its lifetime is bounded, and a failed command is never retried as an unsandboxed duplicate. Existing fixed Exim queue/query operations also use transient services. If the systemd bus is unavailable, the existing wrapper attempts direct execution once; inside the service, restricted writes then fail and are reported. A scope unit is unsuitable because it inherits the calling process’s sandbox.
Limits and validation
These restrictions reduce accidental and misdirected writes. CSM remains a privileged root daemon: fanotify, firewall administration, BPF, process inspection and remediation require elevated privileges. The unit retains its current capabilities, and root access to the system bus permits transient services. This is not a complete boundary against arbitrary code execution as root. Reducing capabilities or separating privileged operations behind a separately permissioned broker requires its own feature and kernel compatibility review.
scripts/systemd-account-roots-test.sh runs the actual packaged sandbox in a
disposable systemd environment. It checks unrelated /etc writes return EROFS,
managed directory atomic replacement, ModSecurity overrides, and Exim helper
apply/remove/rollback using a fixed command fixture. It also exercises account
quarantine, restore and process signaling. The Exim fixture validates the helper
transport and filesystem behavior; it does not validate a real cPanel rebuild.
The cPanel release tests exercise real daemon restarts and require a licensed image
and a live run to establish that platform’s acceptance.
Upgrading
Use the signed APT/DNF repository for maintained package upgrades. Standalone
deploy scripts require successful detached signature verification for current
releases. They verify with OpenSSL 3.0+ where available, otherwise with the installed
CSM binary’s csm verify-release, otherwise with python3-cryptography, so
upgrades keep working on EL8/CloudLinux 8 (OpenSSL 1.1.1) including the first
upgrade to a build that provides csm verify-release. When no verifier is
present the upgrade stops; disabling CSM_REQUIRE_SIGNATURES does not enable
unsigned current upgrades. Set CSM_VERIFIER_BINARY to select a specific CSM
binary. See
Release signing for the historical-release exception.
Package installations (recommended)
Use the same signed repository that installed CSM:
sudo apt update && sudo apt install --only-upgrade csm # Debian/Ubuntu
sudo dnf upgrade csm # RHEL family
The package preserves the operator config and state, updates runtime assets, re-signs integrity metadata, and restarts the daemon when it was already active.
Standalone installations
sudo /opt/csm/deploy.sh upgrade
The upgrade refuses to install a release older than the one running. To roll back deliberately, run it with CSM_ALLOW_DOWNGRADE=1.
The helper:
- Downloads and verifies the binary and supporting assets
- Stages the UI, rules, PAM files, and deploy helper before downtime
- Stops the daemon and keeps the previous binary and assets as rollback material
- Activates the staged release and rehashes the config once
- Restarts the daemon, checks sustained liveness, and runs
csm doctor
The health gate waits 20 seconds by default. Set CSM_UPGRADE_HEALTH_SETTLE to an integer from 1 to 3600 seconds for a longer or shorter observation window. Invalid values fail the gate. Doctor warnings do not trigger rollback; failed checks do.
If activation, rehash, startup, or the health gate fails, the helper stops the new daemon, restores the previous binary and assets, re-signs the restored binary hash, and starts the previous version. It exits nonzero and reports where it retained recovery material. If recovery itself fails, it reports an incomplete rollback that needs operator attention.
Optional nightly upgrades
Packages ship a disabled sample at /opt/csm/configs/cron/csm-auto-upgrade. It runs only after an operator installs it into /etc/cron.d/:
csm version && /opt/csm/deploy.sh check
sudo install -m 0644 /opt/csm/configs/cron/csm-auto-upgrade /etc/cron.d/csm-auto-upgrade
Confirm the host runs a published release first. The checksum check reports any different build as an available update. The downgrade guard rejects a lower version, but a development build with the same version can still be replaced by the published release.
The job invokes the standalone deploy helper every night between 03:30 and 04:30, with a random delay and a nonblocking lock. This updates runtime files directly; it does not update APT/DNF’s installed package version. Hosts that need package-manager ownership and version tracking should use their package upgrade automation instead.
Output goes to /var/log/csm/auto-upgrade.log. Failures also produce cron mail to root; configure and test root’s mail alias before enabling the job. This notification does not depend on CSM’s alert channels or on successful daemon recovery. Successful runs and an already-held lock produce no mail. Staggering spreads upgrades over an hour but does not stop a bad release from reaching the fleet.
Troubleshooting
“store: opening bbolt: timeout” – Most operator commands that need live state now route through the control socket at /var/run/csm/control.sock. This error should only appear from commands that intentionally open the bbolt file directly, such as csm store compact, csm store import, csm store reset-bot-verify, csm db-clean --drop-object, or a second daemon start while one daemon already owns the database.
Fix: stop the daemon before direct-store maintenance commands, then retry:
systemctl stop csm
csm store compact
systemctl start csm
If systemctl says CSM is stopped but bbolt still times out, find the process holding /var/lib/csm/state/csm.db and stop that process after review. Do not delete csm.lock; it is only the daemon instance guard and does not release bbolt’s file lock.
“csm: daemon not running” – CLI commands that talk to the daemon exit 2 with this message when the control socket is missing. This includes csm run*, csm check*, csm baseline, csm status, csm firewall ..., csm store export, csm export --since, and csm phprelay .... Start the daemon with systemctl start csm. Bootstrap commands that run before the daemon exists (csm install, csm validate, csm config schema, csm verify, csm rehash) do not require it.
Never delete csm.db – it contains all historical findings, firewall state, email forwarder baselines, and per-account data. If you delete it, the web UI will show empty data until the next full scan cycle (up to 60 minutes for deep scan findings). Restore from backup when possible; for an intentional reset, run csm baseline --confirm rather than removing the database by hand.
Config changes require rehash – After editing a restart-required field in csm.yaml or any conf.d drop-in, run csm rehash once, validate, then restart. Hot-reload-safe changes can use systemctl reload csm; the daemon validates and re-signs the accepted config itself. A restart that fails with conf.d hash mismatch means a drop-in changed since the last signing: rehash if the change was intentional, and list fragments an integration rewrites on its own under confd.integrity_exempt so they stop needing one. csm doctor reports the mismatch while the daemon is still up.
FHS migration (state, config, drop-ins, and profiles)
Current packages use FHS paths for state, config, drop-ins, and shipped profiles. Legacy main configs continue to work during the transition.
| Concern | Legacy path | Current path |
|---|---|---|
| Drop-in fragments | n/a | /etc/csm/conf.d/*.yaml |
| State directory | /opt/csm/state | /var/lib/csm/state |
| Shipped profiles | n/a | /usr/lib/csm/profiles |
| Binary | /opt/csm/csm | /opt/csm/csm (unchanged) |
| Main config | /opt/csm/csm.yaml | /etc/csm/csm.yaml |
| Legacy config path | n/a | /opt/csm/csm.yaml symlink |
The package postinstall creates the FHS directories with the right ownership. If /opt/csm/csm.yaml is a real file and /etc/csm/csm.yaml is absent or still the shipped placeholder, the package copies the legacy config into /etc/csm/csm.yaml and then replaces the old path with a symlink. If both paths are real files with different operator content, CSM refuses the implicit default path until you move one aside or pass --config <path>. Copies that differ only in the integrity hashes CSM writes itself (binary_hash, config_hash, confd_hash) are treated as the same configuration: the daemon starts from /etc/csm/csm.yaml, and the next csm rehash replaces the legacy copy with the symlink.
The daemon copies a non-empty legacy /opt/csm/state/ into the new state directory on first start, but only when the new directory is empty (so a partial migration cannot corrupt it). The legacy directory is left in place; remove it after you have verified the new install.
Operators upgrading by manual binary swap (without re-running the package postinstall) keep the legacy state path if state_path: /opt/csm/state is pinned in the existing csm.yaml. To move state to the FHS layout, either reinstall the package or create the directories by hand and remove the state_path: override.
systemd Type=notify drop-in
The packaged unit file is Type=notify with WatchdogSec=300. The daemon signals READY=1 after watchers attach and pings WATCHDOG=1 on schedule, so systemctl is-active reflects truth and the watchdog kills a hung daemon.
Older units shipped Type=simple. The watchdog still functions because the daemon pings regardless of unit type, but systemctl status only sees the process, not “watchers attached.” If you need the new behavior on an older unit, drop in:
# /etc/systemd/system/csm.service.d/notify.conf
[Service]
Type=notify
NotifyAccess=main
Then systemctl daemon-reload && systemctl restart csm. Verify with systemctl show csm -p Type -p StatusText.
Auto-response dry-run safety default
auto_response.dry_run defaults to true when the key is absent. The daemon records every IP it would have blocked but does not touch nftables. If your auto_response: block sets enabled: true and block_ips: true but does not set dry_run, add dry_run: false explicitly before relying on auto-block. Verify with:
csm status --json | jq '.auto_response_dry_run, .dry_run_blocks'
csm firewall status # check that "Recently Blocked" picks up new entries after the restart
Manual csm firewall ... operations bypass dry-run and always apply.
CLI Commands
Packages and the standalone installer expose /usr/sbin/csm, which points to /opt/csm/csm. Commands that talk to the control socket require the daemon to be running; direct-store maintenance commands say explicitly when it must be stopped.
Global flags
| Flag | Description |
|---|---|
--config <path> | Override the main config path. Default: /etc/csm/csm.yaml, with fallback to /opt/csm/csm.yaml on legacy installs. |
--config-dir <path> | Override the conf.d directory. Default: /etc/csm/conf.d. Wins over CSM_CONFIG_DIR when both are set. Override paths must be absolute, trusted, and not group- or world-writable; loaded fragments must meet the same write-safety check. |
Daemon
| Command | Description |
|---|---|
csm daemon | Run as persistent daemon (fanotify + inotify + PAM + periodic checks). Signals systemd READY=1 after watchers attach and pings WATCHDOG=1 on the configured interval. The systemd notification variables are read once at startup and removed from the environment, so no command the daemon runs inherits them. |
Notification writes time out if systemd stops reading its socket, so they cannot block the daemon indefinitely. ModSecurity rule actions are checked every five minutes using file contents, including linked rule files. Unchanged rule trees skip parsing; empty or incomplete builds keep retrying, and platform detection resumes when the selected directories no longer yield rules. A temporarily empty tree keeps the last loaded actions until a replacement is available.
Checks
| Command | Description |
|---|---|
csm run | Run all checks now via the daemon, send alerts |
csm run-critical | Critical checks now via the daemon (the daemon also schedules critical checks internally every 10 min) |
csm run-deep | Deep checks now via the daemon (the daemon also schedules deep checks internally every 60 min) |
csm check | Run all checks via the daemon, print findings to stdout, no alerts / auto-response |
csm check-critical | Test critical checks only (dry-run via daemon) |
csm check-deep | Test deep checks only (dry-run via daemon) |
csm scan <user> [--alert] | Scan a single cPanel account (capped quick scan). --alert sends alerts for the findings. |
csm scan <user> --full [--wait] [--json] [--respect-ignores] [--quarantine] | Uncapped deep scan of one account, routed through the daemon (bypasses the per-account file cap). It is report-only by default; --quarantine acts on critical file findings the way the scheduled auto-response does (WordPress core, plugin and theme files are cleaned in place; directories, symlinks and lower-severity findings are left for review); --wait polls to completion and prints the report; --respect-ignores honors ignore_paths. |
csm scan --all --full [--wait] [--json] [--respect-ignores] | Server-wide uncapped full scan across every account. Quarantine per account after review (--quarantine is rejected with --all). |
csm scan --status [id] [--json] | List full-scan jobs, or show one job by id. |
csm scan --report <id> [--json] | Print the stored report for a completed full-scan job. |
csm scan --cancel <id> [--json] | Cancel a running full-scan job. |
Management
| Command | Description |
|---|---|
csm install | Deploy config, systemd, auditd rules, logrotate, WHM plugin |
csm uninstall [--purge] | Remove the executable and CSM-owned service, audit, logrotate, webserver, and ModSecurity integrations. Config, drop-ins, state, logs, signature rules, and quarantine data are preserved by default. --purge removes all CSM-owned data. Operator ModSecurity rules are never removed. |
csm baseline | Full server scan via the daemon, records current state for change tracking. Dangerous privileged accounts or WHM root tokens can still be reported on first scan. Takes 5-10 min on large servers. Required on first install. Add --confirm when existing history would be cleared. The daemon must be running. |
csm rehash | Re-sign the binary, csm.yaml and conf.d hashes without scanning. Use once after editing restart-required config, changing a drop-in by hand, or replacing the binary; a stale hash makes the next restart refuse to start. It also applies integrity.immutable to the installed binary. |
csm status | Show current state, last run, active findings, and automation rollout state. Add --json for the full health snapshot (watchers, severity counts, store health, blocklist size, capabilities, version, hashes, automation). |
csm verify-release <key.pem> <file.sig> <file> | Verify a release artifact’s Ed25519 signature without OpenSSL; the supported verification path on EL8 and CloudLinux 8. |
csm systemd-roots | Print a validated systemd drop-in granting write access to configured account trees; see custom account roots. |
csm privileges [--json] [--markdown] | Print every operation that needs privilege beyond reading CSM’s own files, with what it writes and the config key that stops it. Reads nothing from the host, so it answers “what would this do to my server” before installing. See capability matrix. |
csm actions [--since <when>] [--op <id>] [--limit N] [--json] | Print what CSM did to this host: quarantines, cleans, process kills and firewall changes, with the digest before and after a file change. Reads the log directly, so it answers after a crash. See action log. |
csm selftest [--json] | Scan a bundle of samples with known verdicts and report what the installed rules catch, including the gaps the bundle records. Reads no account data. See self-test. |
csm doctor | Config + integrity + daemon + watchers (including the YARA-X scanning worker, which fails while it is down or keeps crashing) + store sanity check, including service write access for custom account roots. Config checks include the firewall lockout warnings for inbound and outbound policy; the integrity check reports a binary, csm.yaml or conf.d hash mismatch before the next restart refuses to start, with the remedy. On CloudLinux with PHP Shield enabled it also reports whether the Shield event mount has been applied to the cages, names up to five missing cages, and gives per-account remount commands. Unresolved accounts appear as uid:N; resolve their account names before remounting. For more than five missing cages, it also offers a remount of all cages during a maintenance window. The correlation attribution line names the checks whose findings sit in the active set without a hosting owner (WARN) or confirms every eligible finding carries one (OK), with the cumulative count since start either way. csm doctor challenge checks challenge public URL, TLS, port gate, webserver snippets, configtest, and the live /challenge/gate endpoint. Add --json for machine-readable output. |
csm validate | Validate config (--deep for connectivity probes) |
csm config show [--no-redact] [--json] | Display config. Secrets are redacted unless --no-redact; --json emits JSON instead of YAML. |
csm config schema | Print a JSON Schema reflected from the Config struct. Use for CI validation of conf.d drop-ins or panel-side editor schemas. |
csm config apply-immutability | Reapply integrity.immutable to /opt/csm/csm. Unsupported filesystems or missing tooling produce a warning and a successful exit; permission and other unexpected errors fail. Package transaction hooks use this after upgrades. |
csm verify | Verify binary and config integrity |
csm version | Version and build info |
csm incidents ... | List, show, and update correlated security incidents (list, show <id>, status <id> <state>, bulk-status). See Incidents. |
csm forensic-snapshot <account> --out <archive.tar.gz> | Evidence archive for incident handoff (triggers, admins, sessions, file mtimes). |
csm webserver-integration <install|upgrade|status|validate|remove> | Install, upgrade, or remove the challenge reverse-proxy snippets for the detected web server. |
csm pam <install|uninstall|status> | Install or remove the pam_csm.so PAM hook (csm pam --help). |
csm report enroll | Generate an abuse-reporting node key pair. |
csm doctor also lists protection queue health:
waiting and running work, losses and lag. A sustained backlog or drop rate
fails the named queue check and includes recovery guidance. Queues whose work
is best effort warn instead of failing, so they never change the exit status.
The same evidence appears in csm status --json and the HTTP status response.
Backup & restore
| Command | Description |
|---|---|
csm backup <path> | Bundle csm.yaml, /etc/csm/conf.d/, and the state directory into a tar.gz at <path>. Runtime lock files, in-progress export staging, and unconfirmed firewall rollback artifacts are omitted. The daemon must be stopped; the command verifies both the control socket and state lock before reading data. |
csm restore <archive> | Validate and stage the complete archive before atomically replacing the live csm.yaml, conf.d, and state directory. Restore rolls back failed replacements, removes stale state entries, and refuses a running or starting daemon. Stop the daemon first. |
csm store export / csm store import (below) is the lower-level alternative: tar+zstd, sha256-verified, finer-grained --only= flags. csm backup/restore is the convenience wrapper most operators want.
Backup and restore use the state lock to exclude daemon access. Restore stages all data beside its destination before replacement, so validation failures leave the live installation unchanged. Uninstall is intentionally non-destructive unless --purge is supplied.
Restore accepts the format produced by csm backup: one tar archive in one gzip member, with no data or padding after the tar end markers and no trailing compressed bytes or additional members. It validates the gzip checksum and length before replacing any destination, and rejects corrupt or truncated archives.
Both commands stream file contents and limit the total uncompressed tar archive to 16 GiB by default, including headers, manifest, and padding. There is no separate limit on a state file. Use --max-bytes <positive byte count> on both commands to raise or lower the budget; restore also accepts older large backups when they fit the selected budget. An over-limit backup fails without replacing an existing archive. Restore keeps its manifest read limited to 64 KiB.
Backup checks space before copying the state database snapshot. Restore checks before each staged file and each replacement copy, including copies onto a different destination filesystem. These checks leave 64 MiB free beyond the file being copied. Allow space for the old installation, the full extracted archive, and a second copy of the restored files during replacement. Space checks cannot reserve storage against other writers; write failures still abort staging before live replacement. If a limit or space check fails, retain the original archive, provision enough storage, and retry with the same or a higher explicit budget.
Hardening
Operator-driven mitigations applied to the host. Run csm harden with no arguments to print the available subcommands on the current host (the audit detects kernel build, panel, and existing mitigations and only offers what’s relevant). Background, full list, and live-detection details: CVE Mitigations.
| Command | Description |
|---|---|
csm harden | Print the hardening menu for this host. |
csm harden --copy-fail | Apply the CVE-2026-31431 (Copy Fail) modprobe mitigation: blacklist algif_aead + af_alg, unload them. Refuses on built-in-AF_ALG kernels. |
csm harden --copy-fail-seccomp | Apply the CVE-2026-31431 seccomp mitigation: write systemd RestrictAddressFamilies=~AF_ALG drop-ins for LiteSpeed, Apache/Nginx, every PHP-FPM pool, cron, and mail units. The right path on built-in-AF_ALG kernels (typical cPanel/CloudLinux 8). |
Remediation
| Command | Description |
|---|---|
csm clean <path> | Clean infected PHP file (backs up original) |
csm db-clean --option <account> <option_name> [--preview] | Sanitize malicious WordPress option values (e.g. injected siteurl / home) |
csm db-clean --revoke-user <account> <user_id> [--demote] [--preview] | Revoke or demote a compromised WordPress admin and invalidate their sessions |
csm db-clean --delete-spam <account> [--preview] | Delete published posts whose title or content matches a high-confidence spam keyword. Matches are confirmed on a Unicode-aware word boundary, so a post is never removed for containing a keyword inside a longer word. Keywords that also occur in ordinary writing are reported by scans but never deleted. |
csm db-clean --drop-object <account> <schema> <type> <name> [--preview] | Drop a MySQL trigger / event / stored procedure / stored function, capturing its CREATE SQL into the db_object_backups bbolt bucket first. <type> must be trigger, event, procedure, or function. <schema> must match a database discovered for <account>. Daemon must be stopped. |
csm virtual-patch [--apply] | Re-scan web roots and preview reversible access-file deny rules. Requires root and manual or auto mode; a timed-out partial scan applies only findings already confirmed reachable and exits nonzero. |
csm enable --php-shield | Enable PHP runtime protection |
csm disable --php-shield | Disable PHP runtime protection |
State database
| Command | Description |
|---|---|
csm store compact | Reclaim unused space in the bbolt state file (atomic rename over the live DB). Requires the daemon to be stopped (systemctl stop csm) because bbolt holds an exclusive file lock while running. |
csm store compact --preview | Snapshot into a temp file next to the live DB and print src/dst sizes without replacing anything. Use to estimate reclaim before scheduling a maintenance window. |
csm store export <path> | Write a tar+zstd backup containing the bbolt store, the state directory, and the signature-rules cache. Private in-progress staging files and unconfirmed firewall rollback artifacts are omitted. A sibling <path>.sha256 companion file holds the archive hash for verification. Daemon must be running. |
csm store import <path> | Restore from a backup archive. Daemon must be stopped. Default restores everything; --only=baseline restores only state JSON files (file hashes); --only=firewall merges only firewall buckets into the existing bbolt; --force-platform-mismatch allows restoring an archive captured on a different OS / panel / web server. |
csm store reset-bot-verify | Drop cached bot PTR verification results and missing-PTR records so the next scan re-runs reverse DNS checks. Requires the daemon to be stopped because bbolt holds an exclusive file lock while running. |
csm export --since <when> | Dump audit-log events for SIEM backfill. <when> is RFC 3339 (2026-04-01T00:00:00Z) or a duration relative to now (24h, 7d). One JSON event per line on stdout, in the same v=1 schema the live audit_log sinks emit. Pipe to a file or directly into a log shipper. Daemon must be running. |
Updates
| Command | Description |
|---|---|
csm update-rules | Download latest signature rules |
csm update-geoip | Update MaxMind GeoLite2 databases |
csm update-bot-ranges | Refresh built-in AI-crawler IP ranges from vendor feeds |
PHP-relay (mail abuse, cPanel only)
Operator controls for the email PHP-relay detector. Talks to the daemon’s control socket; the daemon must be running. See Real-time detection for what the detector fires on, and Auto-response for the freeze action.
| Command | Description |
|---|---|
csm phprelay status | Print the detector’s current state as JSON: enabled, platform, effective dry-run + source (runtime/bbolt/csm.yaml), Path 2b effective account limit, scripts/IPs/accounts tracked, msgID-index size, active ignores. Use to confirm the watcher is wired on a fresh install. |
csm phprelay ignore-script <scriptKey> [--for-hours N] [--persist] [--reason ...] | Suppress all 4 paths for a host:/path scriptKey. Default TTL 168h (7d). --persist writes to the bbolt phprelay:ignore bucket so the suppression survives daemon restarts; without it the entry is in-memory only. <scriptKey> is the value the daemon prints in email_php_relay_abuse findings (e.g. shop.example.com:/wp-admin/admin-ajax.php). |
csm phprelay unignore <scriptKey> [--persist] | Remove an active ignore. --persist also deletes the bbolt row. |
csm phprelay ignore-list | List all active ignores as JSON: scriptKey, expiresAt, addedBy, reason. |
csm phprelay dry-run on|off|reset [--persist] | Override the auto-freeze dry-run state at runtime. on = freeze findings emitted but no exim -Mf runs; off = live freezes; reset clears the runtime override and falls back to bbolt or csm.yaml. Precedence: runtime > bbolt > yaml. --persist writes the on/off choice to the bbolt phprelay:settings bucket so it survives restarts; on reset --persist the bbolt row is also deleted. |
csm phprelay thaw <msgID> | Manually thaw a frozen Exim message. Wraps exim -Mt with msgID validation (rejects anything that isn’t [A-Za-z0-9-]{16,32}) and writes a thaw entry to the auto-freeze JSONL audit at /var/log/csm/php_relay_audit.jsonl. |
Firewall
See Firewall for the full reference.
csm firewall status
csm firewall deny <ip> [reason]
csm firewall allow <ip> [reason]
csm firewall tempban <ip> <dur> [reason]
csm firewall deny-subnet <cidr> [reason]
csm firewall grep <pattern>
csm firewall flush
csm firewall rollback status|confirm|revert
# ...
Real-Time Detection
CSM detects threats in under 2 seconds using kernel and log watchers running inside the daemon: a fanotify file monitor, inotify tailers on auth/access/mail/FTP logs and the Exim spool, and a PAM brute-force listener.
fanotify File Monitor (< 1 second)
Monitors the filesystems containing /home, /tmp, /dev/shm, /var/tmp, configured account_roots, and detected cPanel document roots.
A filesystem mark covers the whole filesystem, including all its bind mounts,
which makes writes inside CloudLinux CageFS visible. Roots on that filesystem
share one mark. A root being a mount point does not restrict a filesystem mark
to that path: symlinks and bind mounts can name only a subtree. On kernels that
fall back to mount scope, the mark covers the whole containing mount and misses
writes through other bind mounts. The startup log describes these scopes and
which roots share a filesystem mark, and says when a watch root sits on the same
filesystem as /, which is when the mark raises an event for every write on the
machine; failures to inspect or mark a root are reported even when other roots
succeed. Missing roots are skipped.
The daemon filters delivered file events before queueing them for analysis. In
the shared temporary trees (/tmp, /var/tmp, /dev/shm) that decision uses
the event descriptor. Non-directory objects with executable permission bits are
queued, as are watched user crontabs, PHP source, configuration files such as
.htaccess, .user.ini or php.ini, known webshell names, files under .config
directories, and images inside hosted trees. Writes needed for self-deleting-file
tracking are also queued. Other writes are dropped without analysis.
Files retained for tracking keep their normal analyzer route,
including suppression and WordPress checksum checks. Staging names written by an
atomic save are resolved first, so a configuration file saved through one is
still analysed. A descriptor that cannot be inspected is analysed rather than
skipped.
csm_fanotify_events_total, csm_fanotify_events_admitted_total, and
csm_fanotify_events_dropped_total distinguish completed admission decisions,
queued events, and events lost to a full analyzer queue. Subtract both admitted
and dropped events from the received total to count events rejected before
queue admission, including those whose file path could not be resolved.
Atomic-save files enter the same bounded worker queue as ordinary writes. Where supported, create events inspect available content; close-write events inspect the completed write. Scans read the original event descriptor even if the file has been renamed, replaced, or deleted before analysis. No rename event is needed to inspect those bytes. Repeated findings use the normal alert cooldown, and queue overflow schedules a directory rescan of files that remain on disk. That recovery runs one pass at a time, under a time budget and with a minimum gap between passes. Unfinished directories stay queued and retry even after new drops stop. Recovery resumes within a directory and retains its original scan window so deferred files do not age out while waiting. Kernel notification loss still relies on the next deep scan. Shutdown drains the queue under a fixed time budget: events still waiting when it expires are released without being scanned; files still on disk are covered by the next deep scan. Both budgets are checked between scans, so a scan or directory read already in progress must still finish before shutdown completes. A rename-only arrival without a usable create or close-write event is also first examined by the rolling content scan; the watcher does not subscribe to rename events.
When a realtime YARA snapshot exceeds the worker message limit, the worker
scans a sealed memory-backed copy of the captured bytes. This preserves the
realtime read window, including a bounded prefix of a larger file, without
reopening the event path. Alert payload details and the retry fingerprint
describe that snapshot; content outside the read window remains subject to
the scheduled scan limits. A failed retry raises yara_realtime_scan_error.
For WordPress atomic saves, the intended basename can select a core or plugin checksum entry. Only a match against the complete event-file content verifies the file; missing checksums, partial content, and modifications proceed through normal detection. The finding retains the actual event path.
A WordPress update is unpacked under wp-content/upgrade/ before it is moved
into place, and each staged PHP file is judged by hash rather than by location.
A file that matches the official wordpress.org checksum for its package version
is stock and skips detection like an installed stock file. A file whose
checksums are still being fetched is content-scanned now and compared once they
land, even if the package header was written late or WordPress has already
moved the tree into place. An installed plugin header can identify the release
only when its directory is the original staging directory moved into place.
A core update copies files, so the installed version.php identifies the
release only if it was replaced or rewritten after the staged file was seen;
a borrowed release still requires the file’s bytes to be the official ones for
that path. A tree removed before its release could be identified gets a
package warning, not file-by-file comparisons against the old release. This
covers refused updates (for example when the new release needs a newer PHP)
and plugin copies whose staging identity was never captured; the warning does
not establish whether installation succeeded. Reused staging paths follow the
normal alert cooldown rather than suppressing later unidentified uploads. A
file the official package does not ship gets its
own warning, naming the installed path when it still exists. Themes, packages
not published on wordpress.org, a full verification queue, and a package whose
checksums do not arrive within 60 seconds raise one warning per staging
directory. WordPress picks a new staging directory for every upload. Once the
package type and version are known, these warnings are identified by site,
package type, name, version and reason. Uploading the same declared release to
the same site again is a repeat and follows the normal 24-hour reminder. This
groups checksum-availability warnings; it does not establish that the package
contents are unchanged or safe. Incomplete headers keep an upload-specific
warning, and a later change of type, version or reason can raise a new warning
in the same directory. Per-file checksum warnings repeat across uploads only
when their package identity, relative path and complete digest match; files
that could not be hashed keep their upload-specific warning. Content findings
inside staging keep their own per-file identity and normal alert cooldowns;
package-warning dismissal does not dismiss them.
Location warnings for PHP in sensitive WordPress directories are suppressed only for recognized inert content, including comment-only stubs and literal translation or version data. PHP attributes are executable syntax, not comments, and do not qualify a file for that suppression. Content scans still run.
Detects:
- Webshell creation (PHP files in web directories)
- Self-deleting droppers: a PHP or executable created under a document root and unlinked within
thresholds.dropper_unlink_ttl_sec(default 300s), the loader technique that creates a rogue admin then erases itself before any scan. A file whose exact bytes WordPress moved or copied into place before removing the original is not reported: plugin, theme and core packages, language packs, and the version file a core update reads first. Language-pack and version-probe copies also require a complete, stable data-only snapshot with no suspicious content or earlier unsafe write. A core update that stops after reading its version file leaves the old release installed, so no installed copy matches; that version file is not reported when its bytes equal the wordpress.org checksum for the release and locale it declares. Checksum downloads start in the background when capacity permits, and a file still unverified at the deletion probe is reported. Other executable files do not qualify for this copy exception. The script Really Simple Security copies into uploads to test whether PHP runs there, then deletes, is not reported when its content is byte for byte the shipped script; any other content under that name is. A core update unpacks the whole release but installs only what changed and never overwrites the bundled themes and plugins; a staged file it deletes that way is not reported when its bytes equal the wordpress.org checksum for that path in the release now installed. The checksum must describe the same complete, stable snapshot as the retained content. Earlier nonempty content must also be known and unchanged, including content observed at creation; later official bytes or an identical installed copy cannot erase that evidence. The same history rule applies to every staged core, plugin and theme file: a file whose content changed after it was first written is reported even when WordPress then moves it into place or an identical copy is installed. Repeated analysis of the same snapshot does not count as another write. A single snapshot whose complete bytes could not be read can still match a move by file identity; a second nonempty snapshot leaves the unknown earlier bytes unaccounted for. Files too large for a complete snapshot retain their leading bytes for content checks. Other files removed from staging without an identical installed copy are still reported. Upgrade staging, atomic-save temp files, template compile caches, BackWPup job-state files (the running job’s JSON and its folder list, written behind a PHP comment, including a captured head cut inside a CRLF line ending), a path taken over by a newer file, and a file whose original directory was removed are recognized and reported at a lower severity; files removed together with the directory holding them, such as a language pack WordPress unpacked and discarded, are reported once for that directory; a create/delete burst collapses into one lower-severity notice. Grouping is recalculated after ignored paths are removed: remaining files can form directory groups, while a lone remaining file is reported on its own. Off withthresholds.dropper_detection: false. Candidate tracking and findings held for aggregation each accept up to 16,384 entries; additional work is refused without evicting older evidence. Status and doctor report refusals, exhausted probes, overdue work and stalled processing through protection queue health. - Dropper admission and pending findings follow the live
suppressions.ignore_pathslist after reload. Intentionally suppressed candidates consume no tracker capacity; losses of eligible candidates still raise the capacity warning. - PHP in uploads, languages, upgrade directories
- PHP in
.ssh,.cpanel, mail directories (critical escalation) - Executable drops in
.config .htaccessinjection and tampering (auto_prepend, eval/base64 handlers, CGI execution remaps, and ModSecurity disablement).user.initampering andphp.initampering under configured or detected web roots- Obfuscated PHP (encoded, packed, concatenated)
- Fragmented base64 evasion (
$a="base"; $b="64_decode"– function name split across variables) - Concatenation payloads (hundreds of
$z .= "xxxx"lines with eval at end) - Tail scanning: payloads appended to the end of large legitimate PHP files (beyond the 32KB head window)
- CGI backdoors: Perl, Python, Bash, Ruby scripts in web directories (e.g., LEVIATHAN toolkit)
- SEO spam: gambling/togel dofollow link injection in PHP/HTML files
- Phishing pages and credential harvest logs
- Phishing kit ZIP archives
- PHP carried inside a file served as an image. Image writes under an account or configured document root are inspected, including hosted
.configdirectories and roots located in temporary directories. Images up to 128 KiB are inspected in one piece; larger files get a 64 KiB head and a 64 KiB tail scan. Payloads outside those windows require a deep scan, subject tothresholds.full_scan_max_file_mb. A PHP opening tag alone is not enough – a screenshot quoting one in its description chunk stays quiet – and a file that only wears an image name while holding PHP source is reported the same way. Finding:php_in_image_realtime. - A PHP file that includes or requires an image, archive or other non-executable file while also reading request input. That pairing is the loader half of the technique above: the payload lives in the picture and the one-line loader elsewhere. Literal targets ending an include statement, concatenated paths, paths held in a local, encoded paths and suppressed
@includestatements match in either statement order. Ordinary templating that pulls in.html,.tpl,.txtor.svgpartials is outside the extension-based rule. Encoded targets still count because the source hides their extension. - YAML signature matches (PHP, HTML, .htaccess, .user.ini, php.ini)
- YARA-X rule matches (if built with
-tags yara)
Completed upload execution probes remain tracked until the deletion check so combined or out-of-order create and close-write events preserve completion evidence. A probe without a completed write, or with earlier unsafe or uncertain content, remains reportable. Image writes use the existing notification-only fanotify instance, its 16,384 event queue and 4-16 analyzer workers. They create no permission-event holds and do not enter the mail scanner. Thumbnail and WebP bursts fill the same queue as other writes: excess events are closed immediately and their parent directories enter the existing capped recovery tracker. Recovery rescans recent surviving files; kernel queue loss and exhausted recovery coverage still rely on the next deep scan. Queue health and content-read truncation remain visible through the existing metrics. Mail permission holds retain their separate hold budget and watchdog.
Both the real-time and the scheduled WordPress admin-creation signature require an administrator role token plus literal or request-derived credentials, and both accept the same ASCII whitespace, including the vertical tab. An importer that creates users with a generated password does not trigger either one just for reading a login from an import form.
Two divergences between the engines are known and left in place. Go folds the Unicode long s into ASCII s and YARA’s nocase does not, so a token spelled with it can match in real time where a scan declines. Bounded expressions also count differently: the real-time engine counts characters and the scheduled one counts bytes. Neither can occur in real PHP, and both belong in the scanner’s regex compilation rather than in hand-written escapes inside every rule. Strictness parity is not enforced across the rulesets: the parity check compares rule names, not how strict each side is.
Complete blank files are excluded from dropper alerts after a close-write observation. Any content or metadata change during the read, including an unlink, leaves the snapshot uncertain. WordPress data and its checksum are read within the same stat interval. A later complete close-write snapshot cannot rule out code that ran during an earlier unstable read. That uncertainty and code seen in any read remain attached to the tracked file. Concurrent firewall-log writes can therefore still produce a warning when the log is replaced, or a critical finding when it is deleted.
Core-release checksum downloads are non-blocking and bounded, including their retry timers. Repeated misses and streams of distinct release names share an admission budget. Release identifiers longer than 64 bytes are rejected before retention or lookup. When that budget is exhausted, cached checksums remain usable; an uncached file stays unverified and receives the normal dropper assessment. A later observation can request checksums again after the budget recovers.
PHP files are also excluded when their first statement stops the interpreter
(exit, die, or __halt_compiler, with at most a plain literal argument and
no preceding comment). Plugins keep state and firewall data in files of that
shape and rewrite them constantly, and the bytes behind the terminator are
never compiled. The exemption is refused when those trailing bytes are made
only of transport-encoding alphabet: an operator who turns on PHP source
conversion can have a file decoded before it is tokenized, which would make an
encoded tail the program and this header padding. Comment-bearing PHP remains
eligible for the same reason – a comment’s tokens depend on the interpreter’s
source encoding. Executable-mode files are not judged by PHP compilation rules,
because a shell may read them instead. Content or signature findings and
previously observed code always override inert-content filtering.
Features:
- Per-path alert deduplication (30s cooldown)
- Process info enrichment (PID, command, UID)
- Auto-quarantine on high-confidence realtime signature matches (category, size, entropy, and hex/execution validation)
Sensitive System File Watcher
Tracks a fixed set of system-configuration paths: the account and credential databases, sudoers and its drop-in directory, the SSH daemon config and its drop-in directory, the system cron drop-in directories, and per-user crontabs. The set is not operator-configurable – a path an attacker knows is excluded is a free landing pad.
The backend is chosen by detection.sensitive_files_backend. The bpf backend attaches an LSM hook that catches writes as they happen, and re-expands the watchset every detection.sensitive_files_poll_interval (default 5m). The legacy backend content-hashes the watchset on the same interval.
Both backends key on the path, not the inode. Most tools replace a config file by writing a temporary file and renaming it over the target, which gives the path a new inode and hides the write from the LSM hook. The BPF refresh therefore compares content and security metadata per path: a rename-over is reported as a change on the existing file, only a path no previous refresh had seen is reported as newly appeared, and a regular-file rewrite that leaves bytes, permissions, and ownership unchanged is not reported. It also tracks symlink targets, so retargeting a watched path remains visible even when both targets contain the same bytes. Non-regular entries are tracked by path identity without opening them for content reads. The legacy poller compares content digests by path.
Findings are sensitive_file_modified. A write finding successfully delivered by the LSM hook wins only while the path still has the exact state captured for that finding. The refresh reports a later rename-over separately, retries a refresh finding if the alert queue was full, and always evaluates the bytes used for its digest rather than a later read. Writes inside a package-manager window, or by a process whose ancestor is a package manager or the control panel’s own maintenance tooling, are demoted to Warning unless a cron payload carries obvious persistence tokens. This veto applies to live writes as well as added and modified cron files found by polling. It includes the shared crontab persistence patterns and their bounded base64 inspection, in addition to suspicious shell fragments. Control-panel ancestry is read from the executable a parent process actually runs, not from the name it reports, and the executable must be a root-owned file inside the panel’s own directory that no other user can write. Cached executable evidence must come from a resolved procfs link; an exec-event filename or script argument does not qualify. Ancestry is available on every Linux build before monitoring starts. A cache hit avoids a live ancestry walk; when the cache proves nothing, CSM falls back to procfs. CSM’s own managed writes are suppressed by content, not by path.
inotify Log Watchers (~2 seconds)
Tails auth, access, and mail logs in real-time. The exact file paths are chosen per platform at daemon startup – see the platform: ... line in the daemon log.
| Log | Platforms | What it detects |
|---|---|---|
cPanel session log (/usr/local/cpanel/logs/session_log) | cPanel only | Logins from non-infra IPs, password changes, File Manager uploads |
cPanel access log (/usr/local/cpanel/logs/access_log) | cPanel only | cPanel-API auth patterns |
| Auth log | All | SSH logins and failures. /var/log/auth.log on Debian/Ubuntu, /var/log/secure on RHEL family and cPanel |
Exim mainlog (/var/log/exim_mainlog) | cPanel; non-cPanel when the file exists | Mail anomalies, queue issues, SMTP brute force, probe abuse, and cloud relay abuse |
| Apache/LiteSpeed/Nginx access log | All | WordPress brute force (wp-login.php, xmlrpc.php), real-time. Paths: /var/log/apache2/access.log (Debian), /var/log/httpd/access_log (RHEL), /var/log/nginx/access.log (Nginx), /usr/local/apache/logs/access_log (cPanel) |
| Mail log (platform file or journal) | All hosts with Postfix/Dovecot logs | IMAP/POP3/ManageSieve account compromise and mail brute-force |
FTP log (/var/log/messages) | cPanel only | FTP logins and failures |
| ModSecurity error log | All (if ModSec installed) | WAF blocks and attacks. Auto-discovered from the detected web server |
Nginx error log (/var/log/nginx/error.log) | Nginx hosts | General web errors, ModSecurity denies |
Successful FTP logins over loopback do not raise an unfamiliar-address warning. Failed authentication remains reportable over loopback, including through local relays.
The FTP and SSH log watchers read the same files as the periodic ftp_logins
and ssh_logins checks, so both see every login. Each login is reported once:
the watcher and the periodic check build the same finding, and the second one
is recognised as a repeat. When the watcher is not running, on a non-cPanel
host or before the log file appears, the periodic check still reports the
login on its own.
The identity uses the complete log record, including its timestamp and session fields, even when the displayed details are shortened. A later session from the same address is a new finding within the 24-hour reminder window.
A successful FTP login and a cPanel File Manager write are not emailed. On shared hosting every customer connects from an address that is not infrastructure, so both fire on ordinary use of a core feature. They stay on the findings page, in history, in incident correlation and in the attack database, where their value is correlation with other evidence on the same account. Failed FTP authentication, FTP brute force, a login from a brute-force source, and SSH logins are all still emailed.
The phpanel webhook and SSE event stream still receive these successful operations; the email and operator-webhook filter runs after data delivery.
Login check names on upgrade
New findings use ftp_login in place of ftp_login_realtime, and
ssh_login_unknown_ip in place of ssh_login_realtime. Update external
phpanel, export and SIEM rules to accept the merged names. Historical records
and already queued deliveries can still contain the old names; the JSON schema
is unchanged.
Existing alerts.email.disabled_checks entries and saved suppression rules for
either retired name also match its replacement. The Settings page displays and
saves the current names. A mute for one former producer now covers the shared
finding from both producers; it does not mute FTP brute-force escalations.
Pending SSH blocks and restored SSH incident evidence retain their blocking policy. Historical successful FTP activity remains excluded from incident blocking, including events carrying its retired name.
cPanel-only log watchers are not registered on non-cPanel hosts, so you will not see “not found, retrying every 60s” warnings for them on plain Ubuntu or AlmaLinux.
The Postfix/Dovecot file reader polls every two seconds. It reads replacement
files from the start and rewinds when the current file shrinks below its read
position. Truncation also clears buffered bytes from the previous file contents.
With copytruncate, a file that regrows past that position between polls can
hide the truncation and lose events. Use rename/create rotation with the log
writer reopening its file, or journal input, when that loss is unacceptable.
Mail records are emitted only after their newline arrives. A partial record survives temporary EOF up to the 64 KiB limit, including its newline. Longer records are discarded through the next newline even when written across several polls. Rotation and detected truncation clear pending record state.
Mail source attachment retries after failures, starting at one second and doubling to a maximum delay of 30 seconds. The watcher remains unhealthy and reports an unavailable-source finding until a reader attaches successfully. Repeated identical errors are not re-emitted. Retries use the current mail source configuration and start at the current tail, so delayed attachment does not count historical authentication failures as new activity.
With mail_logs.source: auto, each retry chooses the configured or platform
file if present, otherwise the configured journal units. A file missing for
90 seconds triggers a new selection. The old reader stops before a replacement
starts. Journal input follows new records from the selected services, including
services with no prior entries; it does not replay older records on attachment.
Explicit file and journal modes retry their selected source without
switching, and a working reader stays attached until it stops or loses its file.
Journal input requires a build with journal support.
SMTP / Dovecot Brute-Force Tracker
Detects credential stuffing, password spray, and raw SMTP probe storms. Runs as part of the Exim mainlog watcher on cPanel hosts and on non-cPanel Exim hosts where /var/log/exim_mainlog exists.
Four attack patterns:
| Signal | What triggers it | Auto-response |
|---|---|---|
smtp_bruteforce | A single attacker IP exceeds the per-IP failed-auth threshold within the configured window | IP blocked via nftables |
smtp_probe_abuse | A single attacker IP exceeds the raw SMTP connect-rate threshold before AUTH | IP blocked via nftables |
smtp_subnet_spray | Multiple distinct attacker IPs from the same /24 subnet exceed the subnet threshold | Entire /24 subnet blocked via nftables |
smtp_account_spray | Many distinct attacker IPs targeting the same mailbox exceed the account threshold | Visibility finding only. No auto-block, because attackers span many subnets and no single-IP action helps |
Tunable via the thresholds.smtp_bruteforce_* and thresholds.smtp_probe_* keys in csm.yaml. Infrastructure IPs (from infra_ips) are never counted or blocked.
Cloud-Relay Credential Abuse
Detects authenticated outbound Exim deliveries where the same mailbox is sending through public-cloud relay sources. The realtime Exim mainlog watcher evaluates new accepted deliveries, and a bounded startup replay covers recent lines already on disk.
The finding is email_cloud_relay_abuse. Auto-response actions follow the global dry-run and block settings plus the email hold path. Operators with legitimate cloud mailers can opt out specific mailboxes or domains under email_protection.cloud_relay, or use email_protection.high_volume_senders for known high-volume senders.
Mail Auth Brute-Force Tracker
Detects credential stuffing and password spray against IMAP, POP3, and ManageSieve. Runs through the mail_logs reader: file source uses /var/log/mail.log on Debian-family hosts and /var/log/maillog on RHEL-family and cPanel hosts, while journal source reads configured Postfix/Dovecot units. The wrapper composes with the existing geo-based login monitor, so email_suspicious_geo keeps firing for successful logins from novel countries.
Five attack patterns:
| Signal | What triggers it | Auto-response |
|---|---|---|
mail_bruteforce | A single attacker IP exceeds the per-IP failed-auth threshold within the configured window without matching successful mailbox activity | IP blocked via nftables |
mail_bruteforce_suspected | An established good source hits the per-IP failed-auth threshold with a confined stale-password pattern | Visibility finding only. No auto-block |
mail_subnet_spray | Multiple distinct attacker IPs from the same /24 subnet exceed the subnet threshold | Entire /24 subnet blocked via nftables |
mail_account_spray | Many distinct attacker IPs targeting the same mailbox exceed the account threshold | Visibility finding only. No auto-block, because attackers span many subnets and no single-IP action helps |
mail_account_compromised | A successful login comes from an IP that repeatedly failed auth against the same mailbox | Critical findings block immediately. A source with established history on at least two other mailboxes emits a High advisory and is not auto-blocked |
Tunable via the thresholds.mail_bruteforce_* keys in csm.yaml. Independent from the SMTP tracker so the Dovecot noise floor can be tuned separately. Infrastructure IPs are never counted or blocked. Established good sources with a confined stale-password pattern emit mail_bruteforce_suspected instead of an IP block, while wider spraying and compromise from a source without established multi-mailbox history still block. A compromise from an IP with established history on at least two other mailboxes remains visible as a High advisory but does not feed direct, incident, or spray auto-blocking. The mail_bruteforce and mail_bruteforce_suspected alerts name the mailboxes the source was hitting, with per-mailbox failure counts and a count of attempts that named no mailbox, so a real attack on a live mailbox is easy to tell apart from dictionary noise.
When the mail authentication backend itself fails (for example dovecot cannot reach cpdoveauthd), every login fails regardless of password. CSM detects a burst of these backend errors, pauses mail_bruteforce and mail_subnet_spray auto-blocking, and raises a mail_auth_backend_degraded warning so the outage is visible instead of mass-blocking legitimate users. Detection resumes automatically once the backend recovers.
Admin-Panel Brute-Force Tracker
Counts repeated POST requests to high-value non-WordPress admin login endpoints. Runs as part of the web access-log watcher.
Covered endpoints (tight set to avoid false positives on shared hosting):
- phpMyAdmin:
/phpmyadmin/index.php,/pma/index.php,/phpMyAdmin/index.php - Joomla:
/administrator/index.php
When an IP crosses the POST-rate threshold, admin_panel_bruteforce fires and the attacker IP is auto-blocked.
Drupal /user/login and Tomcat Manager /manager/html are intentionally out of scope here. Drupal’s path is too generic on shared hosting, and Tomcat Manager uses HTTP Basic auth (repeated GET requests with 401 responses), not POST form submissions. Both need different detectors and are tracked as follow-up work.
PHP-Relay (Mail Abuse, cPanel Only)
Real-time inotify watcher on /var/spool/exim/input catches WordPress contact-form spam relays where an attacker uses PHPMailer (or similar) with a spoofed From, an external Reply-To, and a script URL that doesn’t belong to the cPanel account. The occonsultingcy incident (2026-04) drove the design: a legitimate site running a vulnerable contact-form plugin became a per-message spam relay through the operator’s own mail account.
The detector runs four paths and only fires email_php_relay_abuse (Critical) when one of them crosses threshold. Paths 1 and 2 are scoped per-script, using the host:/path from the X-PHP-Script Exim header. Path 2b is per cPanel user. Path 4 is per HTTP source IP across distinct scripts. Paths 2 and 4 use a recipient-diversity gate that suppresses only known low-recipient notification mail.
| Path | What triggers it | Why it exists |
|---|---|---|
| Path 1: header score | Per-script: From domain not in the account’s authorised domains AND additional signal (PHPMailer / suspicious Reply-To / suspicious User-Agent), evaluated over a rolling 5-min window once the script has emitted at least header_score_volume_min messages | The shape that matched the original incident: spoofed sender, contact-form-style. FromMismatch is a HARD precondition – the score never accumulates without it |
| Path 2: absolute volume per script | A single script emits at least absolute_volume_per_hour messages in the last hour. If Exim envelope recipients are known, fewer than fanout_distinct_recipients distinct recipients suppresses this path; unknown recipients or fanout_distinct_recipients: 0 fail open. | Catches a compromised script even if the headers themselves are legit-shaped, while leaving plugin notification mail to a fixed admin set alone |
| Path 2b: account log-tail volume | Per cPanel user: more than effective_account_limit outbound messages through the redirect_resolver router in the last hour. The effective limit is auto-derived from /var/cpanel/cpanel.config’s maxemailsperhour (60% of it, clamped to 20-60), capped at 95% of the cPanel limit when an operator override is set | Backstop for when Path 2 misses the window. Reads /var/log/exim_mainlog directly; only fires on lines tagged B=redirect_resolver so forwarders don’t trip it |
| Path 4: HTTP-IP fanout | Per HTTP source IP: one source IP appears in at least fanout_distinct_scripts distinct script keys in fanout_window_min minutes, after excluding loaded HTTP-proxy ranges, loopback, and the host’s own interface addresses. If Exim envelope recipients are known, fewer than fanout_distinct_recipients distinct recipients suppresses this path; unknown recipients or fanout_distinct_recipients: 0 fail open. | Catches one client walking many scripts while avoiding CDN/proxy traffic, local cron or panel callbacks, and fixed-admin notification fanout |
Path 5 (behavioural baseline) is deferred to Stage 2.
The detector starts a one-shot retrospective scan of exim_mainlog at daemon startup so Path 2b can fire on history already on disk. IN_Q_OVERFLOW triggers a bounded recovery walk of the spool (capped at 1000 files; if more were skipped, a email_php_relay_overflow_scan_truncated Critical fires too – Path 2b backstops the missed messages).
Operator suppressions (csm phprelay ignore-script <host:/path>) short-circuit the pipeline before any path scoring runs, so a known-noisy contact form can be opted out individually without disabling the detector. See PHP-relay CLI for the full operator surface.
PAM Brute-Force Listener
Real-time authentication monitoring across all PAM-enabled services.
- SSH login tracking with geolocation
- cPanel, FTP, and webmail authentication
- Credential stuffing / password spray breadth: one source IP failing against many distinct accounts inside
thresholds.multi_ip_login_window_min. The finding iscredential_stuffing; tune the account floor withthresholds.cred_stuffing_distinct_accounts(default 5). - Blocks IPs within seconds of threshold breach
- Integrates with the nftables firewall for instant blocking
Process Context
Exec and outbound-connection findings carry an optional process object with
PID, PPID, UID, user, cPanel account (when known), comm, exe, sanitized cmdline,
and a parent chain up to depth 5. The chain is materialized from an in-memory
LRU+TTL cache (cap 16384 entries, 30-minute TTL) populated from BPF exec
events. Cache misses trigger a bounded async /proc read, so process-context
enrichment does not add blocking work to the connection event loop. When
neither cache nor enricher has data (e.g., a process that exited before
userspace reads its event), the process field is omitted entirely and the
finding still emits.
Counters exposed at /metrics:
csm_process_context_cache_entriescsm_process_context_cache_evictions_total(LRU)csm_process_context_cache_ttl_purges_totalcsm_process_context_cache_misses_total(includes TTL purges)csm_process_context_enrich_queue_drops_totalcsm_process_context_enrich_reads_totalcsm_process_context_enrich_errors_totalcsm_process_context_enrich_stale_totalcsm_process_context_enrich_latency_seconds
Caveats:
started_atis emitted only when the event source supplies a trustworthy start timestamp. CSM does not infer it from procfs directory metadata.- After daemon restart, the
csm_process_context_enrich_*counters may show a smallenqueued - readsdelta. Pending requests in the enricher queue are dropped on shutdown by design. - Hosts without BPF support fall back to
/proc/net/tcp[6]polling. That path has no PID, so emitted findings do not carry aprocessfield.
HTTP Flood, Scanner Profile, UA Spoof, and Distributed Flood
http_request_flood, http_scanner_profile, http_claimed_bot_unverified, http_ua_spoof, and http_distributed_flood are periodic, not real-time. They run inside the same wp_bruteforce scheduled check that scans per-vhost access logs every 10 minutes. A real-time inotify tailer would need to hold per-IP state across log rotations and is out of scope for the initial release (see the plan non-goals). For attack types where sub-minute response matters, the access-log inotify watcher already covers wp_login_bruteforce and xmlrpc_abuse; the periodic scan adds volume-based rate enforcement, pending claimed-bot handling, scanner-profile detection, and per-vhost distributed attack rollups on top.
A verified crawler is dropped from the scan before any counter increments: a source IP whose claimed bot User-Agent passes IP-range or reverse-DNS verification cannot contribute to a flood, scanner, or spoof finding, so verified Googlebot and AI crawler traffic does not produce false positives. Verification is covered in Threat intel.
| Finding | Fires when | Key gates (defaults) |
|---|---|---|
http_request_flood | One source IP makes too many requests inside the rate window | http_flood_threshold requests within http_flood_window_min (5 min). Ships disabled (0) so operators sample baseline traffic first |
http_scanner_profile | One source IP’s in-window traffic is almost all probe-error responses spread across many distinct paths, the shape of URL enumeration hunting for backups, downloadable files, and dormant shells | Three gates must all pass: http_scanner_min_requests volume (ships 0, disabled), at least http_scanner_error_pct (90) of requests on a probe-error status, and http_scanner_min_distinct_paths (10) distinct error paths. Probe-error statuses default to 404 and 403; query strings are stripped so cache-buster URLs on one missing endpoint count once |
http_ua_spoof | One source IP sends non-browser User-Agents | Known scanner agents (nikto, sqlmap, nmap, wpscan, nuclei, and similar) fire on the first hit. Claimed crawler UAs with a cache-confirmed reverse-DNS negative, scripting (curl/python/wget), headless (Puppeteer/Playwright), and empty agents fire at http_ua_spoof_threshold (30); scripting/headless/empty still require their opt-in flags (http_ua_scripting_enabled, http_ua_headless_enabled, http_ua_empty_enabled) |
http_distributed_flood | Many distinct already-abusive source IPs hit one vhost in a single scan window | Opt-in: fires once http_distributed_min_ips distinct IPs (sample 10), each having already crossed a per-IP abuse threshold above, hit the same vhost. Built only from IPs that tripped another finding, so a popular site’s normal visitor spread does not trip it |
Per-IP findings roll up per vhost, so a confirmed scanner that sprays a few probes across many vhosts still feeds the distributed rollup. Full threshold reference and tuning notes are in Configuration.
Direct SMTP Egress
Outbound connections to SMTP ports from non-MTA local processes
emit a direct_smtp_egress finding. See
Direct SMTP egress for the full rule set,
config schema, and metric.
Critical Checks
Critical checks run every 10 minutes. Typical wall-clock cost on a busy shared host is a few seconds; the runner enforces the 10-minute cadence even when a tick takes longer.
The tables below name the finding identifiers CSM emits, grouped by area. Most are standalone scheduled checks; a few (for example the http_* traffic findings and bad_asn_outbound) are produced by a parent scheduled check – the access-log brute-force scan and the periodic connection scan respectively – rather than being independently scheduled.
Process & System
| Check | Description |
|---|---|
fake_kernel_threads | Non-root processes masquerading as kernel threads (rootkit indicator) |
suspicious_processes | Reverse shells, interactive shells, GSocket, suspicious executables |
php_processes | PHP process execution, working dirs, environment variables |
shadow_changes | /etc/shadow modification outside maintenance windows |
uid0_accounts | Unauthorized root (UID 0) accounts |
kernel_modules | Kernel module loading (post-baseline) |
af_alg_socket_use | AF_ALG socket use that may indicate Copy Fail exploit activity |
af_alg_enforcement | AF_ALG hardening policy drift and correction status |
SSH & Access
| Check | Description |
|---|---|
ssh_keys | Unauthorized entries in /root/.ssh/authorized_keys |
sshd_config | SSH hardening (PermitRootLogin, PasswordAuthentication, etc.) |
ssh_logins | SSH access anomalies with geolocation |
api_tokens | cPanel/WHM API token usage |
whm_access | WHM/root login patterns, multi-IP access |
cpanel_logins | cPanel login anomalies, multi-IP correlation |
cpanel_filemanager | File Manager usage for unauthorized access |
Network
| Check | Description |
|---|---|
outbound_connections | Root-level outbound to non-infra IPs (C2, backdoor ports) |
user_outbound | Per-user outbound connections (non-standard ports) |
bad_asn_outbound | Outbound connection whose destination resolves (via GeoLite2-ASN) to a bad or unexpected autonomous system. Config detection.bad_asn_outbound: blocked_asns (always bad) and/or allowed_asns (allowlist mode – anything outside is bad). Classified for every process including root (the periodic connection scan); non-root connections are also flagged in real time by the live BPF tracker. Off by default; the third leg of the host_takeover incident chain |
dns_connections | DNS exfiltration and suspicious queries |
firewall | Firewall status and rule integrity |
Brute Force & Auth
| Check | Description |
|---|---|
wp_bruteforce | WordPress login brute force (wp-login.php, xmlrpc.php) |
http_scanner_profile | Source IPs whose traffic is mostly probe-error responses across many distinct paths in one scheduled scan window |
http_claimed_bot_unverified | High-volume claimed crawler traffic while reverse-DNS verification is still pending; challenge-routed when the challenge subsystem is enabled |
http_ua_spoof | IP hitting the per-IP spoof threshold with a search-engine bot UA that fails reverse-DNS verification, or with scripting/headless/empty UAs when those opt-in flags are enabled |
http_distributed_flood | Many already-abusive HTTP source IPs hitting the same vhost in one scheduled scan window |
http_asn_crawl | A single ASN distributing expensive, uncacheable requests across many addresses while one account’s PHP worker pool is saturated. Reverse-proxy ASNs and operator allowlists are excluded; Critical findings can carry the specific source CIDRs eligible for temporary blocking. |
ftp_logins | FTP access patterns and failed auth |
webmail_logins | Roundcube/Horde access anomalies |
api_auth_failures | API authentication failure patterns |
| Check | Description |
|---|---|
mail_queue | Mail queue buildup (spam outbreak indicator). Exim aborts when it cannot write its own log, and the daemon’s sandbox makes /var/log read-only, so the queue query runs as a transient unit forked by PID 1. Hosts without systemd-run query exim directly. |
mail_queue_unavailable | The queue depth could not be read, so buildup detection is inactive. Reported instead of assuming the queue is empty. |
mail_per_account | Per-account email volume spikes |
All Exim queue reads and actions started by the daemon run as transient services outside its read-only filesystem sandbox. This includes queue composition, safe backscatter flushing, and PHP relay freeze or thaw actions. If systemd-run is unavailable, CSM runs them directly; a host without Exim skips the Exim-only queue check. The target command is attempted at most once, and cancellation during the wrapper probe stops it before execution.
Data & Integrity
| Check | Description |
|---|---|
crontabs | Suspicious cron jobs and scheduled commands |
mysql_users | MySQL user accounts and privileges |
database_dumps | Database exfiltration attempts |
exfiltration_paste | Connections to pastebin/code-sharing sites |
The MySQL audit exempts mysql@localhost only when MariaDB confirms its stock
socket authentication setup with password authentication disabled. Modified
authentication, other host literals, and accounts whose setup cannot be verified
remain reportable.
Threat Intelligence
| Check | Description |
|---|---|
ip_reputation | IPs against external threat databases and optional rspamd history. Passive HTTP/cPanel sightings are High; SSH and mail-auth activity is Critical |
local_threat_score | Aggregated score from internal attack database |
modsec_audit | ModSecurity audit log parsing |
The local threat score retains evidence attributed to the server itself. Such records can represent forwarded attacks or a compromised local process; firewall protection against blocking the server does not establish that traffic is safe.
Performance
| Check | Description |
|---|---|
perf_load | CPU load average thresholds |
perf_php_processes | PHP process count and memory |
perf_memory | Swap usage and OOM killer activity |
Health
| Check | Description |
|---|---|
health | Daemon health, binary integrity, required services |
Platform Support
Runs on every supported platform unless noted below. The daemon auto-detects OS and panel at startup and silently skips cPanel-specific checks on plain Linux hosts (no “not found” spam).
cPanel-only (skipped on plain Ubuntu/AlmaLinux):
api_tokens,whm_access,cpanel_logins,cpanel_filemanager– read WHM API and cPanel session logswp_bruteforce– iterates/home/*/public_html/*/wp-login.phpand per-domain access logs. The domlog pass ranks recent logs first and honorsthresholds.domlog_max_files,thresholds.domlog_tail_lines, andthresholds.domlog_max_age_min.webmail_logins– parses cPanel Roundcube/Horde logsmail_per_account– reads/var/log/exim_mainlog
Plain Linux equivalents that still provide coverage:
mail_queueruns on any host where Exim is installed and is skipped when there is no Exim queue.- Access log brute-force detection (
wp_login_bruteforce,xmlrpc_abuse) runs against the detected web server’s access log (/var/log/nginx/access.logor/var/log/httpd/access_log), so WordPress brute-force alerts still fire on non-cPanel hosts – they just rely on the live log watcher rather than per-domain domlog scanning. modsec_auditruns on any host with ModSecurity installed.ssh_logins, SSH brute force, PAM listener, firewall, kernel modules, RPM/DEB integrity, and threat intelligence all run on every supported platform.
Deep Checks
Deep checks run every 60 minutes and cover thorough filesystem, CMS, email, and database scans.
Scheduled and account scan checks share one budget within the daemon, sized
from the machine’s core count with a minimum of two and a maximum of five
concurrent checks. Web UI scans and CLI jobs submitted with csm scan --full
share this budget. Standalone CLI scans run in a separate process. A cancelled
or timed-out check retains its slot until it actually exits, while its caller
can return promptly. Waiting for slots occupied by another scan does not count
as stalled dispatch in queue health.
The real-time scanner has a separate pool of two to sixteen workers, also sized from the core count. Its two-worker minimum lets other events progress while one content scan is slow. Both the packaged and installer-generated service units give CSM a lower CPU weight than the default. Under contention, this favors sibling services with default weights while allowing scans to use spare CPU on an idle host; it is not a CPU quota.
Scan coverage alerts
PHP, JavaScript, YARA, host-wide WordPress database coverage, and unfinished email password verification warnings keep one alert identity while their counts change. They follow the daily reminder window unless acknowledged. When the owning scan returns without that condition, its acknowledgment clears, so a later recurrence can alert again. An unrelated scan, throttled check, timeout, or cancellation does not prove recovery. Disabling a scanner clears its current coverage condition. Existing per-install database limit acknowledgments are kept when a partial scan cannot inspect those installations.
PHP analyzer crashes and stalls are tracked by the failing file paths, contents, and failure statuses, including files outside the displayed examples. Different failing inputs can therefore alert even after an earlier failure was dismissed. Check crash alerts ignore stack-trace churn, keep account scans separate by account, and re-arm after the same check returns successfully.
Filesystem
| Check | Description |
|---|---|
filesystem | Backdoors, hidden executables, suspicious SUID binaries |
webshells | Known webshell patterns (c99, r57, b374k, etc.) |
htaccess | .htaccess injection (auto_prepend_file, eval, base64 handlers) plus nine hardened per-pattern detectors – htaccess_php_in_uploads, htaccess_auto_prepend, htaccess_user_agent_cloak, htaccess_spam_redirect, htaccess_filesmatch_shield, htaccess_header_injection, htaccess_errordocument_hijack, htaccess_cgi_handler_abuse, htaccess_security_disabled. Auto-cleaning gated by auto_response.clean_htaccess. |
file_index | Indexed file listing to detect new/unauthorized files |
php_content | Suspicious PHP functions (exec, eval, system, passthru) |
group_writable_php | World/group-writable PHP files (privilege escalation) |
js_keylogger_dataflow | JavaScript keystroke exfiltration found by AST data-flow analysis: a key-event value that reaches fetch, sendBeacon, XHR, WebSocket, or a resource .src sink, including flows laundered through variables that regex rules cannot express. Runs inside the scheduled deep scan on the same file snapshots as the YARA pass but independent of the YARA backend, with its own rolling cursor. Complete sources up to 2 MiB are analyzed; a file the analyzer could not complete (oversize, parse failure, resource limit) is reported in the Warning-level js_taint_scan_incomplete aggregate and its prior finding is preserved, never silently cleared. Recognized PHP and HTML documents (including HTML-first PHP templates), complete JSON/source maps, CSS stylesheets, and gettext catalogs are skipped rather than counted as JavaScript parse failures. Classification uses content, not file extensions, and valid candidate JavaScript is always analyzed even when comments or strings contain document markers. Embedded JavaScript in templates or data files is not extracted or analyzed; PHP analysis and content rules do not provide equivalent JavaScript data-flow coverage. Ambiguous sources and malformed JavaScript retain their coverage warning, and oversize sources keep the size-limit warning. Disable with js_keylogger_dataflow (or the owner ID js_taint_deep) in disabled_checks; disabling the YARA scan does not disable this analyzer. |
symlink_attacks | Symlink-based privilege escalation attempts |
exposed_files | Web-downloadable sensitive files under document roots: version-control directories (.git/, .svn/, reported as web_exposed_repo_metadata and virtual-patched by denying the whole directory), database dumps, full-site backup archives, config/credential backups, PHP source-code backups, and phpinfo.php diagnostics. Each candidate is reported only after a headers-only reachability probe pinned to the vhost’s configured serving IP confirms the server serves it (HTTP 200/206, non-executed body) – files the server already blocks (403) and shipped samples such as wp-config-sample.php are never flagged. The domain is preserved for HTTP Host and TLS SNI, while DNS and front-end proxies are bypassed. The probe reads status and content type only, never the file body. The one exception is phpinfo.php: a 200 response alone cannot prove a dump, so CSM reads a bounded portion of that one response and reports web_exposed_phpinfo only when it contains real phpinfo output (the PHP Version banner in a dump-sized body); stub responses are not findings. Findings: web_exposed_db_dump, web_exposed_backup_archive, web_exposed_config_leak, web_exposed_source_backup, web_exposed_phpinfo, web_exposed_sample_sql. A plain SQL file with a sample/schema-specific name under framework/vendor scaffolding (examples/, docs/, vendor/, unpacked *-main/ or *-master/ directories) is reported as the lower-severity web_exposed_sample_sql warning. Archived, renamed, customer-named, and other ambiguous dumps stay Critical. Descent depth is bounded by thresholds.exposed_file_scan_depth (default 2, maximum 10). Optionally, auto_response.virtual_patch_exposed_files (off/manual/auto) writes reversible .htaccess Require all denied rules – manual applies confirmed findings via csm virtual-patch --apply, while auto applies every confirmed class except warning-only sample SQL and honors dry_run. Recognized backup storage directories are denied as a unit so regenerated archives stay blocked. A zip whose name carries no backup wording is classified from its central directory instead, so an archive named after the site it holds is reported when it contains a site-root or structured CMS configuration path, a regular Joomla configuration file at any depth whose same case-sensitive folder also holds a regular Joomla core file, a document-root directory, or a database dump at the archive root; plugin and theme bundles offered as ordinary downloads stay quiet. The complete entry list is inspected within a fixed metadata budget, without extracting or decompressing payloads; an archive that exceeds the budget leaves the scan incomplete instead of silently clearing an earlier finding. Each change records a rollback entry; restore refuses to overwrite a later customer edit. Findings in this family can be re-checked, and the version-triggered startup sweep re-checks existing findings after verifier changes. A finding clears only after a complete vhost map provides an unambiguous current serving address, both origin protocols answer there, and the shared detection rule no longer confirms an exposure; phpinfo re-checks repeat the bounded body confirmation too. A missing local file alone is not enough. The re-check never falls back to DNS, so a domain that has migrated to another host leaves the finding open rather than clearing on a stranger’s answer, and incomplete routing data or an unreachable or half-answered probe leaves it open too. |
The webshell, .htaccess, phishing and filesystem scans run on every deep cycle alongside the file index, including while the realtime monitor is active: the monitor reports neither a file renamed into place nor a setuid bit being set. The exposed-file scan runs in that tier too, on its own interval, because it confirms each candidate with a live request to the site rather than by reading the file. A cycle that skips it on that interval leaves its existing findings in place. The interval starts only after a complete exposure scan; an incomplete attempt can retry on the next cycle, including after switching between full and reduced scans.
Account discovery failures, unreadable directories or files, and truncated backdoor candidate lists leave the affected scan incomplete. Its earlier findings remain active until a complete scan can replace them; detections from the incomplete attempt are still added. These rules apply to both deep tiers.
The file index runs on every deep cycle, including while the realtime monitor is active. It checks new paths and rechecks indexed files with active findings; being present in the baseline does not clear an alert. Failed content reads preserve prior findings for that file. After a large deletion, scans rewalk the directories until the smaller baseline is adopted, so cached entries cannot restore removed paths or interrupt the consecutive-scan shrink guard.
The PHP content scan reuses a clean result only when the file’s identity and timestamps still match a stable read. Recently changed files, files that change during inspection, and files without usable identity metadata are read again on the next visit. Older cached results are refreshed as files are visited after upgrading; interrupted scans retain progress. Periodic host scans bypass the cache, and explicit full-content scans always read the files they visit.
PHP execution heuristics distinguish attribute metadata and multiline string contents from executable calls. Literal examples do not establish callable bindings or invoke them, and scanning continues through code after attributes. Attribute declarations do not qualify as comment-only stubs for demotion.
Re-verifying Findings
Repeated YARA worker buffer-scan failures are logged at most once per minute per tracked error, with suppressed counts reported when the error recurs after that window. If many distinct failures exhaust the tracking limit, additional errors share a bounded summary instead of resetting suppression. Scan errors still reach callers and preserve incomplete-scan reporting.
A finding records what was true when it was raised. The condition behind it is often resolved by someone else – an operator cleans a file, a virtual patch denies an exposure – and none of that moves CSM’s own rules, so re-verification runs once per deep-scan cycle as well as when the re-check logic changes.
A scan retires a finding it did not raise again only for files it actually examined. Coverage is tracked per file and scanner: a gap that names a file – one past the scan size limit, or one that could not be opened – keeps that scanner’s findings for the file and nothing else, while a gap with no path, such as a directory the walk could not enter or a stat that may hide a subtree, keeps every finding the scanner owns, because the unscanned range is unknown. This matters more than it sounds: a single oversized log that will never fit under the limit is a permanent gap, and treating it as a whole-scanner gap froze every finding on the host indefinitely.
Each family is re-checked by re-running the test that raised it, giving four outcomes:
- Cleared. The condition is provably gone: the file was removed, or the server no longer serves it as an exposure.
- Demoted to Warning. The flagged content is gone but the file changed since detection, so the cleanup cannot be proven byte-for-byte. The finding is kept – an attacker must not be able to retire one by editing the file – but it stops ranking beside live threats. Only a replacement proven inert qualifies: an empty file, or a comment-only PHP stub with a valid opening tag and no closing tag. A malformed opening tag can leave the file serving page content, so it does not qualify for demotion. This is an allow-list on purpose. A deny-list of dangerous shapes would make every omitted include, callback or inline script a way to buy a lower severity.
- Restored. A demoted finding whose replacement stops being inert returns to the severity it came from, so a second edit into a detection gap cannot leave live content sitting at Warning.
- Left alone. Anything uncertain: an unreachable or half-answered probe, a domain no longer served by this host, a replacement that changed while it was being read, a file that still matches.
A finding keeps its identity throughout; only its severity changes, because a finding’s key is derived from its details. Every store mutation is conditional on the snapshot verification actually examined, so a scan or realtime alert that refreshes the same key while a re-check is in flight is never overwritten by the older verdict. An unconfirmed demotion also survives a scan that does not raise the finding again, since that scan is weaker evidence than the verifier that reads the file.
PHP Remote-Source Taint Analysis
This analyzer runs as part of the scheduled deep-content scan, alongside the YARA and JavaScript consumers, and reports php_remote_taint. Disable it with php_taint_deep (or php_remote_taint) in disabled_checks.
It runs in a separate process, and that is not an implementation detail. A single file can put the underlying PHP parser into a state it never returns from, which no in-process deadline can interrupt. Analysis therefore happens in a supervised child process that is killed outright when a file exceeds its deadline; that file is reported as reduced coverage and the scan continues. Repeated failures are rate-limited so no single account can consume the scan budget, and the analyzer recovers without operator action. Whenever a file cannot be examined – including when no worker is available – it is reported as unexamined, never as clean. The byte check that decides whether a file could hold such a flow at all needs no parser, so it runs in the daemon itself: a file without a PHP open tag, a code-execution construct and a remote-fetching call is ruled out there and never reaches the worker. Cancellation during this byte check is still reported as unexamined.
CSM includes a PHP analyzer that looks for code fetching content from a remote server and then executing it, even when the fetch and the execution happen in different functions. This is a flow that regular expressions cannot express: matching it requires binding the value a fetch returns to the value handed to execution, which is beyond what YAML pattern rules or YARA-X can do.
The analyzer parses PHP source and tracks whether a value returned by a remote-fetching call reaches a code-execution construct (eval, include, include_once, require, require_once, create_function, or assert given a string argument), following the value through variable assignment, string concatenation, decoding calls, across function and method boundaries, and through a file_put_contents write into an include of the same path expression. curl_exec, curl_multi_getcontent, wp_remote_get, wp_remote_retrieve_body, and fsockopen are always treated as a remote source, whatever their argument. file_get_contents and fopen are dual-use: only these two calls have their argument inspected, and only when that argument carries an HTTP, HTTPS, FTP, FTPS, php://input, or data:// scheme are they classified as a remote source rather than a local read. Not every php:// stream counts as remote: php://input is request-controlled and treated as a source, while php://memory, php://temp, and a local php://filter resource are not – though a filter wrapping a remote resource still carries its nested remote scheme and is classified accordingly. A URL used directly as a sink’s own argument, with no acquiring call in between – include 'http://evil/x.php' – is out of scope: the pre-filter requires a source keyword before parsing is even attempted, and a bare remote include has none, so it is reported not_candidate, not analyzed.
Only two outcomes mean a file was actually examined: analyzed (parsing and the data-flow pass both completed) and not candidate (a fast pre-check proved the file cannot contain a reportable flow, without needing to parse it). Every other outcome – oversize, a parse failure, a parser recovery that produced only a partial tree, an internal resource limit, cancellation, or an internal error – is a coverage gap, not a clean result, and must never be read as “nothing to see here.”
An oversize file is judged once more before it counts. Only a file whose leading bytes could be source of that language is reported. Both analyzers are handed every readable file the walk produces, and their own pre-check rejects the rest instantly – but that pre-check never runs on a file too large to send, so without this step every large image, archive and compiled catalog on the host arrived as source the scan had failed to examine. PHP is judged by an opening tag anywhere in the inspected prefix, including after binary content. Binary content carrying PHP tokens can therefore remain reportable; the prefix alone cannot establish that it is inert. JavaScript has no such marker. A prefix without NUL is admitted; one containing NUL gets a conservative lexical check, since binary characters are legal in JavaScript literals and comments. Incomplete tokens and ambiguous syntax remain reportable. Neither gate looks for taint source or sink keywords in the prefix. A read that fails answers yes either way, because a file the scan could not examine is exactly what the report exists to name.
A finding’s severity follows how firmly the source was shown to be remote: a decoder-confirmed remote fetch reaching execution is Critical, a fetch carrying a remote URL is High, and a dual-use call whose argument could not be resolved either way is a Warning for review. When one file contains several flows, its strongest flow sets the severity. The last group is where legitimate template compilers and cache layers land, so it is kept visible without paging anyone.
Each flow in a finding’s details reads source -> sink (confidence, basis), for example curl_exec -> eval (high, always-remote) or file_get_contents -> include (low, unresolved). The basis says how the source was identified: always-remote (the call can only read over the network), literal (the argument text carries a remote scheme), decoded (the scheme appears only after escape or builtin decoding), request (a requester can supply the start of the path), call-argument (a call site in the same file passes the remote argument), or unresolved (the analyzer could not decide whether the argument is local or remote). The basis describes the strongest proof that reaches the sink, which can come from a different source call than the one the flow names. Some basis values appear only in later versions.
A finding’s identity is its file, its severity and the source and sink pairs of its reported flows. Changes to the details wording or the basis do not show a dismissed finding again. Any change in those flows or severity, or a new file, does. When evidence is capped, basis ranking does not change which equal-confidence flows are retained.
The WordPress database scan reports what it could not fully inspect as
db_content_scan_incomplete, counting affected installs against the number
discovered and naming one bounded example config path per reason:
unreadable_config, missing_credentials, unresolved_table_prefix,
query_failed or incomplete_content. Example paths use ASCII escapes for
control characters, non-ASCII bytes, quotes and backslashes, with the existing
length limit applied after escaping. Content gaps include truncated or
unusable query results. A server regular-expression timeout is a statement-local
failure: independent checks continue, while the failed check keeps coverage
incomplete and preserves earlier findings. Hidden-link selection examines only
the leading and trailing parts of each value that the parser reads, so large
values cost no more than the parsed sample. When the parser joins those parts
into a complete value, selection examines them together too. Ordinary CSS
declarations are matched with literal searches; guarded expressions handle
commented and encoded styles without admitting unrelated page text. Comment
bodies are preserved during matching so normalization cannot change their
meaning. If the server stops a regular expression, the selection is repeated
without commented styles, so plain
and encoded styles stay covered while the coverage gap is still reported. Each
install counts once; installs sharing a failed database count as affected without
retrying its queries. Incomplete discovery
is reported separately in the details because additional installs may be
missing from the total. With no attributed reasons, including when discovery
stops before reaching any install, the generic three-cause sentence remains
the fallback. Multisite safety limits keep their own account-specific warning
and do not count again in the summary or hide unrelated failures. Installation
findings carry an opaque database scope; the host-wide coverage summary remains
unscoped. A complete installation scan retires
its resolved findings even when another installation fails. Failed or
undiscovered installations retain their findings in the same atomic store
transaction. Older findings without a database scope stay until reobserved or
until the whole scanner completes; CSM does not guess their database from
message text. A partial multisite scan keeps findings for its entire network.
Incomplete or interrupted scans protect earlier findings from eviction when
new results fill the active list, including when every database or discovery
attempt fails, the scanner times out, or an internal panic stops execution.
That protection applies to findings already in the active list. Newly detected
conditions compete for the remaining space under the normal priority order,
so repeatedly incomplete scans cannot grow the list beyond its cap. Retained
findings can still refresh their details without losing their first observation.
Credential aliases sharing a database scope must all complete before that scope
can retire findings, regardless of scan order. This includes an alias whose
database and table prefix are known but whose login credentials are missing.
Query diagnostics include the detector stage, failure class and numeric MySQL error code. Repeated errors are counted together, with bounded detail when many causes occur. Raw SQL, server error messages and credentials are never included. Known statement errors, such as a missing table or column, leave the scan incomplete but allow independent detectors to continue. Connection, authentication and unknown failures stop further queries for that database.
Coverage the scan could not reach is reported as php_taint_scan_incomplete, which names how many files were affected and why – a per-file status such as a timeout or a worker failure, or a location the walk could not read at all, where the affected files cannot even be listed. Panics and timeouts are reported in a separate aggregate so hard analyzer failures remain visible beside routine coverage limits. A file that had a finding and later becomes unexaminable keeps its previous finding rather than having it cleared.
Separately from the outcome, an analyzed file can still report reduced precision. Constructs that defeat static variable identity – extract(), compact(), variable variables, a call dispatched through a value, an assignment target the analyzer cannot name, or a value a closure or arrow function captures from its enclosing scope – are recorded alongside the result. A recorded loss means tracking stopped at that point and the file may hold a flow that was not followed; it is never left implicit. The capture case is recorded only when the captured value was itself tainted, so it marks a real loss rather than the mere presence of a closure.
The parser supports PHP syntax up to version 8.1. A file written against a newer PHP version may use constructs the parser does not recognize; when that happens, parsing recovers what it can but the result is incomplete, and the file is reported as reduced coverage (a partial parse) rather than analyzed, so an incomplete view is never presented as a complete one.
WordPress
| Check | Description |
|---|---|
wp_core | Core file integrity via official WordPress.org checksums. A command timeout does not mark files as verified, and a timed-out re-check stays unresolved. Repeated verification failures produce wp_core_unverified. |
wp_plugin_inventory | Reports persistent plugin inventory failures as wp_plugin_inventory_unverified. Shares the existing refresh and interval with the outdated and vulnerable plugin checks. |
nulled_plugins | Cracked/nulled plugin detection |
outdated_plugins | Plugins behind the latest release, graded by version gap |
vulnerable_plugins | Installed plugins whose version matches a curated known-vulnerable feed (CISA-KEV + confirmed in-the-wild CVEs). Fires only from a fresh shared inventory when a parseable version is inside the affected range; patched, stale, and unparseable versions stay silent. Matched versions are High or Critical regardless of version gap, including inactive plugins because their files remain reachable. Alert-only (never disables a plugin). Toggle detection.vulnerable_plugin_scanning; accept one reviewed build via detection.vulnerable_plugin_allow (slug@version, case-insensitive). For an active install, a host engine that is Off or DetectionOnly, or a disabled account or vhost scope, marks the finding unprotected and raises it to Critical: no request filtering applies and no modsec audit record is written. When CSM ships a virtual patch for that CVE (virtual_patch in the feed) the alert says that patch cannot run; otherwise it claims only the filtering gap. An inactive install is never annotated – its files are why it is reported, but WordPress does not load the vulnerable code path, so a missing filter says nothing about reachability. An addon domain is linked only to its exact cPanel-associated subdomain; unrelated parked, main, and addon sites that happen to share a document root do not inherit one another’s disabled flags. Off cPanel, where CSM deploys no virtual patches and no per-vhost ModSecurity state exists, findings are left as the detector graded them. |
vulnerable_timthumb | Bundled TimThumb (timthumb.php / thumb.php) image-resizer scripts, the abandoned library whose remote-code-execution bug (CVE-2011-4106) is a recurring WordPress entry point. Files are confirmed by TimThumb’s own constants (not just filename) so generic thumbnail helpers are never flagged. A version below the last patch (2.8.14), an unparseable version, or an enabled WebShot / external-fetch feature is reported High; a patched-but-deprecated copy is a Warning to remove. The final release is suppressed only when ALLOW_EXTERNAL, ALLOW_ALL_EXTERNAL_SITES, and WEBSHOT_ENABLED are defined as false in executable PHP and none is also defined as true. Alert-only – TimThumb is never auto-quarantined, since deleting it would break the theme. |
db_content | Database injection (an external script loader on an ordinary HTTPS host that carries no attacker marker is reported once, as a Warning db_options_new_external_script, the first time that host appears in an option after the site’s first scan; plugin status options that WordPress renders as dashboard notices – LiteSpeed Cache’s CDN setup summary and its stored message list – are read by name instead and reported Critical as db_options_plugin_notice_injection when they carry markup that executes, because those rows hold plugin state and never site content, so neither host reputation nor the first-seen baseline applies and a payload stored before the site’s first scan is still reported. The named notice rows are read without a script-only SQL filter; complete values up to 64 KiB are inspected, while oversized or malformed query rows mark the scan incomplete and retain prior findings. Notice findings are alert-only; stored JSON and serialized data are not rewritten by this check. If another finding triggers option cleanup, cleanup refuses to write while any executable notice markup remains, even when its host has no reputation marker. A value WordPress stored as JSON has every slash escaped, so script URLs are normalised before a host is extracted, in options, in posts, and in page-builder content), siteurl hijacking, rogue admins, spam, a sudden publishing flood measured against the site’s own history (db_post_volume_burst; only on sites older than 400 days, with at least 50 recent posts and more than 5x everything published before – a genuine content migration looks the same, so it is alert-only), spam categories, tags and other taxonomy terms left behind when spam posts are removed (db_spam_taxonomy; a term named after a URL is High, spam vocabulary alone is Warning), PHP-capable snippets stored in the database by WPCode (db_stored_code_execution, Critical when the snippet is published and therefore running, lower when it is a draft or trashed; stored code is invisible to filesystem scanning), stored snippets that both defeat request caching (true DONOTCACHEPAGE, false WP_CACHE, WordPress no-cache headers, or LiteSpeed’s no-cache response header) and inspect the user agent for a search or SEO crawler (db_stored_cloak_logic, High when published and Warning otherwise; direct, concatenated, and ROT13 crawler names are recognised, while a snippet that already matched a malware signature carries the cloak as a bounded note rather than raising a second finding), outbound links wrapped in a container the page hides from readers (db_hidden_link_injection; an off-canvas container – a large negative left/top/text-indent – is High on its own, while display:none, visibility:hidden and fully transparent opacity need corroboration from a second linked domain in the same container or spam vocabulary, because themes legitimately hide panels that link out. Links back to the site’s own registrable domain are ignored, and the scan is skipped when siteurl/home cannot be read, since every absolute link would then look external. Markup is tokenized rather than parsed into a tree, so nesting the injection past a tree parser’s open-element limit does not hide it), autoloaded options named by a 32-character hex digest whose value base64-decodes to one complete PHP-serialized array (db_hostname_keyed_option; kits key the digest to the site’s own hostname so one payload serves many sites, and both halves are required because plugins do write hashed cache keys and base64 alone is ordinary), rewrite rules pairing an exact sitemap<N>.xml route with feed=xmlsitemap<N> in the same rule (db_doorway_sitemap_routes; sitemap plugins add numbered rewrite rules too, so the matching number is what distinguishes a doorway cluster from paginated sitemaps). Bounded option reads that omit any stored bytes mark the scan incomplete, and that row does not produce a finding. The scan also reports published posts attributed to a user that does not exist (db_phantom_post_author, Critical at 100 posts or more and Warning below). Small orphan groups can result from direct SQL user deletion; large groups can identify doorway content whose invented author IDs are hidden from dashboard queries. Direct WordPress installs are covered at every document root the panel serves, at every addon-domain directory in an account home, and one directory below a document root, so a root the panel has stopped serving is still examined, and every finding states which of the two it came from – served roots are reachable now, unserved ones hold a live database that publishes again as soon as a domain is pointed at them, and the statement is omitted entirely when the panel’s domain map could not be read rather than guessing; account backup, cache, staging and metadata directories are not, and installs deeper than one directory below a document root are not discovered recursively. Discovery is shared with the object, admin-overlap, credential-reuse, core-integrity and plugin checks and with their fixers and re-checks, so a finding raised here can always be re-located by the code that resolves it. A site address is reported when its shape cannot be one – a missing host, an invalid port, a backslash, a query string, a fragment, a script for a path, or a non-web scheme – because WordPress builds every asset URL from it. Invalid-address findings are alert-only; only explicit code injection enters database auto-response. Serving a site under a domain hosted elsewhere is ordinary and is not reported on its own – but a root the panel is serving right now, whose address names a domain absent from that account’s panel domains, is reported once per site as db_siteurl_foreign_host (High). Ownership follows the panel’s most-specific exact or wildcard domain mapping, so a delegated subdomain belongs to its own account rather than the account holding its parent. A migrated site is no longer served here, so the served state is what separates the two; an unserved root, an incomplete domain map, or a value the shape check already reports stays silent rather than saying it twice. Multisite-aware: when wp-config.php declares define('MULTISITE', true), up to 100 active secondary blogs (wp_<N>_options / wp_<N>_posts for blog IDs from wp_blogs) are scanned alongside the unprefixed main-site tables; each blog’s PHP-capable WPCode snippets and taxonomy are scanned, and its posts are checked against the network-wide users table. Larger networks emit db_content_scan_incomplete and retain prior findings. |
db_content_joomla | Joomla database content scanning. Discovers installs via configuration.php containing class JConfig, parses credentials from public $...; assignments. Scans <prefix>extensions params, <prefix>content article bodies, and joins <prefix>users with <prefix>user_usergroup_map for Super User detection (group_id=8). Findings: joomla_extensions_injection, joomla_content_injection, joomla_admin_injection. Administrator rows are baselined on the first pass; only an administrator that appears later is reported, once, as High. Every query is capped at 200 rows and runs under the scan deadline; installs under addon-domain document roots are discovered too. |
db_content_drupal | Drupal 8+ database content scanning. Discovers installs via sites/default/settings.php plus the core/lib/Drupal.php marker. Credentials parsed from the $databases array. Scans config, node_revision__body, and users_field_data joined with user__roles (administrator role). Findings: drupal_settings_injection, drupal_content_injection, drupal_admin_injection. Administrator rows are baselined on the first pass; only an administrator that appears later is reported, once, as High. Every query is capped at 200 rows and runs under the scan deadline; installs under addon-domain document roots are discovered too. Drupal 7 not yet covered. |
db_content_magento | Magento 1.x and 2.x database content scanning. Discovers installs via app/etc/env.php (M2, preferred) or app/etc/local.xml (M1). Credentials parsed via encoding/xml for M1 (CDATA-aware) or field-level regex for M2. Scans core_config_data, catalog_product_entity_text, cms_block, cms_page, and admin_user (with the configured db.prefix). Findings: magento_settings_injection, magento_content_injection, magento_admin_injection. Administrator rows are baselined on the first pass; only an administrator that appears later is reported, once, as High. Every query is capped at 200 rows and runs under the scan deadline; installs under addon-domain document roots are discovered too. |
db_content_opencart | OpenCart database content scanning. Discovers installs via the config.php + admin/config.php pair both containing define('DB_DRIVER'. Credentials parsed from DB_HOSTNAME / DB_USERNAME / DB_PASSWORD / DB_DATABASE / DB_PREFIX defines. Scans <prefix>setting (config_url / config_ssl are canonical hijack targets), <prefix>product_description, <prefix>information_description, and <prefix>user (admin/staff). Findings: opencart_settings_injection, opencart_content_injection, opencart_admin_injection. Administrator rows are baselined on the first pass; only an administrator that appears later is reported, once, as High. Every query is capped at 200 rows and runs under the scan deadline; installs under addon-domain document roots are discovered too. |
db_objects | MySQL persistence mechanisms: triggers, events, stored procedures, stored functions. Critical when the body matches known-malware patterns (sys_+exec, INTO OUTFILE, LOAD_FILE, etc.); Warning when an object exists at all (vanilla CMSes ship none). Toggle with detection.db_object_scanning; suppress Warnings via detection.db_object_allowlist. Manual drop via csm db-clean --drop-object. |
admin_overlap | WordPress administrator email overlap across cPanel accounts. Reports when the same admin email appears on the configured number of accounts, with reviewed emails and domains suppressible in detection. |
credential_reuse | WordPress administrator password-hash reuse across cPanel accounts. Groups identical hashes with an in-memory fingerprint and reports only the affected accounts and count. |
supply_chain | Composer and npm lockfile advisory matching against the local advisory database. Silent when no advisory file is present. |
Core verification and plugin inventory report persistent failures separately from malware or integrity findings. A failed attempt first appears in status; repeated failed attempts in distinct scan cycles raise a Warning naming the installation, last attempt and a bounded reason category. A missing wp-cli or an unreachable checksum service stops every installation on the host at once, so when one reason covers more installations than the per-cause limit a host scan collapses the warnings into a single finding naming that reason, the total and a sample of the paths. The limit applies across accounts for each kind of verification and cause. Account scans keep installation-specific warnings. A host summary belongs to an account only when every affected installation shares it; sampled paths do not determine account attribution. Raw wp-cli output is never stored in this history. These warnings do not enter incidents or automatic remediation. The history survives restarts, and successful checks clear the warnings through the normal completed-scan merge. Overlapping scans preserve attempt order and failure streaks even when they finish out of order. New discovery alone does not discard an attempt still finishing in another scan.
Cache reads do not count as attempts. Plugin coverage uses
thresholds.plugin_check_interval_min; partial refreshes keep the existing
refresh policy. disabled_checks controls these scheduled checks, and plugin
coverage performs no refresh when both inventory consumers are disabled.
Cancellation and incomplete discovery retain earlier evidence. Complete discovery
removes installations that are no longer present, scoped to the account scanned.
A wp-cli refusal explicitly identifying a non-WordPress directory is recorded as
not applicable rather than an outage. A completed core check that finds modified
or missing files retains its existing integrity findings and is counted separately.
Only the completed checksum summary establishes a negative verification result;
partial warnings without that summary remain unverified.
csm status, csm status --json and /api/v1/status report the last observed
coverage for core verification and plugin inventory. These counts describe
previous attempts, not a fresh scan of every site. See
verification status.
WordPress content findings use the account, database server, database name, and table prefix to stay separate. Changes to the panel’s document-root map do not create duplicate findings. Publishing a previously inactive suspicious snippet, or an orphaned-post group growing into a Critical content farm, produces a fresh alert even when the earlier condition was baselined or dismissed. Ordinary count changes within the same orphaned-post severity tier keep the existing identity.
Hidden-link findings are tracked per account, database host, database, and table prefix. More affected rows or a different row order do not raise another alert for the same destinations and concealment strength. Off-screen concealment raises a new High even after an earlier Warning was baselined or dismissed. Changes to destinations outside the displayed sample also produce a new finding. The corrected identity can show existing hidden-link findings once more after an upgrade; previous dismissals do not transfer to the new identity.
Hidden-link corroboration counts registrable domains within one hidden container, not hostnames across a whole row. Multiple subdomains of one linked domain count as one target, and both the WordPress home and site addresses count as local.
Joomla, Drupal, Magento, and OpenCart configuration reads accept regular files up to 1 MiB. Configuration symlinks and special files are rejected, including during Joomla and OpenCart marker probes. Drupal’s version marker must also be a regular file. Reads use the opened file throughout, reject changes observed during the read, and stop when the scan is canceled. Read failures and missing required credentials mark that CMS scan incomplete; manual re-checks keep the finding unresolved when its configuration cannot be inspected.
Administrator baselines for Joomla, Drupal, Magento, and OpenCart are scoped to the hosting account, CMS, database host, database name, and table prefix. Two sites under one account keep separate baselines when they use different databases or prefixes. Paths that share the same database and prefix share the same administrator set. Upgrading from account-wide baselines starts a fresh baseline for each installation on its first complete administrator query; later additions produce one High finding per new administrator. Finding details identify the affected database and prefix.
Database errors, discovery errors, and configuration or query limits keep the affected CMS check incomplete. Earlier findings remain until that CMS completes a scan; another CMS can still complete and clear its own resolved findings. Queries inspect at most 200 rows and request one extra row to detect overflow. Known statement errors allow independent queries to continue; connection failures and overflow stop further queries for that installation. Administrator baselines and recorded IDs change only after a complete result; a successful empty administrator result also establishes a baseline. New IDs in a partial result can still be reported against an existing baseline.
CMS Scanner Support Policy
New CMS scanner work targets upstream-supported major versions. EOL versions are best-effort when the existing scanner covers them through the same low-risk layout or schema. Adding a new EOL-only scanner needs operator fleet data and an explicit security reason.
Current scanner scope:
- WordPress single-site and multisite.
- Joomla installs using the common
configuration.php/JConfiglayout and standard content/user tables used by supported Joomla releases. - Drupal 8 and newer. Drupal 7 is not a planned support target.
- Magento 1 and 2.
- OpenCart installs using the standard storefront and admin config pair.
The supported kinds are declared once in internal/cms. The database
scanners and the PHP taint analyzer’s knowledge of each CMS’s bootstrap path
constants are tested against that table. The tests also require every typed
CMS kind constant, including local declarations and aliases, to have exactly
one descriptor. Membership drift fails the tests. The clean-corpus manifest
is validated against the same table: every supported CMS is either pinned as
a source or listed as pending with a reason. Support does not by itself mean
a CMS has clean-corpus false-positive evidence (see
the clean corpus). Database object scanning (db_objects)
discovers WordPress installs only.
Phishing & Malware
| Check | Description |
|---|---|
phishing | 8-layer phishing detection (kit directories, credential harvesting) |
email_content | Outbound email body scanning for credentials and suspicious URLs |
System Integrity
| Check | Description |
|---|---|
rpm_integrity | System binary verification via rpm -V |
open_basedir | open_basedir restriction validation |
php_config_changes | Security-weakening .user.ini and php.ini files under account web roots |
DNS & SSL
| Check | Description |
|---|---|
dns_zones | Security-sensitive DNS zone changes (delegation, mail, apex, and wildcard records) |
ssl_certs | SSL certificate issuance (subdomain takeover) |
waf_status | WAF mode, staleness, bypass detection. On cPanel, staleness follows vendor configuration files that WHM reports as active and ignores retired trees that remain on disk; the warning means at least one loaded vendor has not refreshed in over a month. Other platforms keep the conservative oldest-artifact check so an unused fresh tree cannot hide stale loaded rules. |
Email Security
| Check | Description |
|---|---|
email_weak_password | Email accounts with weak passwords; in-process verification with supported hash formats and cost limits |
email_password_audit_incomplete | Password verification was interrupted or encountered a hash outside the supported audit formats or limits |
email_forwarder_audit | Forwarders redirecting to external addresses, piping mail to a command, or discarding it. Pipes that run cPanel’s autoresponder, BoxTrapper or Mailman list software are not reported. Scheduled and real-time scans interpret command quoting and escaped destination quotes using Exim’s rules. |
email_mail_filters | Exim mail filters and dovecot/Roundcube Sieve scripts that copy mail to an external address while keeping a local copy, forward externally, pipe to a command, or blackhole all mail. Sieve is what webmail-managed rules actually execute, so both are scanned. A forward that leaves the mailbox its own copy is what a webmail forward rule produces, so on its own it reports as a Warning for review; it is Critical when an independent forwarding layer on the same mailbox, mail the mailbox never receives, or the same destination across accounts corroborates it. |
Performance
| Check | Description |
|---|---|
perf_php_handler | PHP handler configuration (DSO vs CGI vs FPM) |
perf_mysql_config | MySQL my.cnf optimization |
perf_redis_config | Redis configuration |
perf_error_logs | Error log file growth (bloat) |
perf_wp_config | WordPress wp-config.php settings |
perf_wp_transients | WordPress database transient bloat |
perf_wp_cron | WordPress cron scheduling (missed crons) |
Platform Support
The deep checks are the most cPanel-biased part of CSM because they iterate account home directories and per-user public_html trees. On plain Ubuntu/AlmaLinux the account-scan based checks do not run today:
cPanel-only (skipped on plain Linux):
htaccess,file_index,php_content,group_writable_php,symlink_attacks– iterate/home/*/public_html/**wp_core,outdated_plugins,vulnerable_plugins,db_content,db_objects,admin_overlap,credential_reuse– find WordPress installs through the shared discovery: the panel document-root map,/home/*/public_html, one directory below it, and addon-domain directories in an account home. Unresolved document-root aliases retain prior findings and cached plugin inventory instead of treating a partial walk as a clean result.supply_chain– scanscomposer.lockandpackage-lock.jsonunder/home/*and/home/*/public_htmlphishing,email_content– scan user home directories and Exim spooldns_zones,ssl_certs– read cPanel’s DNS zone store and SSL installation recordsemail_weak_password,email_forwarder_audit– read/etc/valiases, Dovecot/Courier auth databasesemail_mail_filters– read per-mailbox Exim filters under/home/*/etc/<domain>/<localpart>/filterand domain filters under/etc/vfiltersopen_basedir– reads EA-PHPphp.iniunder/opt/cpanel/ea-php*/php_config_changes– recursively scans.user.iniandphp.inibelow account web roots; incomplete walks emit a coverage finding and preserve prior findingsperf_wp_config,perf_wp_transients,perf_wp_cron,perf_php_handler– WordPress and PHP handler introspection via cPanel’s EA-PHP layout; the WP-Cron check also uses cPanel’s domain map for addon and subdomain roots
Runs on every platform:
filesystem,webshells– fanotify and file-tree scans over/home,/tmp,/dev/shmrpm_integrity– dispatches torpm -Von RHEL family ordebsums/dpkg --verifyon Debian familywaf_status– detects ModSecurity on Apache, Nginx, and LiteSpeed across all supported distrosperf_mysql_config,perf_redis_config,perf_error_logs– rely on standard service locations
Operators on plain Linux can point perf_error_logs, perf_wp_config, perf_wp_transients, and perf_wp_cron at generic web roots with the account_roots glob list (see configuration.md). The remaining account and CMS scans still assume the cPanel /home/*/public_html layout.
Package integrity rechecks retain a modification that the package verifier still reports for the flagged file. Changing its executable mode or removing the file does not by itself resolve that finding.
Hidden temporary files
The filesystem check reports hidden regular files in the shared temporary directories when they have executable permissions, ELF magic, or a leading shebang, long PHP opening tag, or short echo tag. Aliases of the same inode produce one finding; distinct files remain separate even when their names match. Failed content reads leave the filesystem scan incomplete and preserve earlier alerts.
An inert staged blob without these signals is outside this heuristic. There is no guaranteed alternate detection before it is made executable or interpreted; signature and runtime detection depend on recognizable content and activity. Changing its permissions to executable makes it eligible for the next filesystem scan. This check does not establish that every unreported temporary file is safe.
Auto-Response
When enabled, CSM automatically responds to detected threats. All actions are logged in the audit trail.
Actions
| Action | Description |
|---|---|
| Kill processes | Fake kernel threads, reverse shells, GSocket. Never kills root or system processes. |
| Quarantine files | Moves webshells, backdoors, phishing to /opt/csm/quarantine/ with full metadata (owner, permissions, mtime). Restoreable from the web UI. |
| Block IPs | Adds attacker IPs to the nftables firewall with configurable expiry. Rate-limited by auto_response.max_blocks_per_hour (default 50/hour). |
| Clean supported malware | Applies bounded PHP and .htaccess cleaners with pre-clean backups. Database cleanup has a separate opt-in. |
| Drop malicious DB objects | When clean_database is on, confirmed-malicious stored triggers/events/procedures/functions are dropped after a SHOW CREATE backup is recorded, so the drop is reversible. Detection runs regardless; the drop is gated on the operator opt-in. |
| PHP shield | Blocks PHP execution from uploads/tmp directories and inspects directly executed wp-content scripts for request-fed command sinks and packed eval loaders. |
| PAM blocking | Instant IP block when one address breaches thresholds.pam_bruteforce_threshold failures inside pam_bruteforce_window_min minutes, or fails against cred_stuffing_distinct_accounts distinct accounts. |
| Subnet blocking | Auto-blocks IPv4 /24 or IPv6 /64 when 3+ IPs from the same range were blocked within netblock_window (7 days by default). Currently blocked addresses count regardless of age, including operator and permanent blocks. Ended blocks count while their latest observed block start is inside the window, unless an earlier subnet block already answered them. A returning operator block starts fresh history after its absence is observed. Whitelist and clear actions, including bulk whitelist, forget the address. |
| Permblock escalation | Promotes temporary blocks to permanent after N repeated offenses. |
| Auto-freeze (PHP relay) | On cPanel, freezes active Exim messages attributed to a high-confidence PHP-relay finding. It has its own dry-run control and action-rate limit. See PHP-relay CLI. |
Subnet history is pruned hourly and saved only when it changes. If the file cannot be read it is left untouched and escalation counts only the addresses blocked right now until it is repaired or removed; status and doctor report the failure. Clear, whitelist and flush actions report history cleanup failures so operators can retry them.
Process termination
CSM opens a kernel process handle before verifying ownership, executable, start time, or the file referenced by a process. It checks that the captured process is still alive after verification and sends the signal through that handle. An exited process cannot redirect the signal to a replacement with the same PID. Automatic and manual malware termination reject root credentials, including effective and saved root IDs. The separate opt-in AF_ALG reaction requires the current credentials and executable to match its recorded event.
Safe signaling needs pidfd_send_signal (Linux 5.1). Where pidfd_open
(Linux 5.3) is absent – EL8 and CloudLinux 8 ship 4.18 kernels without it –
CSM pins the target through its /proc/<pid> directory descriptor, which
pidfd_send_signal accepts and which refers to the same kernel process. Both
paths reject a recycled PID; CSM never falls back to numeric PID signaling.
On a kernel without pidfd_send_signal, or when service restrictions deny the
call, termination stays disabled: csm doctor reports process termination supported as failed and the health status becomes degraded whenever
auto_response.kill_processes is enabled, so an inoperative protection is
visible before an incident needs it. Each health snapshot probes this capability
again, so transient resource failures clear once signaling becomes available.
Automatic termination failures are logged while the original detection remains.
Manual kill-and-quarantine reports a termination failure even if the file was
successfully quarantined; inspect both the process and recovery entry before
retrying. Manual request cancellation is checked before sending a signal.
Automatic file response limits
Realtime quarantine, scheduled quarantine, PHP cleaning and automatic .htaccess
cleaning share one rolling-hour budget. auto_response.max_file_actions_per_hour
caps host-wide attempts (default 50), max_file_actions_per_account_per_hour caps
attempts for one account (default 10), and max_file_action_failures_per_hour
pauses these responses after repeated failures (default 3). Zero or an omitted
key uses the default; negative values and values above 10000 are rejected.
Each attempt is reserved before touching the file. Successful, failed and
interrupted attempts all consume capacity. Reservations live in
<state_path>/file-response.json and survive configuration reloads and daemon
restarts. Entries expire one hour after admission. A backwards clock adjustment
keeps future-dated reservations charged until their window has passed.
Account budgets come from the target’s account-home path, not finding text. Paths outside recognized account homes share an unknown-account budget. One account reaching its limit does not stop other accounts unless the host or failure limit is also reached. Automatic directory and special-file quarantine is refused because one directory move can affect an unbounded number of files. Manual remediation remains available after reviewing the original detection.
A busy safety lock refuses that attempt without waiting behind another file
operation. Unreadable, incomplete or unwritable safety state also refuses
mutations. The original detections remain visible and a deduplicated
auto_response_paused warning reports the cause. Account-limit notices are
grouped at host scope so a fault across many accounts cannot flood the alert
budget. A pause does not create a retry job; new eligible detections can act
after capacity returns. Review outstanding findings and recovery evidence before
manual remediation. Do not delete safety state to clear a pause; repair storage
faults and let reservations expire.
Duplicate detections of one path share a single response attempt in each batch. Within one daemon run, alert delivery does not repeat a file response already evaluated by a scan or admitted to the realtime safety gate, including budget refusals. A realtime detection rejected using sampled content remains eligible for full-file validation during delivery. The original findings still reach alerts and history; a new detection can be evaluated again. Findings still queued at shutdown retain the existing restart replay behavior: the next daemon run evaluates them again under the same persisted limits.
A failed PHP cleaner leaves the file and any pre-clean backup for manual review. It no longer escalates to whole-file quarantine. A cleaner that recognizes no injection or declines an unsupported target refuses the file instead of failing. Sources that change or disappear before mutation are also refusals, including socket replacements and parent paths replaced after quarantine copying. These attempts still use capacity but do not count toward the failure pause. Read, write, backup and durability errors still count as failures. Quarantine and cleaners retain their descriptor-based identity checks, and automatic actions revalidate the file after saving the reservation. Opening a replacement special file cannot block response processing.
Manual full scans leave cleaner refusals for review instead of reporting a
failed remediation. csm clean also distinguishes a refusal from an action
failure; both return a nonzero exit status when the file was not cleaned.
These limits use the existing enabled, quarantine_files and clean_htaccess
opt-ins. Observe mode still forbids automatic changes. dry_run continues to
control IP blocking and web-exposed-file virtual patches; it does not preview
file quarantine or cleaning. Other response families have their own controls.
Restoring quarantined files
Regular-file quarantine and pre-clean backups write and sync the private content copy, metadata, and directory entries before removing or changing the original. Directory quarantine syncs the tree and metadata before moving it on the same filesystem. A directory move across filesystems is refused and leaves the source in place. Storage failures do not count as successful remediation.
New quarantine and pre-clean sidecars record the original modification time, owner, group, permissions, and size. Restore reapplies the saved attributes to the opened destination before syncing it. Linux preserves modification times at the precision supported by the destination filesystem. An ownership or timestamp failure keeps the recovery evidence and reports an error.
Older sidecars may use quarantine_at; listing and restore also read that
historical spelling. Entries without a saved quarantine date sort last. Older
entries have no recorded original modification time, so restore leaves the new
file’s modification time in place instead of treating the archive’s timestamp
as the original. Some older access-file cleanup backups also lack trustworthy
ownership and permission data; check those attributes when restoring them.
Configured account roots participate in manual remediation and restore. Set up service write access for roots outside the packaged grants before enabling these operations.
Web UI restore refuses symbolic links in destination directories and does not replace an existing file or directory. If a destination changes during restore, CSM reports a conflict and retains the quarantine entry. Check the original location before retrying; a failed or interrupted file restore can leave a partial file in the directory that was opened for restoration. CSM keeps this file because removing it could discard a concurrent replacement.
Restore syncs the replacement and its containing directory before deleting quarantine evidence. A failure after a move or unlink reports partial completion and retains recovery metadata where possible. Inspect both the original and quarantine locations named in the error before retrying; a copy may already be restored while quarantine cleanup remains incomplete. These guarantees depend on the filesystem and storage device honoring sync requests.
Virtual-patch rollback uses the same directory confinement. It replaces or removes the access file only when the saved content, owner, and permissions still match, preserving later customer edits.
Virtual-patch backups are reused only when their content and saved attributes, including modification time, match. A later edit with identical bytes but a new modification time gets a separate recovery point.
Rollback isolates displaced entries in a private directory. If another writer changes the destination, CSM preserves the live replacement and reports any retained recovery directory in the error. Check that directory beneath the original destination or quarantine location before retrying. If another writer moves a restored directory, CSM cannot recover it from its old name and will not substitute a different entry into quarantine.
Configuration
auto_response:
enabled: true
kill_processes: true
quarantine_files: true
max_file_actions_per_hour: 50
max_file_actions_per_account_per_hour: 10
max_file_action_failures_per_hour: 3
block_ips: true
block_expiry: "24h" # positive temp block duration; omitted defaults to 24h
max_blocks_per_hour: 50 # per-IP blocks per hour; 0/omitted uses default
netblock: true # enable subnet blocking
netblock_threshold: 3 # IPs from same IPv4 /24 or IPv6 /64 before subnet block; minimum 2
netblock_window: "168h" # blocked IPs count this far back, expired blocks included; omitted defaults to 168h
permblock: true # promote temp blocks to permanent
permblock_count: 4 # temp blocks before promotion; minimum 2
permblock_interval: "24h" # positive counting window; omitted defaults to 24h
# Response to http_scanner_profile findings: "challenge" (default)
# routes the IP to the PoW challenge when challenge.enabled is true,
# falling through to a firewall block when it is not; "block" always
# hard-blocks without offering a challenge.
http_scanner_action: "challenge"
# Deny HTTP access to confirmed exposed files with reversible .htaccess
# rules. off: alert only; manual: `csm virtual-patch --apply`; auto:
# apply confirmed findings except warning-only sample SQL when dry_run is false.
virtual_patch_exposed_files: "off"
# SAFETY DEFAULT: dry_run defaults to TRUE when this key is absent.
# In dry-run, BlockIP records the intended block to bbolt but does
# NOT touch nftables. Manual operator commands (`csm firewall ...`)
# bypass via BlockIPForce and always apply. Flip to false only after
# verifying the policy in dry-run.
dry_run: true
# Advisory verdict callback. CSM POSTs each impending auto-block
# to the panel before applying. The panel can downgrade to "allow"
# (audit-only), attach `tenant_id` for downstream correlation, or
# add a reason. CSM fails open on hook errors. Wire contract:
# docs/verdict-callback-contract.md.
verdict_callback:
enabled: false
url: "" # POST target
hmac_secret: "" # signing secret, or use hmac_secret_env
hmac_secret_env: ""
allow_unsigned: false # true only for staged unsigned rollouts
require_response_signature: true # reject unsigned callback replies
timeout_sec: 2
# PHP-relay auto-freeze (cPanel only). Off by default; opt in
# explicitly. dry_run defaults to true even when freeze=true so an
# operator who enables freeze without thinking gets a dry-run.
php_relay:
freeze: true # enable the exim -Mf hook
dry_run: true # safe default; flip with `csm phprelay dry-run off`
max_actions_per_minute: 60 # rolling 60s window cap on exim -Mf invocations
Exposed-file virtual patches
Each applied deny keeps a rollback copy under /opt/csm/quarantine/pre_clean/. Re-applying the same rollback state reuses that copy; content, ownership, permissions, or remove-versus-replace changes keep separate restore points.
All-in-One WP Migration rewrites the access file inside wp-content/ai1wm-backups, so CSM also denies only .wpress filenames from the parent wp-content access file. It does not add parent-wide .zip or .gz rules because those extensions can be legitimate downloads elsewhere under wp-content.
Dry-run safety default
auto_response.dry_run defaults to true when the key is absent. This is deliberate: an operator who turns on block_ips: true without reviewing policy gets recorded-but-not-applied blocks. The dry-run count surfaces in csm status --json and /api/v1/status so dashboards can verify the policy before flipping live. CSM clears those records when auto-response starts or reloads in live mode, and ages out records older than a week while dry-run remains enabled.
This setting is not a universal simulation mode. It gates automatic firewall and related network enforcement paths plus web-exposed-file virtual patches. File quarantine, cleanup, permission changes, and process termination are controlled by their individual auto_response flags. Leave those flags off while evaluating block policy.
IP auto-blocking still requires firewall.enabled: true. The firewall engine owns both live nftables mutations and dry-run block records; with the firewall disabled there is no engine to call, so csm validate warns on auto_response.enabled: true plus block_ips: true.
Verify dry-run state explicitly:
csm status --json | jq '.severities, .blocklist_size'
csm firewall status # "Recently Blocked" entries with timestamps after the restart confirm live mode
To go live: set dry_run: false, then run systemctl reload csm. The field is hot-reload-safe, and a successful reload validates and re-signs the config. For a planned restart instead, run csm rehash once before restarting.
Verdict callback (advisory)
When verdict_callback.enabled: true, every auto-block call POSTs a
signed JSON request to the panel before mutating nftables. CSM refuses
to start without hmac_secret or a non-empty hmac_secret_env value
unless allow_unsigned: true is set for a staged unsigned rollout.
Without that opt-in, an unsigned allow response is rejected and the
default block continues.
When a secret is configured, CSM also requires the panel to sign the
response body unless require_response_signature: false is set for a
staged rollout. With that opt-out, CSM still checks any echoed nonce
or timestamp when a secret is configured; a legacy response that
omits both keeps working. The panel can return {"verdict": "block"}
(apply), {"verdict": "allow"} (audit-only; CSM logs the decision and
skips nftables), or attach metadata (tenant_id, note). The callback
runs after local validation and infra-IP safety checks, and before the
dry-run gate, so panels can observe dry-run decisions too.
CSM fails open on hook errors (timeout, non-2xx, malformed body): the block continues as if the hook were disabled, or is recorded as dry-run when dry-run is active. The failure is written to the daemon log. Full request/response schema: docs/verdict-callback-contract.md.
Infrastructure IP DNS guard
Hostnames listed in top-level infra_ips or firewall.infra_ips are resolved every 5 minutes and their current addresses feed the infra auto-block guard. If a hostname stops resolving, the daemon emits an infra_ips_unresolvable Warning finding and keeps the last known addresses protected during the grace period (default 10 min). The finding auto-clears when resolution recovers.
Findings that always trigger IP block
When auto_response.block_ips: true and the firewall is enabled, qualifying findings in this list block the source IP. Per-row severity and challenge exceptions apply. The dry-run gate still applies if dry_run: true. Suppression rules do not stop these blocks; allowlist an address to exempt it.
Suppression rules also leave incident auto-blocking, credential-spray containment and central threat responses active. With database response enabled, suppression stops database cleanup and session revocation while keeping session IP blocking eligible. Suppress the action’s own check type to mute its notification; this does not disable enforcement.
Incidents and central threat responses receive new findings, including suppressed findings and checks that do not notify operators. Duplicate observations within a batch count once. Cross-account correlation uses only unsuppressed sources, and its derived alerts can be muted with their own suppression rules without removing them from enforcement.
| Finding | Description |
|---|---|
wp_login_bruteforce | WordPress login flood via wp-login.php |
xmlrpc_abuse | XML-RPC endpoint flood |
http_request_flood | Per-IP HTTP request volume exceeds threshold (disabled by default; enable by setting thresholds.http_flood_threshold > 0) |
http_scanner_profile | Random-URL probe pattern from one source IP (disabled by default; enable by setting thresholds.http_scanner_min_requests > 0; routed to the PoW challenge first unless auto_response.http_scanner_action: "block") |
http_claimed_bot_unverified | High-volume claimed crawler traffic while reverse-DNS verification is pending (routed to the PoW challenge first when challenge is enabled) |
http_ua_spoof | IP exceeding the UA anomaly threshold, including confirmed search-engine bot UA spoofing (periodic; see configuration.md for opt-in flags) |
ftp_bruteforce | FTP authentication flood |
smtp_bruteforce | SMTP authentication flood |
smtp_probe_abuse | Raw SMTP connect-rate flood before AUTH |
mail_bruteforce | IMAP/POP3/ManageSieve authentication flood without matching successful mailbox activity |
mail_account_compromised | Successful login from an IP that repeatedly failed auth on the same mailbox. Only Critical findings block; the established multi-mailbox High advisory is visibility only |
admin_panel_bruteforce | phpMyAdmin or Joomla admin POST flood |
ssh_login_unknown_ip | SSH login from an IP with no prior history, whether the realtime watcher or the periodic scan saw it first |
c2_connection | Outbound connection to a known C2 server |
ip_reputation | IP flagged by AbuseIPDB / rspamd / upstream threat-intel |
local_threat_score | IP crosses the aggregated internal attack-history threshold |
modsec_block_escalation | ModSecurity deny escalation |
modsec_csm_block_escalation | CSM-internal ModSecurity deny escalation |
waf_attack_blocked | WAF high-volume attacker |
email_compromised_account | Email account compromise indicator |
email_cloud_relay_abuse | Cloud relay abuse |
Successful cPanel, FTP, webmail and PAM login audit events and authenticated File Manager writes do not trigger direct blocks or challenges. Incidents containing only these audit events cannot request a block, even when they retain a higher severity from an older version. Independent attack evidence can still justify a response.
block_cpanel_logins gates cPanel multi-IP logins, API authentication failures,
webmail brute force and realtime FTP authentication failures. This includes
webmail challenge routing, since an unanswered challenge can become a block.
Queued block candidates are checked against the current check policy and login-blocking setting before retrying. Legacy queue entries without a check identity are discarded on upgrade; new eligible findings can queue them again. Existing firewall blocks and permanent evidence are not removed by this change.
Distributed HTTP flood rollups do not trigger a direct IP block because they describe one targeted vhost, not one source IP. The per-IP findings that feed the rollup still drive normal block decisions.
http_asn_crawl is also handled separately because one finding can carry several offending CIDRs rather than one source IP. Only Critical findings with confirmed PHP worker saturation can tempban those CIDRs, using auto_response.http_asn_crawl_tempban. Dry-run and subnet safety guards still apply.
mail_bruteforce_suspected and the established multi-mailbox High form of
mail_account_compromised are visibility only. They do not feed direct,
incident, or spray auto-blocking. A separate blockable finding can still block
the same source.
Every auto-response block source - scan findings, challenge-timeout
escalations, central-intel corroborated blocks, and incident spray
containment - records the same evidence: a temporary threat-DB row, a
blocked-IPs tracker entry, an auto_block finding visible to the block
digest and alerting, and a step toward permanent-block escalation.
Challenge, central, and incident blocks are not limited by
auto_response.max_blocks_per_hour; that budget applies to scan-driven
blocks only.
The resulting auto_block findings are output evidence, not new local
corroboration for central intelligence or incident correlation. They are
deduplicated before the digest, attack database, history, and alert sinks
receive them.
Safety Guards
- Never kills root processes, system daemons, or cPanel services
- Infrastructure IPs (
infra_ipsin config) are never blocked - Subnet blocks refuse the default route and any range that covers infrastructure, local host, allowed, or port-specific allowed IPs
- Quarantined files preserve full metadata for restoration
- Every regular file is copied from its verified open descriptor into a private quarantine inode before the detected name is removed. Other hard links are reported after removal; a file swapped into the detected path is reported as a refused remediation, with the captured copy kept as evidence and the replacement left untouched
- Realtime signature auto-quarantine requires high confidence: category
webshellordropper, file size at least 512 bytes, and either Shannon entropy >= 5.5 or hex density > 20% with an obfuscated-execution signal. This prevents legitimate WordPress plugins from being quarantined. - IP block rate limited by
auto_response.max_blocks_per_hour(default 50/hour) to prevent runaway blocking - CRITICAL alerts and threat-intel reputation sightings always bypass the operator email/webhook rate limit (default 30/hour); lower-severity findings batched with them still count against it
- Trusted countries (
trusted_countries) suppress login alerts from expected geolocations
What CSM Detects in Real-Time
Beyond standard malware patterns, CSM detects advanced evasion techniques:
- Fragmented function names: attackers split
base64_decodeacross variables ($a="base"; $b="64_decode") to evade simple string matching - Appended payloads: malicious code added to the end of large legitimate files, beyond typical scan windows. Realtime PHP checks scan the first and last 32KB, and periodic PHP content analysis scans a larger head window plus the tail.
- Non-PHP backdoors: Perl, Python, Bash CGI scripts in web directories (detects toolkits like LEVIATHAN)
- SEO spam injection: gambling/togel dofollow link injection into theme files
- WordPress brute force: real-time access log monitoring for wp-login.php and xmlrpc.php floods (blocks within seconds, not the 10-minute periodic scan)
- Admin-panel brute force: same access-log path, tracks POSTs to
/phpmyadmin/index.php,/pma/index.php,/phpMyAdmin/index.php, and Joomla/administrator/index.php. Emitsadmin_panel_bruteforceand auto-blocks the IP. Path matcher is intentionally tight to avoid false positives on shared hosting; Drupal and Tomcat Manager use different attack shapes and need separate detectors. - SMTP brute force and probes: tails
/var/log/exim_mainlogon cPanel and non-cPanel Exim hosts where the file exists. Emitssmtp_probe_abuseandsmtp_bruteforce(per-IP, auto-blocks),smtp_subnet_spray(per-/24, auto-blocks the whole subnet), andsmtp_account_spray(per-mailbox, visibility only). - Mail brute force: tails
/var/log/maillogfor direct IMAP, POP3, and ManageSieve auth failures. Composes with the existing geo-login monitor soemail_suspicious_geokeeps working. Emitsmail_bruteforce,mail_bruteforce_suspected,mail_subnet_spray,mail_account_spray,mail_account_compromised, andmail_auth_backend_degraded. Established good sources with a confined stale-password pattern emit the suspected advisory without auto-blocking. A compromise finding from an IP established on at least two other mailboxes stays visible as High without auto-blocking. Wider spraying and compromise without that standing still block. When the auth backend is degraded,mail_bruteforceandmail_subnet_sprayauto-blocking pause until backend errors age out. - Mail auth backend probe (cPanel): independently of the log signals above, CSM opens the
cpdoveauthdsocket on a short interval. dovecot keeps answering its IMAP/POP3 ports during a cpdoveauthd outage, so cPanel’s own service checks do not notice, yet every login fails regardless of password. When the probe finds the socket unreachable CSM raisesmail_auth_backend_degradedand pauses both mail and SMTP brute-force auto-block. Withauto_response.mail_auth_recovery.restart_enabled, CSM restarts the mail service once the backend has been continuously down pastdown_grace(default 10m, rate-limited), so a brief blip during maintenance never triggers a needless restart. Changes to mail auth recovery settings require a daemon restart.
Network dry-run precedence
Three settings combine for direct-SMTP BPF enforcement. Any applicable dry-run value that is true wins; live network denial requires all three to be false.
| Layer | Knob | Default | Effect when true |
|---|---|---|---|
| Auto-response | auto_response.dry_run | true | Record automatic block decisions without applying firewall denial |
| Detector | detection.direct_smtp_egress.dry_run | true | Suppress detector-scoped action |
| Kernel | bpf_enforcement.dry_run | true | BPF program emits decision but allows traffic |
The kernel knob is consulted by the BPF program itself; the others gate userspace action paths. All three default to true on a first install so a configuration mistake cannot start blocking traffic.
Observe mode
mode declares what CSM is allowed to do to the host it runs on.
| Value | Meaning |
|---|---|
enforce (default) | Every subsystem acts under its own switch. This is how CSM has always behaved. |
observe | Detection, correlation, alerting and the audit sinks run. Automatic host remediation and integration updates are disabled. |
mode: observe
Use observe mode to evaluate detection quality on a real host before granting CSM the ability to act, or to run CSM permanently as a reporting sensor next to another response tool.
What observe mode stops
Observe mode skips host integration work that otherwise runs automatically:
- The auditd rules file is written and
augenrulesis run, so CSM’s audit layers stay current across package upgrades. - The host integration files are refreshed: the WHM plugin CGI and its AppConfig registration, the CSM section of the ModSecurity user config, and the deploy script.
- Legacy challenge snippets and managed webserver integration snippets are refreshed, with webserver validation and reloads when needed.
- The WAF check refreshes stale vendor rules and deploys custom ModSecurity rules during periodic scans.
- The Exim forward guard is reconciled, including removing an installed guard when its switch is disabled.
- The AF_ALG kernel mitigation is re-applied on each critical tick when an opted-in hardening marker is present.
Existing host protections are left in place. Remove or change them explicitly before switching modes if that is the intended posture.
Resolve pending firewall apply/rollback windows in enforce mode before switching to observe. Observe startup refuses a pending settings or ruleset rollback: restoring it would change the host, while ignoring it would abandon the rollback deadline.
Contradictory settings are refused, not rewritten
A config that sets mode: observe and still enables a subsystem that changes
host state is rejected at load, naming every conflicting key at once:
mode: observe forbids changing host state, but these keys still enable it:
auto_response.enabled (set false), firewall.enabled (set false)
The keys checked are auto_response.enabled, firewall.enabled,
php_shield.enabled, bpf_enforcement.enabled,
email_protection.forward_guard.enabled, email_av.quarantine_infected,
auto_response.php_relay.freeze,
auto_response.mail_auth_recovery.restart_enabled, and
auto_response.virtual_patch_exposed_files set to auto. It also refuses
auto_response.copy_fail_kill_process: true and, when the antivirus scanner is
enabled, email_av.fail_mode: tempfail, both of which act on a host
independently of auto_response.enabled.
CSM refuses rather than silently turning those switches off in memory, because
the config re-signing path marshals the in-memory config back over csm.yaml:
an in-memory override would eventually be written into the operator’s file.
What still runs
Detection is unchanged. Real-time watchers, scheduled checks, correlation,
incidents, alerts, webhooks, the SSE stream and the audit-log sinks all behave
as they do under enforce, except that checks report host drift without fixing
it. CSM still writes its own configured state, log, cache and quarantine trees,
and creates its runtime sockets under /run/csm (/var/run/csm). Explicit
config edits and reloads can update the config integrity hashes.
Manual operator commands still work. csm firewall deny, csm clean,
csm virtual-patch --apply, csm db-clean and csm harden are explicit
actions an operator takes, not daemon behaviour, so observe mode does not block
them. auto_response.virtual_patch_exposed_files: manual is accepted for the
same reason.
Signature and GeoIP updates still run: they write only inside CSM’s own directories.
Two host writes remain, and the capability matrix marks both as not configurable:
- On a BPF build, capability discovery loads and briefly attaches probe programs to find out what the kernel supports.
- The Copy Fail check runs
kcarectl --patch-infoto see whether a livepatch covers the host. kcarectl refreshes its own cache under/var/cache/kcareon every run, including a read-only query.
Neither changes the host’s security configuration, and skipping them would cost the detection they exist for. They are listed here so “observe” is not read as a promise of zero writes.
Confirming the posture
csm doctor # "operating mode: observe (no automatic remediation or integration changes)"
csm status # mode: observe
csm status --json # .mode
curl .../api/v1/status # .mode
The capability string mode.observe.v1 reports that a build understands the
setting.
Changing mode requires a restart (systemctl restart csm), not a reload.
Doctor compares the running mode with the configured mode and warns if they
differ. When the daemon is unreachable, it reports only the configured mode.
Capability matrix
Every operation CSM performs that needs privilege beyond reading its own files, or that writes outside its own directories, with the privilege it needs and the setting that stops it.
Print it on any host, before or after install:
csm privileges # table
csm privileges --json # machine-readable
The command prints a static inventory. It does not load CSM configuration or probe host privileges, so it can be used before installing the service. Output failures produce a nonzero exit status in text, JSON and Markdown formats.
How to read it
Needs is what the operation requires from the kernel. root means uid 0
rather than one capability: the operation reads or writes files owned by many
different accounts, or drives a panel tool that assumes root. Today the daemon
runs as a single root process; splitting it into a small privileged helper plus
a reduced-privilege main process is planned, and this inventory is its first
stage. Capability names describe the privileged kernel interfaces as well as
the root filesystem access. For BPF, tracing and LSM programs need
CAP_BPF plus CAP_PERFMON; cgroup socket programs need CAP_BPF plus
CAP_NET_ADMIN. CAP_SYS_ADMIN is the broader fallback on older kernels.
These requirements follow the kernel program-load checks.
Trigger is automatic when the daemon may start the operation on its own,
and operator when it runs only in response to a command or a button. Not
running the command is how an operator operation is turned off.
Writes lists what the operation writes while it runs, including writes made
by a tool it invokes. Paths starting with / are filesystem paths; everything
else is a resource named <kind>:<name>. “nothing (read-only)” means the
operation only reads. Paths describe default locations; configured account roots,
state, rule, log, spool and policy paths replace those defaults and may require
service drop-ins. Shared CSM state and audit-log writes are grouped in the state
rows. This is not a list of every possible path written by an external panel tool.
Turn it off is the config key and YAML value that stops the named operation.
Dotted keys describe nested YAML mappings; they are not literal top-level keys.
Apply changes before restarting the daemon: several controls are read only at
startup. A row saying “not configurable” has no single config switch that stops
all its callers. disabled_checks takes a list and only affects scheduled checks.
Stopping new work does not undo installed rules, hooks, quarantines or blocks;
use the subsystem removal commands for that. Disabling the forward guard removes
its existing Exim configuration, which is itself a host write.
Action record says whether the operation writes to the action log. “no” does not mean the operation is silent; it means it is not on that stream yet and the daemon log is where it appears.
Without the privilege is what an operator loses by withholding it, so the row reads as a decision rather than a demand.
Relationship to the systemd sandbox
The daemon runs under ProtectSystem=strict with an explicit writable-path
allow-list; see service confinement. The two lists are
compared by tests: every declared in-daemon filesystem write must fit a packaged
grant, and every grant must have an in-daemon writer. Empty inventories and grant
lists fail. Paths are compared on directory boundaries, and an operation outside
the sandbox cannot justify a daemon grant. The unit parser rejects invalid paths
and unsupported syntax, including continuations, specifiers and directory flags,
instead of guessing which locations are writable. Tests cannot discover omitted operations
or prove the privilege and disable claims; those require tracing the runtime code.
Operations marked “outside the systemd sandbox” run as standalone operator commands or transient services forked by PID 1. The latter use direct execution on hosts without a reachable systemd service manager. Mixed operations have separate rows for their in-daemon writes, such as the AF_ALG marker repair and forward-guard lookup refresh. The service sandbox does not bound external tools.
Turning most of it off at once
mode: observe stops automatic remediation and integration deployment, and
refuses configurations that enable those subsystems. It still maintains CSM state,
loads temporary BPF capability probes in BPF builds, and invokes the KernelCare
probe, which can update its own cache. It is not a guarantee of zero host writes. See
observe mode.
The matrix
The risk tier says what can go wrong if an operation acts on the wrong
target: 0 changes nothing outside CSM’s own directories, 1 records a
recommendation or dry-run decision, 2 is a low-risk change such as a probe,
challenge gate or mail hold, 3 quarantines, blocks or denies a target, and 4
signals processes, reloads services, or rewrites content or configuration.
Rows describe maximum live effect, so preview-only tier 1 has no separate
row today. A tier does not promise automatic rollback; see recovery gaps
below and the current contracts in the JSON inventory. The TIER column of
csm privileges and the JSON inventory use the same 0 to 4 numbers.
| Operation | Needs | Trigger | Risk tier | Writes | Turn it off | Action record | Without the privilege |
|---|---|---|---|---|---|---|---|
state.control_socketcreate the private control socket used by operator commands | root | automatic | 0 | /var/run/csm | not configurable | no | CLI commands cannot reach the daemon |
state.mail_relay_policiesload operator-supplied mailer classes and proxy ranges for the PHP-relay detector | none | automatic | 0 | nothing (read-only) | email_protection.php_relay.enabled: false | no | only policy files readable by the daemon uid can be loaded |
state.php_shield_eventsreceive PHP Shield events over a socket and append their local archive | root | automatic | 0 | /var/log/csm-php-shield | php_shield.enabled: false | no | PHP runtime events are not collected |
state.sign_configrewrite CSM’s own integrity hashes into csm.yaml after an approved change | root | operator | 0 | /etc/csm | do not run the command | no | the integrity gate cannot be re-signed, so the next restart refuses to start after any config edit |
state.update_forgedownload and verify YARA Forge rules independently of YAML updates | root | automatic | 0 | /opt/csm/rules | signatures.yara_forge.enabled: false | no | YARA Forge rules are not refreshed |
state.update_signaturesdownload and signature-verify YAML malware rule updates | root | automatic | 0 | /opt/csm/rules | signatures.update_url: "" | no | detection freezes at the ruleset shipped with the installed package |
state.write_deploy_scriptrefresh the embedded upgrade script in CSM’s own directory at startup | root | automatic | 0 | /opt/csm/deploy.sh | mode: observe | no | the packaged upgrade helper is missing and upgrades are run by hand |
state.write_logswrite the daemon log and the audit-log sinks that feed a SIEM | root | automatic | 0 | /var/log/csm | not configurable | no | no local record of findings or actions |
state.write_storewrite the bbolt state database, baselines, incidents and scan reports | root | automatic | 0 | /var/lib/csm, /opt/csm/state | not configurable | no | CSM cannot run: without state there is no baseline, no dedup and no incident history |
detect.account_databasesread account MySQL credentials and scan databases for injected content and stored objects | root | automatic | 0 | nothing (read-only) | not configurable: db_object_scanning only stops stored-object scans; content and requested scans remain | no | no database detection: injected admin users, poisoned options and malicious triggers stay invisible |
detect.af_alg_socketsdeny AF_ALG sockets through BPF LSM, or observe socket use through the audit-log fallback | CAP_BPF, CAP_PERFMON, root | automatic | 3 | kernel:AF_ALG socket denial | detection.af_alg_backend: none | no | the live monitor is gone; the periodic critical check still reports an exposed kernel |
detect.audit_rulesquery loaded audit rules with auditctl to check detection coverage | CAP_AUDIT_CONTROL, root | automatic | 0 | nothing (read-only) | not configurable | no | audit-rule coverage cannot be verified |
detect.bpf_probeload and briefly attach BPF LSM, tracepoint and cgroup programs to discover kernel support | CAP_BPF, CAP_PERFMON, CAP_NET_ADMIN, root | automatic | 2 | kernel:temporary BPF programs and maps | not configurable: capability discovery runs independently of monitor settings in BPF builds | no | BPF capabilities report unavailable; configured automatic backends use their fallbacks |
detect.filesystem_eventsfanotify stream over account roots and world-writable temp directories | CAP_SYS_ADMIN | automatic | 0 | nothing (read-only) | not configurable | no | no real-time file detection; scheduled scans still run, so a webshell lives until the next cycle |
detect.kernel_livepatch_proberun kcarectl –patch-info to see whether a KernelCare livepatch covers Copy Fail; kcarectl rewrites its own cache on every run | root | automatic | 2 | /var/cache/kcare | not configurable: the startup kernel probe and hardening audits also invoke kcarectl, including in observe mode | no | a patched kernel reads as vulnerable, so the AF_ALG finding cannot be cleared |
detect.kernel_oomread the restricted kernel message buffer with dmesg to find recent OOM kills | CAP_SYSLOG, root | automatic | 0 | nothing (read-only) | not configurable | no | swap statistics remain available, but OOM kills are not reported |
detect.mail_queue_probequery the Exim queue through a transient unit, because Exim opens its logs even to answer a read | root | automatic | 2 | /var/log/exim_mainlog, /var/log/exim_paniclog, /var/log/exim4 (outside the systemd sandbox) | disabled_checks: [mail_queue] | no | queue depth is unknown, so a queue-size spam outbreak may go undetected |
detect.outbound_connectionswatch outbound connections through BPF cgroup hooks, or by polling /proc/net/tcp | CAP_BPF, CAP_NET_ADMIN, root | automatic | 2 | kernel:BPF programs and maps | detection.connection_tracker_backend: none | no | auto mode uses polling; an explicitly selected BPF backend stays unavailable |
detect.pam_eventsreceive authentication attempts from the pam_csm.so hook over a private socket | root | automatic | 0 | /var/run/csm | not configurable | no | no immediate brute-force or credential-stuffing block; log parsing lags by up to a minute |
detect.process_execwatch process execution through a BPF tracepoint, or by walking /proc | CAP_BPF, CAP_PERFMON, root | automatic | 2 | kernel:BPF programs and maps | detection.exec_monitor_backend: none | no | auto mode uses periodic /proc sampling; an explicitly selected BPF backend stays unavailable |
detect.read_service_logsread mail, authentication, web server and ModSecurity logs | root | automatic | 0 | nothing (read-only) | not configurable | no | no mail abuse, brute-force or WAF detection: these logs are root-readable only |
detect.scan_account_filesread every account’s files for scheduled, real-time and on-demand scans | CAP_DAC_READ_SEARCH, root | automatic | 0 | nothing (read-only) | not configurable: disabled_checks only suppresses scheduled checks; realtime and requested scans remain | no | only files readable by CSM’s own uid are scanned, which on a shared host is close to nothing |
detect.sensitive_file_writeswatch writes to /etc/shadow and comparable files through a BPF LSM hook | CAP_BPF, CAP_PERFMON, root | automatic | 2 | kernel:BPF programs and maps | detection.sensitive_files_backend: none | no | sensitive-file changes are found by periodic hashing instead of at the moment of the write |
integrate.auditd_ruleswrite CSM’s auditd rules and reload them, so audit-backed detection survives package upgrades | CAP_AUDIT_CONTROL, root | automatic | 4 | /etc/audit, kernel:audit rules | mode: observe | no | audit-backed detection layers stay inactive after an upgrade |
integrate.challenge_port_gateinstall the separate nftables port gate for the public challenge listener | CAP_NET_ADMIN, root | automatic | 2 | nftables:challenge port gate | challenge.port_gate.enabled: false | no | the listener remains reachable without the challenge port filter |
integrate.challenge_snippetrefresh the web server snippet and rewrite maps that route challenged visitors to CSM’s proof-of-work listener, and reload the web server when the snippet changed | root | automatic | 4 | /etc/apache2/conf.d, /etc/apache2/conf-enabled, /etc/httpd/conf.d, /etc/nginx/conf.d, /usr/local/lsws/conf/templates, /var/cache/csm, service:web server reload | mode: observe | no | suspicious visitors are blocked outright instead of being offered a challenge |
integrate.firewall_rulesetbuild and load CSM’s nftables table, including the operator’s port policy and rate limits | CAP_NET_ADMIN, root | automatic | 4 | nftables:csm table | firewall.enabled: false | yes | no firewall management; another tool owns the host’s packet policy |
integrate.modsec_sectionrefresh the managed ModSecurity section at startup and during periodic WAF checks, preserving operator configuration | root | automatic | 4 | /etc/apache2/conf.d, /usr/local/apache/conf | mode: observe | no | CSM’s virtual-patch WAF rules are not installed |
integrate.panel_plugindeploy the WHM plugin CGI and its AppConfig entry, and register it with the panel | root | automatic | 4 | /usr/local/cpanel/whostmgr/docroot/cgi, /var/cpanel | mode: observe | no | no panel plugin; the web UI is still reachable on its own port |
integrate.php_shieldinstall the PHP runtime hook and register its shared event directory for an operator-scheduled CageFS remount | root | operator | 4 | /opt/csm, /var/log/csm-php-shield, /etc/cagefs, /etc/csm, /opt/cpanel, /opt/alt, /usr/local/lsws, php:runtime configuration (outside the systemd sandbox) | do not run the command | no | no PHP runtime blocking of uploads and temp-directory execution |
integrate.waf_vendor_rulesask the panel to update stale ModSecurity vendor rulesets when a periodic check finds them out of date | root | automatic | 4 | /etc/apache2/conf.d, /usr/local/apache/conf, /var/cpanel, service:web server reload | mode: observe | no | stale WAF rules are reported but not refreshed |
operate.export_archivesread protected configuration, state or account evidence and export backup or forensic archives, temporary snapshots and checksum sidecars | root | operator | 4 | filesystem:operator-selected archive destinations (outside the systemd sandbox) | do not run the command | no | protected configuration, state and account evidence cannot be exported completely |
operate.harden_hostapply a supported CVE mitigation: a modprobe blacklist, or seccomp drop-ins for the services that need one | CAP_SYS_MODULE, root | operator | 4 | /etc/modprobe.d, /etc/systemd/system, service:restart, kernel:modules (outside the systemd sandbox) | do not run the command | no | the mitigation is applied by hand from the command the audit prints |
operate.install_serviceinstall or remove CSM itself: the systemd unit, the PAM hook, logrotate and the panel integrations | CAP_AUDIT_CONTROL, CAP_LINUX_IMMUTABLE, root | operator | 4 | /opt/csm, /etc/csm, /var/lib/csm, /var/log/csm, /etc/systemd/system, /etc/pam.d, /etc/logrotate.d, /etc/audit, /usr/sbin/csm, /lib64/security, /usr/lib64/security, /lib/security, /lib/x86_64-linux-gnu/security, /usr/lib/x86_64-linux-gnu/security, /lib/aarch64-linux-gnu/security, /usr/lib/aarch64-linux-gnu/security, /usr/local/cpanel/whostmgr/docroot/cgi, /var/cpanel, service:csm and panel integrations, /etc/cron.d, /var/cache/csm, /opt/cpanel, /var/run/csm, /etc/apache2/conf.d, /etc/apache2/conf-enabled, /etc/httpd/conf.d, /etc/nginx/conf.d, /usr/local/apache/conf, /usr/local/lsws/conf/templates (outside the systemd sandbox) | do not run the command | no | CSM cannot be installed as a service |
operate.manual_firewallblock, allow, tempban or flush addresses on request, and roll a firewall apply back | CAP_NET_ADMIN, root | operator | 3 | nftables:csm sets | do not run the command | yes | firewall changes are made with nft or the panel’s own tooling |
operate.manual_remediationclean or quarantine files and spool messages, truncate malicious crontabs, kill verified malware processes, or change database objects on request | CAP_KILL, root | operator | 4 | /home, /tmp, /var/tmp, /dev/shm, /var/spool/cron, /var/spool/exim/input, /var/spool/exim4/input, /opt/csm/quarantine, mysql:account databases, process:signal | do not run the command | no | remediation is done by hand over SSH |
operate.rehashre-sign configuration, converge legacy config copies, set binary immutability and refresh the launcher, service and log rotation | CAP_LINUX_IMMUTABLE, root | operator | 4 | /opt/csm, /etc/csm, /etc/systemd/system, /etc/logrotate.d, /usr/sbin/csm, service:daemon reload (outside the systemd sandbox) | do not run the command | no | hash signing and upgrade integration refresh cannot complete |
operate.restore_backupstage and restore protected configuration, drop-ins and state from a backup while the daemon is stopped | root | operator | 4 | /etc/csm, /var/lib/csm, /opt/csm, /tmp (outside the systemd sandbox) | do not run the command | no | configuration and state cannot be restored |
operate.truncate_error_logempty an account error log that has grown large enough to threaten the filesystem | root | operator | 4 | /home | do not run the command | no | a bloated log is reported and truncated by hand |
respond.af_alg_enforceunload the AF_ALG kernel modules again when an opted-in mitigation marker is present | CAP_SYS_MODULE, root | automatic | 4 | kernel:modules (outside the systemd sandbox) | auto_response.disable_enforce_af_alg: true | no | a module reload silently reopens the Copy Fail exposure |
respond.af_alg_killkill a verified AF_ALG socket caller through the separate Copy Fail response setting | CAP_KILL, root | automatic | 4 | process:signal | auto_response.copy_fail_kill_process: false | no | the process is reported but remains running; this path is independent of kill_processes |
respond.af_alg_markerrestore a changed AF_ALG mitigation marker after the operator has opted in | root | automatic | 4 | /etc/modprobe.d | auto_response.disable_enforce_af_alg: true | no | a changed module blacklist is reported but cannot be repaired |
respond.block_ipadd an attacker address or subnet to the firewall’s deny sets | CAP_NET_ADMIN, root | automatic | 3 | nftables:csm sets | auto_response.block_ips: false | yes | attacks are reported but not stopped; dry-run records what would have been blocked |
respond.bpf_deny_egressdeny matched outbound connections in the kernel through a BPF cgroup hook | CAP_BPF, CAP_NET_ADMIN, root | automatic | 3 | kernel:bpf cgroup program | bpf_enforcement.enabled: false | no | direct-to-MX spam egress is detected but not stopped at the source |
respond.clean_filestrip injected code from a PHP or access file, keeping a pre-clean backup | root | automatic | 4 | /home, /tmp, /var/tmp, /dev/shm, /opt/csm/quarantine | auto_response.enabled: false | yes | injected files are reported and left in place |
respond.database_cleanuprevoke a rogue CMS admin, sanitize poisoned options, and drop confirmed malicious stored objects after recording their definition | root | automatic | 4 | mysql:account databases | auto_response.clean_database: false | no | database persistence survives file cleanup and re-infects the account |
respond.enforce_permissionschmod a world-writable or group-writable PHP file back to 644 | root | automatic | 3 | /home | auto_response.enforce_permissions: false | no | writable code files are reported only |
respond.fix_wp_crondisable WP-Cron for an account and install a per-user system cron entry instead | root | automatic | 4 | /home, /var/spool/cron, /tmp | auto_response.fix_wp_cron: false | no | runaway WP-Cron load is reported only |
respond.forward_guardinstall or remove the managed Exim block that holds forwarded spam and backscatter, and rebuild the Exim configuration | root | automatic | 4 | /etc/exim.conf.local, /etc/exim.conf, /var/lib/csm, exim:rebuild artifacts and service reload (outside the systemd sandbox) | mode: observe | no | forwarded spam keeps damaging the host’s sending reputation |
respond.forward_guard_lookuprefresh the forward-guard bad-sender lookup inside the daemon sandbox | root | automatic | 0 | /var/lib/csm/forward_guard | email_protection.forward_guard.enabled: false | no | Exim continues using the last successfully written sender lookup |
respond.freeze_mailfreeze queued Exim messages attributed to a confirmed PHP-relay finding | root | automatic | 2 | /var/spool/exim/input, /var/spool/exim4/input, /var/log/exim_mainlog, /var/log/exim_paniclog, /var/log/exim4 (outside the systemd sandbox) | auto_response.php_relay.freeze: false | no | a compromised script keeps sending until an operator freezes the queue |
respond.hold_outgoing_mailrequest a cPanel account outgoing-mail hold through whmapi1 after sustained mail abuse | root | automatic | 2 | cpanel:account outgoing mail hold | auto_response.enabled: false | no | the mail abuse finding is reported but the account can continue sending |
respond.kill_processsignal a malicious process through a kernel process handle, never a recycled PID, never root | CAP_KILL, root | automatic | 4 | process:signal | auto_response.kill_processes: false | yes | reverse shells and miners keep running until an operator kills them |
respond.mail_delivery_gatedefer Exim delivery with fanotify permission responses when tempfail policy requires it | CAP_SYS_ADMIN, root | automatic | 2 | fanotify:mail delivery decisions | email_av.enabled: false | no | mail AV falls back to notifications and cannot defer delivery on scan failures |
respond.quarantine_filemove a confirmed malicious file out of an account tree into CSM’s quarantine, preserving owner, permissions and mtime | root | automatic | 3 | /home, /tmp, /var/tmp, /dev/shm, /opt/csm/quarantine | auto_response.quarantine_files: false | yes | malware is reported and left in place |
respond.quarantine_mailmove an infected message out of the Exim spool after an antivirus match | root | automatic | 3 | /var/spool/exim/input, /var/spool/exim4/input, /opt/csm/quarantine | email_av.quarantine_infected: false | no | infected mail is reported and delivered |
respond.restart_mail_authrestart the panel’s mail authentication service after a sustained outage | root | automatic | 4 | service:mail authentication | auto_response.mail_auth_recovery.restart_enabled: false | no | an authentication backend outage is alerted but not repaired |
respond.virtual_patchwrite a reversible deny rule into an account access file to close an exposed file | root | automatic | 3 | /home, /opt/csm/quarantine | auto_response.virtual_patch_exposed_files: off | no | exposed backups, dumps and configs stay reachable until the account owner fixes them |
Regenerate the table with go run ./cmd/csm privileges --markdown. A gate in
internal/ci fails when the page and the inventory disagree.
Risk tiers and recovery coverage
csm privileges --json reports each operation’s risk tier as a number from 0
to 4. A row with no tier would report -1; the test suite refuses such an
inventory, so a release never ships one. The tier is the operation’s maximum
live effect on a wrong target. It is not the severity of the finding that
triggered the operation.
Seven operations carry a safety contract in the JSON inventory:
respond.block_ip, operate.manual_firewall, integrate.firewall_ruleset,
integrate.challenge_port_gate, integrate.challenge_snippet,
respond.quarantine_file and respond.clean_file. Each states the authority it
needs, how the target is revalidated before the change, how the change is
reversed, and the limit that bounds it. The contracts describe current
behaviour, including what it does not cover. Every other operation that changes
the host states the recovery this inventory does not yet cover:
| Operations | Recovery gap |
|---|---|
detect.bpf_probe, detect.outbound_connections, detect.process_exec, detect.sensitive_file_writes, detect.af_alg_sockets, respond.bpf_deny_egress | This inventory does not yet specify verified detach, map restoration and crash recovery for these kernel hooks. |
detect.kernel_livepatch_probe, detect.mail_queue_probe | This inventory does not specify recovery of incidental external cache or log writes by probe commands. |
respond.mail_delivery_gate, respond.hold_outgoing_mail, respond.freeze_mail, respond.quarantine_mail | This inventory does not yet specify identity checks for releasing or restoring mail, or recovery after a restart. |
respond.af_alg_enforce, respond.af_alg_marker, respond.forward_guard, respond.fix_wp_cron, integrate.auditd_rules, integrate.modsec_section, integrate.panel_plugin, integrate.php_shield, integrate.waf_vendor_rules, operate.harden_host, operate.install_service, operate.rehash, operate.restore_backup | This inventory does not specify an action-wide snapshot and verified rollback of configuration, service and external-tool side effects. |
respond.af_alg_kill, respond.kill_process, respond.restart_mail_auth | Process termination and restart cannot restore lost process state; this inventory does not yet specify a full recovery contract. |
respond.database_cleanup, respond.virtual_patch, respond.enforce_permissions, operate.manual_remediation, operate.truncate_error_log | Current remediation may retain local recovery evidence, but per-operation identity-checked undo and partial-failure recovery are not specified by this inventory. |
operate.export_archives | Export can replace an existing operator-selected archive; removing the new archive does not restore overwritten bytes. |
Self-test
csm selftest scans a bundle of samples whose verdicts are known and reports
what the installed rules catch. It reads no account data and writes nothing, so
it is safe to run on a production host, and it answers the question an
evaluator actually has before pointing a scanner at real sites.
csm selftest
csm selftest --json
Exit status is 0 when every sample matched what the bundle records for each available engine, and 1 when anything needs attention. Incomplete or empty rule loads, scan errors, invalid samples and an empty bundle fail the run. A build without YARA-X can pass its realtime checks, but explicitly reports that YARA coverage was not tested.
The command uses the configured rules directory and disabled-rule list for
both engines. Disabling a detection can turn a sample into MISSED. A missing
or unreadable configuration fails the command rather than measuring a full
packaged ruleset that the host may not run. Pass --config and --config-dir
when the daemon uses custom paths.
What the bundle contains
Adversarial samples and benign controls, in one list. The controls are the half that decides whether a rule set is usable: a scanner that flags an ordinary WordPress plugin is worse than one that misses a shell.
Samples are stored base64-encoded and decoded only in memory. Endpoint antivirus deletes files that look like web shells, and a decoded copy on disk would make the bundle disappear from the machine it is meant to test.
The split-assert sample represents legacy PHP string assertions, not PHP 8 behaviour. The chr-built sample constructs a callable command-execution function; the test suite checks that it is callable without running its payload.
Reading the output
| Verdict | Meaning |
|---|---|
detected | An adversarial sample the rules caught. |
clean | A benign control the rules left alone. |
known gap | An adversarial sample the shipped signature rules do not catch, recorded in the bundle. |
MISSED | An adversarial sample that should have been caught and was not. A regression. |
FALSE POSITIVE | A benign control the rules fired on. |
GAP CLOSED | A recorded gap that now fires. Good news; the bundle needs updating. |
ERROR | The sample could not be decoded or scanned. The error is printed below it. |
The summary counts missed samples, false positives, closed gaps and errors
separately. A skipped engine has no sample results; JSON reports the reason in
its skipped field.
If configuration disables every rule in an engine, the command still measures its samples, reports zero loaded rules and the resulting misses, and exits with failure. Missing or malformed rule files remain load errors.
Known gaps are recorded, not hidden
A gap is a measurement, and the test suite fails when one closes as well as when one opens. A rule change that starts catching a recorded gap breaks the build until the bundle is updated, so no gap can quietly become permanent and no improvement goes unnoticed.
A gap is a gap in the signature engines. CSM’s other layers – PHP taint analysis, the behavioural checks, PHP Shield, and the correlation that turns findings into incidents – are not measured by this bundle. A sample listed as a known gap may still be caught in production by one of them.
Engines
CSM ships two rule sets, and the command reports each separately:
realtime, the YAML rules the real-time watchers use.yara, the YARA-X rules used by scheduled and email scanning. Present only in builds compiled with YARA-X; a build without it reportsSKIPPEDwith a reason. A compiled engine whose rules cannot load fails the command.
The regular CI test job runs the realtime gate. The required test:production
job runs scripts/production-tests.sh portable, which includes all packages
with the yara,journal,bpf tags and the pinned YARA-X library, so it runs the
YARA bundle gate too. Tagged builds and lint alone do not execute that test.
To measure the bundle locally, install PHP CLI for the fixture validity check, then run:
go test -count=1 ./internal/selftest
go test -count=1 -tags yara ./internal/selftest
The second command requires the pinned YARA-X library. The untagged command does not measure YARA detection.
Incidents
CSM groups related findings into Incident objects so operators see one escalating story per account, mailbox, or process instead of a stream of unrelated findings. Original findings are not mutated or suppressed – the Incident is layered on top.
Classification
Each incident is assigned a kind at create time. Inbound attacks are keyed
on the attacker source IP, not on the victim they name:
web_attack– a WAF block, scanner probe, or login brute-force. When the finding names a victim domain or account, that is the attack target, not evidence the account is compromised, so the incident keys on the attacker IP and one attacker’s hits across many vhosts collapse into one incident.mailbox_bruteforce– failed mailbox logins, mail brute-force bursts, and pre-auth SMTP probes from one source. A failed login is an attack attempt, not a takeover, so it keys on the attacker IP rather than the targeted mailbox.
Compromise kinds are reserved for evidence that an attack succeeded:
web_account_compromise (on-disk or behavioural signals such as a webshell or
suspicious PHP) and mailbox_takeover (post-authentication abuse such as
outbound spam, cloud relay, or a compromised-account signal). Inbound-attack
kinds carry short attacker-grade retention; compromise kinds get the longer
review window.
Two further classes of finding never open an incident at all. Findings that record an action CSM took, or how well it can see, are excluded so CSM’s own output cannot re-enter its decision path. Findings that report a standing weakness in installed software – a known-vulnerable or outdated plugin – are excluded because nothing has happened yet and there is nothing to contain. Known-vulnerable plugin findings remain alertable. Routine outdated-plugin notices remain informational, and both classes appear in the findings list.
Lifecycle
| Status | Meaning |
|---|---|
open | Active. New findings for the same correlation key keep merging in. |
contained | Operator marked under control. Findings still merge in window. |
resolved | Closed. Future findings start a new incident. |
dismissed | False positive. Future findings start a new incident. |
Resolved and dismissed incidents are pruned 30 days after their last
update when an operator closed them, and 7 days after their last update
when the daemon closed them (closed_by starting with auto:). Older records
without close attribution keep the 30-day period. Confirming an automatic
closure, changing a closed status, or recording a block on a closed incident
gives that decision 30 days of retention. Recorded operator decisions on older rows with
stale automatic attribution also keep this longer period. Reopening clears
close attribution; the next closure determines retention again.
Open and contained incidents are never auto-pruned by the retention loop, but they may be auto-resolved by the per-kind idle threshold described under “Auto-close” below. Retention sweeps use bounded transactions so a large backlog does not hold the store writer for the entire cleanup. Pruning frees space for reuse; it does not shrink the state file. Finding history has its own retention, so a retained incident can refer to findings already evicted from history.
Auto-close
To stop the open-incident backlog from growing without bound on busy
hosts, the daemon scans Open / Contained incidents shortly after startup
and then once an hour, auto-resolving any whose updated_at exceeds the
per-kind idle threshold. A live sweep closes at most 1000 stale incidents
at a time; if more stale incidents remain, follow-up sweeps run every 30
seconds until the backlog drains. Dry-run sweeps still scan the full set
so the counters show every would-close decision. Auto-resolved incidents
carry closed_by: "auto:stale" and an incident_auto_closed action in
their timeline so reporting can distinguish them from operator closes.
Defaults (configurable in csm.yaml):
incidents:
auto_close:
enabled: true # set false to disable
dry_run: false # set true to log decisions without writing back
by_kind:
mailbox_takeover: 24h
mailbox_bruteforce: 24h
credential_spray: 24h
web_attack: 24h
web_account_compromise: 168h
Kinds absent from by_kind are never auto-closed. The default map
omits host_integrity_risk, host_takeover, and post_exploit_process
because those host-level incidents should stay open until an operator
reviews them. host_takeover is the compound escalation raised when any
two of three host-takeover legs (a new uid-0 account, a planted suid
binary, an outbound connection to a bad ASN) are correlated for the same
host inside the merge window.
If a fresh finding for the same correlation key arrives after the auto-close, the merge-window stale-binding logic creates a new open incident – nothing about auto-close blocks re-detection. History is preserved on the closed record.
Tuning on high-volume hosts. Each by_kind threshold is the idle
time before a kind auto-resolves; they are independent and operator-set.
A host under sustained brute-force keeps a large open set mostly from the
longer-lived kinds (web_account_compromise defaults to 168h). If the
open-incident count is higher than you want to triage, shorten the
relevant by_kind entry (e.g. web_account_compromise: 72h) rather than
disabling auto-close. Untouched auto-resolved records are retained 7 days,
measured from when the incident resolves, so shortening the threshold also
moves the eventual prune point earlier relative to the last finding.
Auto-close still keeps a resolved record for follow-up instead of deleting
history at close time.
Metrics: csm_incidents_auto_closed_total and
csm_incidents_auto_close_dry_run_total.
Credential-spray suppression
Failed mailbox logins already collapse onto the attacker IP as a single
mailbox_bruteforce incident (see “Classification”). Spray suppression
extends that across protocols: it tracks the distinct-mailbox/account set
per source IP for the configured per_check detectors (mail and PAM)
across the merge window and, once an IP exceeds distinct_mailboxes,
opens a single credential_spray super-incident keyed on the IP with
breadth-based severity escalation. Subsequent findings from that IP
attach to the spray incident’s timeline.
Defaults (configurable in csm.yaml):
incidents:
spray_suppression:
enabled: false # default OFF; opt-in
dry_run: true # default ON; counters move, routing unchanged
distinct_mailboxes: 10 # threshold to trip
severity_escalate_at: 50 # bump severity to CRITICAL at this many
per_check:
- email_auth_failure_realtime
- pam_bruteforce
- credential_stuffing
max_tracked_ips: 10000
block_at_severity: "" # "" detection-only, "high" block on open,
# "critical" block on escalation
Setting block_at_severity hands the source IP to the firewall as soon
as the spray detector trips at the chosen tier, once
spray_suppression.dry_run is false. The detector also requires
auto_response.enabled and auto_response.block_ips; the firewall still
honors auto_response.dry_run, so a dry-run host logs the would-be block
without applying nftables rules. Live accepted requests are recorded on
the incident timeline as a credential_spray_block_requested action.
Non-live outcomes (dry-run, verdict-allow, already blocked) and failed
attempts do not latch the incident, so a later finding can retry after
blocking is live again. Concurrent findings for the same incident share
one in-flight firewall call, and resolved or dismissed spray incidents do
not make new block decisions.
Visibility-only findings do not make a spray incident blockable by themselves.
This includes mail_bruteforce_suspected and the High
mail_account_compromised advisory for an established multi-mailbox source.
A separate blockable finding in the same spray incident can still trip the
configured severity gate.
Whitelisted IPs (entries in reputation.whitelist and the live bbolt
whitelist updated via the Web UI) are skipped from spray detection so
internal mail relays, NAT egresses, and known-good infrastructure
never produce a spray incident.
Choosing block_at_severity:
""(default) – detection-only. Spray incidents open, no firewall hand-off. Use during dry-run validation and on hosts where blocking is owned by a separate system.high– block at thedistinct_mailboxestrip. Recommended once the dry-run counter looks clean. Trips on the first sustained burst before the source IP goes idle for longer than the merge window.critical– block only after severity escalates, i.e. one IP hitsseverity_escalate_atdistinct mailboxes before the source IP is idle for more than the merge window. A low-and-slow attacker that stays below that count before each idle reset never escalates and never blocks. Pick this only when you have strong shared-NAT exposure and accept that slow sprayers evade the gate.
Rollout:
- Ship the daemon with
enabled: false, dry_run: true. The detector tracks per-IP mailbox sets and incrementscsm_credential_spray_dry_run_totalwhenever the threshold would have tripped, but routing stays on the legacy per-mailbox path. - Validate the counter on your own infrastructure for 24h. If a
trusted IP shows up in the dry-run trips, add it to
reputation.whitelist. - Flip
enabled: true, dry_run: false. New attacker IPs route through the spray path; existing per-mailbox backlog drains via the auto-close path. - After another 24h, set
block_at_severity: high. The firewall hand-off runs on every spray decision (open + merge), so an incident opened before the flag was armed still blocks on the next finding from the same IP.
Metrics: csm_credential_spray_opened_total,
csm_credential_spray_suppressed_mailbox_takeover_total,
csm_credential_spray_dry_run_total,
csm_credential_spray_tracked_ips.
Incident auto-block
spray_suppression only handles the credential_spray super-incident
kind. Low-and-slow scanners that never trip a per-detector window
(modsec escalation, mail brute-force, smtp probe) still produce
web_attack, mailbox_bruteforce, mailbox_takeover, or
web_account_compromise incidents but never get firewalled. The
incidents.auto_block block adds a generic incident-driven firewall
hand-off:
incidents:
auto_block:
enabled: false # default OFF; opt-in
block_at_severity: "" # "" / "high" / "critical"
kinds: [] # empty = any non-spray kind with one source IP
When the gate trips, the correlator hands the source IP to the firewall
through the same dry-run / block_ips gate as the spray path. A live
accepted request records incident_block_requested; non-live outcomes
(dry-run, verdict-allow, already blocked) do not latch the incident, so
an operator who arms auto_block AFTER an incident has already crossed
the gate still gets a block on the next finding while the incident is
open or contained. Incidents with multiple source IPs are left for manual
review.
If a long-running incident’s timeline was truncated and the source IP is
not part of the incident key, auto-block also stays off because the
remaining visible timeline may not contain every source IP.
The same visibility-only findings are excluded from generic incident
auto-blocking. In particular, setting block_at_severity: high does not turn
an established multi-mailbox mail_account_compromised advisory into firewall
evidence. A Critical compromise finding remains blockable.
credential_spray is explicitly excluded from this path; the dedicated
spray hand-off owns it. Set kinds to narrow the surface (e.g. only
web_account_compromise) if you do not want every CRITICAL
mailbox_takeover incident to block its source IP.
This pairs naturally with the ModSecurity escalation thresholds
(thresholds.modsec_escalation_hits,
thresholds.modsec_escalation_window_min) – raising the window from
the shipped default of 10 minutes to e.g. 4 hours lets the modsec
detector promote paced scanners to a Critical escalation finding,
which then trips the generic auto_block gate.
ModSecurity escalation is confidence-gated. Each deny is classified as
high-confidence (a specific attack/probe rule – SQLi, RCE, traversal,
URL-encoding abuse, CSM custom, or an OWASP CRS rule from an
APPLICATION-ATTACK rule file, which is how LiteSpeed logs, since it
omits the rule message and tags), low-confidence (policy/anomaly scoring
rules such as COMODO content-type 210710 or anomaly-points 214930,
and OWASP CRS anomaly-score rules), or unknown. A burst escalates to a
firewall ban at the normal hit count only when it contains a
high-confidence deny or an unknown deny; a low-confidence-only burst
emits one non-actioned modsec_low_confidence_burst finding for
visibility instead of banning, then stays quiet until that active
low-confidence window drains. This stops false bans of legitimate
traffic (e.g. an unusual checkout that only trips content-type/anomaly
rules) without blinding CSM to real attacks, which trip the
high-confidence rules. A determined source that floods only
low-confidence rules is still banned once it reaches the
thresholds.modsec_low_confidence_escalation_hits backstop (default
30). Unknown blocking rules are escalation-eligible (fail-secure) and
raise a modsec_classifier_gap finding once per rule for the whole host
(repeated daily while the rule stays unclassified) so a new vendor rule
pack is noticed rather than silently given a no-ban path.
Kinds
web_account_compromise– findings attributable to a hosted account or script (PHP relay, webshell, account-scoped login bruteforce, etc.).web_attack– an inbound web attack or remote-IP reputation/threat-score signal. When the finding includes a victim domain or account, that field is recorded as target context, not as the correlation key. Keyed on the source IP and given attacker-grade retention (default 24h) so defended probes and attacker-IP reputation hits do not inflate the account-compromise count or its longer review window.mailbox_bruteforce– failed mailbox authentication, mail brute-force bursts, account-spray signals, and pre-auth SMTP probes. Keyed on the attacker source IP with attacker-grade retention (default 24h).mailbox_takeover– post-authentication mailbox abuse such as suspicious geo, credential leaks, cloud relay abuse, spam outbreaks, and confirmed compromised-account signals.post_exploit_process– process exec from/tmp,/var/tmp,/dev/shm.host_integrity_risk– daemon/kernel-level signals (sensitive file changes, fake kernel threads, binary/config tampering). Periodic binary/config tamper findings join the local host incident even without account or IP attribution. Startup verification still alerts and refuses to start before normal incident processing is available.host_takeover– any two of a new uid-0 account, a planted suid binary, and an outbound connection to a bad ASN, seen for the same host inside the merge window.credential_spray– one source IP brute-forcing many distinct mailboxes/accounts inside the merge window. Keyed on the source IP rather than per-mailbox, so a scanner spraying thousands of usernames produces one super-incident instead of thousands of mailbox_bruteforce rows. Findings from the same IP after the trip attach to this incident’s timeline. See “Credential-spray suppression” below.
The host-integrity set, all five compound sets, kind selectors and identity
exclusions are checked against the detector registry. The independent test
fixture internal/incident/testdata/check-policy.json records an explicit
classification role and selecting-set membership for every registered check,
including checks with no named override. Tests reject unknown names, missing
eligible members, new tables without a contract and new checks without a
policy decision. They also exercise each check with account, mailbox, process
and source-IP attribution to test classification precedence.
These tests live in the external incident_test package, which can import the
check registry without adding a production dependency from incident to
checks. That dependency would cycle through checks -> control -> incident.
Changes to the fixture require a review of the detector’s emitted evidence;
do not regenerate it from the selecting tables. Classification coverage does
not calibrate correlation weights or prove that broader compound membership
is safe.
Severity policy
Severity escalates only. Each incident keeps the highest severity any
joined finding has carried. Findings themselves are never re-emitted at
a higher severity. The audit trail records an
incident_severity_changed action when an incident’s severity bumps.
Correlation window
15 minutes by default. Findings outside the window for the same key start a new incident. The window is a named constant in code; not yet exposed via config.
Open threshold
Non-Critical findings normally need at least two correlated sightings inside the merge window before an incident opens. The first sighting is held in a pending bucket and counted toward the threshold; the second promotes both into a new incident with a two-event timeline. A mail-filter exfil finding remains a first-hit incident even when an uncorroborated copy-forward is graded Warning, so the operator review path is not lost. Stale pending entries are pruned by the daily retention sweep.
Critical-severity findings (account compromise, cloud-relay abuse, modsec rule escalations) bypass the threshold and open immediately so escalations still page on first hit.
The threshold suppresses one-shot scanner noise (a single modsec
deny from a wandering scanner, an isolated mistyped password) without
hiding sustained activity. The current pending-bucket size is exposed
as the csm_incidents_pending gauge.
The stored incident includes the full correlation key, including process PID/UID and remote IP when those are the only available dimensions, so active incidents keep merging after daemon restart.
Cross-account correlation of findings
Separately from incidents, the scan runner, the latest-state merge and the realtime dispatcher derive two findings from the findings they see:
coordinated_attack(Critical) when at least three distinct hosting accounts each carry at least one Critical finding from a check classified as a security event or malware artifact. The checks may differ between accounts. Repeated findings or several installs inside one account never raise the count.cross_account_malware(Critical) when the same malware-artifact check (webshell,new_webshell_file,backdoor_binary,new_executable_in_config) is present on two or more accounts at any severity. Different malware checks on different accounts do not combine.
Over the persisted latest-state set, both aggregates only combine findings
whose condition was first observed in the last hour. A scan re-emits every
finding it still sees and the merge refreshes its report time, so judging by
that would let a months-old condition re-enter the window on every cycle; the
merge therefore carries each finding’s first observation across re-reports and
correlation reads that instead, including when a completed scan replaces its
owned findings. A condition removed by a completed scan starts a new observation
if it is found again later. The window is relative to the merge time, so
the next completed scan clears expired aggregates even if it produces no
findings. Source rows are retained when their aggregates expire, a finding
carrying no first observation falls back to its report time, and a finding
carrying no timestamp at all still counts. Derived timestamps never affect the
window. JSON output omits first_seen when no first observation was recorded.
The first observation is not retroactive. On upgrade, a stored finding adopts its existing report time, because nothing recorded when that condition actually started. A long-lived finding with a recent report can therefore count once after upgrading; it ages out one hour after that saved report time, and the next completed scan clears expired aggregates even if it reports the same conditions again.
The realtime and scan batch paths keep their existing grouping and count all qualifying rows in the batch. A scan can carry forward an older finding for a file it could not examine; its original timestamp does not exclude it from batch correlation or get refreshed by correlation.
Every registered check is classified as a security event, a malware
artifact, ignored with a stated reason, or derived. Derived findings are never
inputs. A nonempty TenantID wins verbatim, including case, over the file
path and text. It must identify the host-local hosting account. Existing
stored values and verdict callbacks use the same precedence; callbacks must
supply the same owner keys as producers. External identifiers can split one
owner or combine distinct owners, and correlation does not validate them.
If that field is empty, correlation uses the account home containing the
cleaned absolute FilePath. Relative paths, root/home-only paths and paths
that escape the home after cleaning do not identify that home. This is a
lexical mapping through the platform’s account roots, including Plesk roots;
it does not resolve symlinks or prove that directory aliases are distinct
owners. Nonstandard content roots need producer-supplied identity.
The final fallback scans Message before Details for an account-root path.
It recognizes <account-root>/<user>/, not (account: user) labels. Within
each field it searches roots in configured order, can match an embedded root
substring, and can select an incidental path. These compatibility limits are
why account-aware producers should supply identity.
Qualifying rows with no identity from any source do not count toward either aggregate. They are counted once per call and logged once per check name per process, using only that call’s row count and no finding text. Ignored, derived, unknown and below-threshold security-event rows produce no diagnostic count. A check with a declared attribution gap is still eligible: when its producer does supply an authoritative owner, that Critical counts.
The health snapshot (csm status --json, /api/v1/status) carries a
correlation_attribution block with two views: current is the per-check
count of unattributed qualifying rows retained in the active set and inside
the correlation window at its latest merge, including any eviction caused by
the size limit. Unstamped legacy rows also count. It clears when a later merge
attributes those rows or they age out of the window; cumulative
sums every unattributed row since the daemon started, across active-set
merges and per-batch derivations, so a producer that recovered stays visible
as having failed. Counters are published together in merge order. The block
is absent until the first merge. csm doctor
reports the same state as correlation attribution: OK when current is
empty, WARN naming the checks and their counts otherwise, with the history
in both cases.
The two per-batch derivations (scan runner, realtime dispatcher) see only
their batch and produce alerts. The latest-state merge derives from the
merged, deduplicated, capped persisted active set under the same lock as the purge,
so evidence from separate scans combines: three accounts compromised in three
different scans still produce a persisted coordinated_attack, and it clears
only when fewer than three accounts still carry a qualifying Critical there.
Each completed runner snapshot replaces only its own checks’ rows; rows kept
for an unscanned owner or a file coverage gap stay as inputs; the previous
derived findings are dropped and recomputed from the merged set. Demotion,
dismissal or re-verification of a contributing row takes effect at the next
nonempty scan merge, and only if fewer qualifying accounts remain; there is
no immediate recompute and no promise that every file is re-verified.
Demoting a malware artifact below Critical cannot clear the all-severity
malware aggregate by itself. There is no time window, no shared-signature or
causal requirement, and no precision multiplier: the widened Critical inputs
accumulate across scans and can include unrelated events and false
positives, which is why the threshold calibration stays open on the roadmap.
Callers initialize platform detection before correlation; the latest-state caller does so before taking the store lock. Correlation reads cached roots, and attribution warnings are reported after the merge releases that lock.
Class membership does not mean an emitter currently reaches Critical: the non-WordPress administrator checks emit High with a stored baseline and Warning without one, and the outbound backdoor-port and bad-ASN variants emit High, so all of them are eligible but contribute nothing until a Critical variant exists.
Eligible producers supply verified identity when available. The test inventory names the producer test for each check, including unresolved branches:
- Database and CMS scanners stamp the owner of the install’s configuration path, resolved through the account roots. An install outside every root keeps its display label but no owner.
- Mail producers (rate windows, mail holds, credential leaks, bulk-service logins, forwarders, filters, mail brute force, cloud relay, geo logins, PHP relay volume) resolve a mailbox or domain to its owning account through the panel’s domain ownership table. Without that table (any panel other than cPanel) the owner stays empty and the row is reported, not counted. A bare account name must resolve to a passwd home directly under an account root. Credential and bulk-service findings use the authenticated identity, never the envelope sender. A sender-domain volume aggregate carries an owner only when every counted arrival proves the same local account through authentication or local submission. Arrivals with a remote ident username after the connecting address stay unattributed: that unquoted text can imitate later authentication metadata. Records whose greeting hides the connecting address also remain unattributed and supply no address. Mixed or unverified aggregates stay unattributed without reducing the volume count. Owner lookups run after tracker locks are released. Mail hold and governor findings require a local mail-server permission decision. Submission identities are read only from reception metadata before the message size; message IDs, subjects, addresses and login names cannot supply, replace or remove an owner or connecting address. Address literals in message IDs and recipients remain message data, and local arrivals do not acquire a peer from later metadata. An optional MAIL AUTH value follows the authenticated identity and does not change that identity or the peer. TCP Fast Open connections retain the same peer and ownership checks. The submission boundary follows Exim’s reception log fields and assumes Exim’s default greeting syntax check.
- Process, login and crontab producers accept a system user as owner only when its home directory sits directly under an account root, so root, service users and unknown uids never become an account. Direct SMTP findings apply this validation to both socket users and enriched process accounts.
- File families (content, phishing, htaccess, file index, core integrity, realtime file events, PHP shield events, self-deleting droppers) carry the judged file’s path, which resolves as described above. The collapsed core-integrity finding has no single path and carries the install owner.
- Socket checks resolve the kernel UID through the same passwd and direct account-home validation. Bad-ASN events also retain that owner when realtime process enrichment misses. Root, service and unknown UIDs stay unattributed. An attributed socket or mail finding includes its owner in the message so different accounts retain separate dispatch and audit identities.
- A sender-domain mail aggregate with mixed or unverified submitters remains a declared attribution gap. The owning account for a mailbox still requires the cPanel domain table. Any unattributed qualifying finding reaches the diagnostic count rather than contributing an invented account.
Correlation policy table
The table below is generated from the check registry: every registered
check with its class, its ignore reason when it is excluded, and its declared
attribution gap. A test compares it with the registry byte for byte and fails
when a row is stale, missing or duplicated; regenerate it with
go test ./internal/checks -run '^TestCorrelationDocumentation$' -args -update-correlation-docs
rather than editing rows by hand. A class says how correlation treats a check
when it fires; it does not say the check currently reaches Critical, that its
owner is available on every panel, or how precise it is.
Regeneration preserves the surrounding prose and marker line endings, even when they use CRLF. The generated block itself uses LF line endings.
Ignore reasons:
account-aggregate: already summarizes several accounts without a single victim identityattacker-side: attacker activity or attempted access, not evidence of compromise of the named victimhost-scope: host-wide condition with no account to attribute; a cross-account count cannot use it even when it is a real compromiseinformational: audit trail or inventory event with no compromise claimperformance: resource usageposture: static configuration, hardening or hygiene state; a Critical means a misconfiguration, not an attack on the accountresponse: record of an automatic action already taken; feeding it back would double countself-health: CSM’s own health, capacity or coverage state
Attribution gaps:
envelope-sender: sender-domain volume aggregate is unattributed when contributing submissions are unverified or belong to different accounts
| Check | Class | Ignore reason | Attribution gap |
|---|---|---|---|
account_scan | ignored | self-health | |
account_scan_error | ignored | self-health | |
account_scan_truncated | ignored | self-health | |
admin_cross_account_overlap | ignored | account-aggregate | |
admin_panel_bruteforce | ignored | attacker-side | |
af_alg_enforcement_corrected | ignored | self-health | |
af_alg_socket_use | security event | ||
api_auth_failure | ignored | attacker-side | |
api_auth_failure_realtime | ignored | attacker-side | |
api_tokens | ignored | informational | |
auto_block | ignored | response | |
auto_response | ignored | response | |
auto_response_paused | ignored | response | |
backdoor_binary | malware artifact | ||
backdoor_port | security event | ||
backdoor_port_outbound | security event | ||
bad_asn_outbound | security event | ||
bpf_ringbuf_error | ignored | self-health | |
bpf_unavailable | ignored | self-health | |
bulk_password_change | ignored | account-aggregate | |
c2_connection | security event | ||
cgi_backdoor_realtime | security event | ||
cgi_suspicious_location_realtime | security event | ||
challenge_route | ignored | response | |
check_panic | ignored | self-health | |
check_timeout | ignored | self-health | |
config_reload_error | ignored | self-health | |
config_reload_restart_required | ignored | self-health | |
coordinated_attack | derived | ||
cpanel_file_upload | security event | ||
cpanel_file_upload_realtime | security event | ||
cpanel_login | ignored | informational | |
cpanel_login_realtime | ignored | informational | |
cpanel_multi_ip_login | security event | ||
cpanel_password_purge | ignored | informational | |
cpanel_password_purge_realtime | ignored | informational | |
credential_log_realtime | security event | ||
credential_reuse | ignored | posture | |
credential_stuffing | ignored | attacker-side | |
crond_change | ignored | host-scope | |
crontab_change | ignored | informational | |
cross_account_malware | derived | ||
csm_health | ignored | self-health | |
database_dump | ignored | informational | |
db_content_scan_incomplete | ignored | self-health | |
db_doorway_sitemap_routes | security event | ||
db_hidden_link_injection | security event | ||
db_hostname_keyed_option | security event | ||
db_magic_token_user | security event | ||
db_malicious_event | security event | ||
db_malicious_function | security event | ||
db_malicious_procedure | security event | ||
db_malicious_trigger | security event | ||
db_options_injection | security event | ||
db_options_new_external_script | security event | ||
db_options_plugin_notice_injection | security event | ||
db_phantom_post_author | security event | ||
db_post_injection | security event | ||
db_post_volume_burst | security event | ||
db_rogue_admin | security event | ||
db_siteurl_foreign_host | security event | ||
db_siteurl_hijack | security event | ||
db_siteurl_invalid | ignored | posture | |
db_spam_cleaned | ignored | response | |
db_spam_found | security event | ||
db_spam_injection | security event | ||
db_spam_taxonomy | security event | ||
db_stored_cloak_logic | security event | ||
db_stored_code_execution | security event | ||
db_suspicious_admin_email | security event | ||
db_unexpected_event | ignored | informational | |
db_unexpected_function | ignored | informational | |
db_unexpected_procedure | ignored | informational | |
db_unexpected_trigger | ignored | informational | |
direct_smtp_egress | security event | ||
dns_connection | ignored | host-scope | |
dns_zone_change | ignored | informational | |
dpkg_integrity | ignored | host-scope | |
drupal_admin_injection | security event | ||
drupal_content_injection | security event | ||
drupal_settings_injection | security event | ||
email_auth_failure_realtime | ignored | attacker-side | |
email_av_degraded | ignored | self-health | |
email_av_encrypted_archive | ignored | self-health | |
email_av_hold_bypass | ignored | self-health | |
email_av_late_verdict | ignored | self-health | |
email_av_parse_error | ignored | self-health | |
email_av_quarantine_error | ignored | self-health | |
email_av_queue_overflow | ignored | self-health | |
email_av_scan_error | ignored | self-health | |
email_av_scanner_panic | ignored | self-health | |
email_av_timeout | ignored | self-health | |
email_cloud_relay_abuse | security event | ||
email_compromised_account | security event | ||
email_credential_leak | security event | ||
email_defer_fail_governor | ignored | informational | |
email_dkim_failure | ignored | posture | |
email_filter_blackhole | security event | ||
email_filter_exfil | security event | ||
email_filter_forwarder | security event | ||
email_filter_pipe | security event | ||
email_mail_filters | ignored | self-health | |
email_malware | ignored | attacker-side | |
email_password_audit_incomplete | ignored | self-health | |
email_phishing_content | ignored | attacker-side | |
email_php_relay_abuse | security event | ||
email_php_relay_account_volume_capped | ignored | self-health | |
email_php_relay_action_dry_run | ignored | response | |
email_php_relay_action_failed | ignored | response | |
email_php_relay_action_skipped | ignored | response | |
email_php_relay_cpanel_limit_unreadable | ignored | self-health | |
email_php_relay_disabled | ignored | self-health | |
email_php_relay_inotify_overflow | ignored | self-health | |
email_php_relay_inotify_overflow_recovered | ignored | self-health | |
email_php_relay_msgindex_persist_failed | ignored | self-health | |
email_php_relay_no_exim | ignored | self-health | |
email_php_relay_overflow_scan_truncated | ignored | self-health | |
email_php_relay_path2b_disabled | ignored | self-health | |
email_php_relay_policies_reload | ignored | self-health | |
email_php_relay_rate_limit_hit | ignored | response | |
email_php_relay_sweep_failed | ignored | self-health | |
email_php_relay_watcher_failed | ignored | self-health | |
email_pipe_forwarder | security event | ||
email_rate_critical | security event | ||
email_rate_warning | security event | ||
email_spam_outbreak | security event | ||
email_spf_rejection | ignored | posture | |
email_suspicious_forwarder | security event | ||
email_suspicious_geo | security event | ||
email_weak_password | ignored | posture | |
executable_in_config_realtime | security event | ||
executable_in_tmp_realtime | security event | ||
exfiltration_paste_site | security event | ||
exim_frozen_realtime | ignored | host-scope | |
fake_kernel_thread | security event | ||
fanotify_kernel_overflow | ignored | self-health | |
fanotify_overflow | ignored | self-health | |
firewall | ignored | host-scope | |
firewall_ipv6_unmanaged | ignored | host-scope | |
firewall_ports | ignored | host-scope | |
ftp_auth_failure_realtime | ignored | attacker-side | |
ftp_bruteforce | ignored | attacker-side | |
ftp_login | ignored | informational | |
ftp_login_after_bruteforce | security event | ||
full_scan_file_too_large | ignored | self-health | |
group_writable_php | ignored | posture | |
htaccess_auto_prepend | security event | ||
htaccess_cgi_handler_abuse | security event | ||
htaccess_errordocument_hijack | security event | ||
htaccess_filesmatch_shield | security event | ||
htaccess_handler_abuse | security event | ||
htaccess_header_injection | security event | ||
htaccess_injection | security event | ||
htaccess_injection_realtime | security event | ||
htaccess_php_in_uploads | security event | ||
htaccess_security_disabled | security event | ||
htaccess_spam_redirect | security event | ||
htaccess_user_agent_cloak | security event | ||
http_asn_crawl | ignored | attacker-side | |
http_claimed_bot_unverified | ignored | attacker-side | |
http_distributed_flood | ignored | attacker-side | |
http_request_flood | ignored | attacker-side | |
http_scanner_profile | ignored | attacker-side | |
http_ua_spoof | ignored | attacker-side | |
infra_ips_unresolvable | ignored | self-health | |
integrity | ignored | host-scope | |
ip_reputation | ignored | attacker-side | |
joomla_admin_injection | security event | ||
joomla_content_injection | security event | ||
joomla_extensions_injection | security event | ||
js_keylogger_dataflow | security event | ||
js_taint_scan_incomplete | ignored | self-health | |
kernel_module | ignored | host-scope | |
local_threat_score | ignored | attacker-side | |
magento_admin_injection | security event | ||
magento_content_injection | security event | ||
magento_settings_injection | security event | ||
mail_account_compromised | security event | ||
mail_account_spray | ignored | attacker-side | |
mail_auth_backend_degraded | ignored | self-health | |
mail_bruteforce | ignored | attacker-side | |
mail_bruteforce_suspected | ignored | attacker-side | |
mail_log_source_unavailable | ignored | self-health | |
mail_per_account | security event | envelope-sender | |
mail_queue | ignored | host-scope | |
mail_queue_unavailable | ignored | self-health | |
mail_subnet_spray | ignored | attacker-side | |
modsec_block_escalation | ignored | attacker-side | |
modsec_block_realtime | ignored | attacker-side | |
modsec_classifier_gap | ignored | self-health | |
modsec_csm_block_escalation | ignored | attacker-side | |
modsec_disabled_vhost | ignored | posture | |
modsec_low_confidence_burst | ignored | attacker-side | |
modsec_warning_realtime | ignored | attacker-side | |
mysql_superuser | ignored | host-scope | |
new_executable_in_config | malware artifact | ||
new_php_in_languages | security event | ||
new_php_in_sensitive_dir | security event | ||
new_php_in_sensitive_dir_clean | ignored | informational | |
new_php_in_upgrade | security event | ||
new_php_in_uploads | security event | ||
new_php_in_uploads_clean | ignored | informational | |
new_suspicious_php | security event | ||
new_webshell_file | malware artifact | ||
nulled_plugin | ignored | posture | |
obfuscated_php | security event | ||
obfuscated_php_realtime | security event | ||
open_basedir | ignored | posture | |
opencart_admin_injection | security event | ||
opencart_content_injection | security event | ||
opencart_settings_injection | security event | ||
outdated_plugins | ignored | posture | |
pam_bruteforce | ignored | attacker-side | |
pam_login | ignored | informational | |
password_hijack_confirmed | security event | ||
perf_error_logs | ignored | performance | |
perf_load | ignored | performance | |
perf_memory | ignored | performance | |
perf_mysql_config | ignored | performance | |
perf_php_handler | ignored | performance | |
perf_php_processes | ignored | performance | |
perf_redis_config | ignored | performance | |
perf_wp_config | ignored | performance | |
perf_wp_cron | ignored | performance | |
perf_wp_transients | ignored | performance | |
phishing_credential_log | security event | ||
phishing_directory | security event | ||
phishing_iframe | security event | ||
phishing_kit_archive | security event | ||
phishing_kit_realtime | security event | ||
phishing_page | security event | ||
phishing_php | security event | ||
phishing_realtime | security event | ||
phishing_redirector | security event | ||
php_config_change | ignored | posture | |
php_config_realtime | ignored | posture | |
php_config_scan_incomplete | ignored | self-health | |
php_dropper_realtime | security event | ||
php_in_image_realtime | security event | ||
php_in_sensitive_dir_realtime | security event | ||
php_in_uploads_realtime | security event | ||
php_remote_taint | security event | ||
php_shield_block | security event | ||
php_shield_eval | security event | ||
php_shield_webshell | security event | ||
php_suspicious_execution | security event | ||
php_taint_scan_incomplete | ignored | self-health | |
protection_queue_degraded | ignored | self-health | |
protection_queue_recovered | ignored | self-health | |
realtime_scanner_panic | ignored | self-health | |
reputation_quota_exhausted | ignored | self-health | |
root_password_change | ignored | host-scope | |
rpm_integrity | ignored | host-scope | |
self_deleting_dropper_overflow | ignored | self-health | |
self_deleting_dropper_realtime | security event | ||
sensitive_file_modified | ignored | host-scope | |
shadow_change | ignored | host-scope | |
signature_match_realtime | security event | ||
signature_update_rescan_queued | ignored | self-health | |
signature_update_rollback | ignored | self-health | |
smtp_account_spray | ignored | attacker-side | |
smtp_bruteforce | ignored | attacker-side | |
smtp_probe_abuse | ignored | attacker-side | |
smtp_subnet_spray | ignored | attacker-side | |
ssh_keys | ignored | host-scope | |
ssh_login_unknown_ip | ignored | informational | |
sshd_config_change | ignored | host-scope | |
ssl_cert_issued | ignored | informational | |
suid_binary | security event | ||
supply_chain_vuln | ignored | posture | |
suspicious_crontab | security event | ||
suspicious_file | ignored | host-scope | |
suspicious_php_content | security event | ||
suspicious_process | security event | ||
symlink_attack | security event | ||
test_alert | ignored | informational | |
threat_feed_stale | ignored | self-health | |
uid0_account | ignored | host-scope | |
user_outbound_connection | ignored | informational | |
vulnerable_plugins | ignored | posture | |
vulnerable_timthumb | ignored | posture | |
waf_attack_blocked | ignored | attacker-side | |
waf_bypass | ignored | posture | |
waf_detection_only | ignored | posture | |
waf_rules | ignored | posture | |
waf_rules_stale | ignored | posture | |
waf_status | ignored | posture | |
web_exposed_backup_archive | ignored | posture | |
web_exposed_config_leak | ignored | posture | |
web_exposed_db_dump | ignored | posture | |
web_exposed_phpinfo | ignored | posture | |
web_exposed_repo_metadata | ignored | posture | |
web_exposed_sample_sql | ignored | posture | |
web_exposed_source_backup | ignored | posture | |
webmail_bruteforce | ignored | attacker-side | |
webmail_login_realtime | ignored | informational | |
webshell | malware artifact | ||
webshell_content_realtime | security event | ||
webshell_realtime | security event | ||
whm_account_action | ignored | informational | |
whm_login_realtime | ignored | informational | |
whm_password_change | ignored | informational | |
whm_password_change_noninfra | security event | ||
whm_unauth_scripts_realtime | ignored | attacker-side | |
world_writable_php | ignored | posture | |
wp_core_integrity | security event | ||
wp_core_unverified | ignored | self-health | |
wp_login_bruteforce | ignored | attacker-side | |
wp_plugin_inventory_unverified | ignored | self-health | |
wp_user_enumeration | ignored | attacker-side | |
xmlrpc_abuse | ignored | attacker-side | |
yara_forge_rollback | ignored | self-health | |
yara_match_realtime | security event | ||
yara_match_scheduled | security event | ||
yara_realtime_scan_error | ignored | self-health | |
yara_scan_incomplete | ignored | self-health | |
yara_worker_crashed | ignored | self-health |
Findings from retired checks
A check name that no version of CSM emits any more stays registered while
older installations can still hold findings under it, because a finding is
only ever cleared when its name appears in the owning runner’s purge list.
Two file-index names, new_php_in_languages and new_php_in_upgrade, are
in that state: findings written by releases before the content-first file
index are cleared by the next completed file_index scan, are kept while
that scan is incomplete, and are kept per file while the scan reports a
coverage gap for that file. Nothing emits them again. php_dropper was
never emitted by any release, is not registered, and no response table lists
it: the manual, automatic and full-scan quarantine sets and the attack
database mapping are each declared once and tested against the registry, so a
renamed or never-emitted name cannot sit inert in a response table. The same
guard removed modsec_block and waf_block from the attack database mapping;
neither was ever emitted, so WAF blocks have never fed local reputation
scoring through that database. Whether the emitted ModSecurity block names
should is an open scoring decision, not something the guard decides.
Directory enumeration and PHP handler-configuration read errors make the file-index scan incomplete, including when only one account root is unreadable. The scanner keeps its previous index and directory cache, preserves its active findings, and still reports new findings from readable directories. A later completed scan clears the retired names; absent optional directories do not prevent completion. The first scan after startup and every retry after an incomplete or interrupted walk enumerate directories again, even if their cached modification times still match.
API
GET /api/v1/incidents– list, newest first. Without query parameters the response is a bare JSON array (compat with the existing wire shape phpanel/SIEM consumers decode against). When?limit=,?offset=, or?status=is present the response switches to an envelope:{"items":[...], "total":N, "offset":N, "limit":N, "status":"..."}. Status accepts the four spec values plusactive(open + contained, the default web UI filter). Limit is capped server-side at a safe maximum.GET /api/v1/incidents/<id>– one incident.POST /api/v1/incidents/<id>/status– transition status.
See api.md for endpoint detail.
Web UI
Open Monitor -> Incidents. The page has three tabs:
- Correlated – the default flat list of incidents with status filter, page size, and detail panel. The detail panel shows the current firewall block state for the incident’s source IP (permanent, temporary, cphulk, or not blocked) when an IP is known.
- Grouped – rolls up incidents by
(kind, source)so a credential spray that produced thousands of mailbox_bruteforce incidents shows as one row per attacker IP. Pageable with the same page-size selector as Correlated. Click a group to see member incidents in the detail panel, which also surfaces the source IP’s firewall block state; clicking a sample id jumps back to the Correlated tab focused on that incident. - Timeline Search – the older IP/account history search across the audit log.
Admin tokens can transition incident status (open / contained / resolved / dismissed); read-scope tokens can browse all three tabs.
Control socket
csm incidents list [--status all|active|open|contained|resolved|dismissed] [--limit N] [--offset N] [--all]
csm incidents show <id>
csm incidents status <id> <open|contained|resolved|dismissed> [details]
csm incidents bulk-status --older-than 24h [--last-seen-before RFC3339] [--status active|open|contained] [--kind K] [--domain D] [--account A] [--mailbox M] [--limit N] [--to resolved|dismissed] [--apply --confirm]
csm incidents list returns the first 100 incidents by default. Use
--offset for the next page, --status active for open + contained
incidents, or --all for an explicit full dump.
csm incidents bulk-status defaults to dry-run. It prints the total
match count and a bounded preview of the incidents that would change.
At least one age guard is required: --older-than, --last-seen-before,
or both. To mutate incidents, pass both --apply and --confirm.
Metrics
csm_incidents_open– gauge of currently open + contained incidents.csm_incidents_created_totalcsm_incidents_severity_changed_totalcsm_incidents_status_changed_totalcsm_incidents_findings_merged_totalcsm_incidents_compacted_totalcsm_incidents_pending– gauge of findings held in the threshold gate, awaiting a second correlated sighting.
Incident Response Runbook
Use this flow when CSM flags account compromise, mailbox takeover, malicious database triggers, or outbound spam on a production cPanel host.
Safety rules
- Do not delete customer files during first response.
- Do not thaw, release, or purge queued mail until affected credentials are rotated or an operator approves the specific queue action.
- Do not close incidents until the account was reviewed, credentials were rotated or explicitly deferred, and a fresh scan is clean.
- Take a CSM backup before upgrading CSM or changing incident state.
1. Verify the deployed binary
Deploy only after the required GitLab pipeline passed and the package was published.
/root/deploy-csm.sh check
/root/deploy-csm.sh upgrade
/opt/csm/csm version
/opt/csm/csm doctor --json
2. Take a backup
mkdir -p /root/csm-backups
systemctl stop csm.service
/opt/csm/csm backup /root/csm-backups/csm-pre-response-$(date +%Y%m%d%H%M%S).tar.gz
systemctl start csm.service
/opt/csm/csm doctor --json
sha256sum /root/csm-backups/csm-pre-response-*.tar.gz
Backup refuses a live or starting daemon and verifies the state lock before reading data. Restart the service immediately after the archive is created. If backup fails, start the service before investigating the failure.
Confirm the archive is readable:
gzip -t /root/csm-backups/csm-pre-response-*.tar.gz
tar -tzf /root/csm-backups/csm-pre-response-*.tar.gz | sed -n '1,80p'
3. Preserve evidence
mkdir -p /root/csm-forensics
/opt/csm/csm forensic-snapshot <account> --out /root/csm-forensics/<account>-$(date +%Y%m%d%H%M%S).tar.gz
sha256sum -c /root/csm-forensics/<account>-*.sha256
tar -xOzf /root/csm-forensics/<account>-*.tar.gz manifest.txt
Check the manifest for private-path exclusions, schema count, capture
errors, and recent_mtimes_status=ok.
4. Map affected accounts
Map incident domains and queued local senders to cPanel users before rotating credentials or changing mail queue state.
/opt/csm/csm incidents list --status open --all
exim -bpc
exim -bp | exiqsumm
grep -E '^example.com:' /etc/trueuserdomains /etc/userdomains
whmapi1 listaccts searchtype=user search=<account> --output=json
Use native cPanel APIs for inventory:
uapi --user=<account> Email list_pops --output=json
uapi --user=<account> Ftp list_ftp --output=json
uapi --user=<account> Mysql list_users --output=json
5. Rotate credentials
Rotate the cPanel account password, FTP accounts, affected mailboxes, WordPress administrator users, database users, and application secrets for the affected account. Prefer WHM and UAPI calls or the control panel over direct file edits.
Do this before releasing mail or marking incidents resolved unless the operator explicitly defers rotation for a documented reason.
6. Review queued mail
Start with read-only summaries:
exim -bpc
exim -bp | exiqsumm
exim -bp
Review headers before any queue action:
exim -Mvh <message-id>
Group messages into:
- safe to remove: frozen bounces, obvious backscatter, duplicate failed delivery notices with no customer value
- do not touch: current customer conversations, invoices, form leads, or any message where the business value is unclear
- needs review: suspicious local sender messages, mixed external bulk mail, or messages tied to an account whose credentials are not rotated
Only remove or thaw message IDs that were reviewed:
exim -Mrm <message-id>
exim -Mt <message-id>
7. Review stale incidents
Preview first:
/opt/csm/csm incidents bulk-status --older-than 72h --status active --kind web_account_compromise --limit 20
/opt/csm/csm incidents bulk-status --older-than 24h --status active --kind mailbox_takeover --limit 20
Apply in bounded batches only after review:
/opt/csm/csm incidents bulk-status --older-than 72h --status active --kind web_account_compromise --limit 100 --apply --confirm --details "operator cleanup after review"
For mailbox incidents, confirm mailbox rotation or explicit operator deferral before applying status changes.
8. Confirm recovery
/opt/csm/csm status --json
/opt/csm/csm doctor --json
exim -bpc
Keep the forensic archives, CSM backup, command notes, and queue decisions with the incident record.
Direct SMTP egress
CSM watches the local mail stack via spool + log scanning. Non-MTA processes that open outbound SMTP connections directly bypass that path. The direct SMTP egress detector catches that at connect time and feeds the incident correlator.
What fires
A finding with check: "direct_smtp_egress" is emitted when:
- A non-root process opens an outbound TCP connection.
- Destination port is one of the configured SMTP ports (default 25, 465, 587).
- Destination IP is not loopback, infra, or in the operator’s
infra_ipslist. - The process user is NOT a known MTA user (mail, mailnull, postfix, dovecot, dovenull, mailman, plus exim on cPanel).
Process names are never a standalone allow condition. A hosted account
renaming malware to smtp or smtpd still emits a finding.
The detector always emits findings when enabled. The dry_run knob does not suppress findings; it participates in the BPF enforcement gate, where any dry_run=true layer keeps kernel denial in observe-only mode.
Configuration
detection:
direct_smtp_egress:
enabled: true
backend: auto # auto / bpf / legacy / none
dry_run: true # safety default for detector-scoped action
ports: # each value must be 1-65535
- 25
- 465
- 587
Backends
auto– allow both BPF and legacy scan paths. Live backend choice still followsdetection.connection_tracker_backend.bpf– emit only from the cgroup/connect4,6 consumer.legacy– emit only from the/proc/net/tcp[6]polling path (live poller or scheduled critical scan). This path lacks PID/comm; MTA matching is user-only.none– detector disabled even whenenabled: trueis set elsewhere; useful for staged rollout.
The generic outbound connection tracker is still governed by
detection.connection_tracker_backend; this setting only gates
direct_smtp_egress findings.
Metric
csm_direct_smtp_egress_findings_total – monotonic counter,
incremented per finding emitted by the BPF connection consumer. The
legacy poller does not bump this counter today; operators who run
backend=legacy should track findings via the audit log.
rDNS enrichment
When the BPF backend is active, finding details include a Domain field populated from a TTL-cached reverse lookup (30 min TTL, 1 second per-lookup deadline). The lookup runs only after the cheap direct-SMTP filters match. On resolver miss or timeout the field is omitted; the finding still fires.
Caveats
2525is intentionally NOT in the default port list. Many operators run unrelated services on it. Add it toportsif your infra uses it for submission.- The detector emits regardless of the dry_run knob. Kernel denial
requires
auto_response.dry_run, this dry_run key, andbpf_enforcement.dry_runto all be explicitly false.
BPF cgroup-deny enforcement
Optional in-kernel denial of outbound connections that match an enabled userspace detector. Direct SMTP egress is currently the only enforcement gate. Every layer starts in dry-run so operators can review telemetry before allowing live denial.
What it does
When bpf_enforcement.enabled=true, direct_smtp_egress=true, the
connection tracker is running on BPF, and all dry-run layers are false:
- The cgroup/connect4 + cgroup/connect6 BPF program inspects each outbound TCP connect.
- If destination port is in the protected set AND the source UID is not in the safe-UID map AND the gated detector matches, the program returns 0 (kernel denies the connect).
- Userspace observes the decision via the
decisionfield on the ringbuf event and emits an audit-log entry.
When any dry-run layer is true (the default), the program emits the decision but always returns 1 (allow). Operators can run dry-run for as long as they need to gather telemetry before flipping to live denial.
What it does NOT do
- It does NOT wait on remote verdict callbacks in-kernel. That would add HTTP latency to every connect. The verdict callback (if enabled) runs in userspace after the BPF decision and enriches the emitted finding; it cannot undo a kernel denial.
- It does NOT enforce on UDP, ICMP, or non-cgroup paths.
- It does NOT replace detection. Findings still emit regardless; enforcement is a separate, layered control.
Configuration
bpf_enforcement:
enabled: false # master switch; default off
dry_run: true # safety default; flip after telemetry review
direct_smtp_egress: false # gate enforcement on the direct SMTP detector
verdict_callback: false # userspace post-decision callback
bpf_enforcement.enabled=true requires at least one feature gate.
Today the only gate is direct_smtp_egress, which itself requires
detection.direct_smtp_egress.enabled=true. The connection tracker
backend must be auto or bpf, and the direct SMTP backend must be
auto or bpf.
Kernel requirements
- Linux >= 4.10 with
CONFIG_CGROUP_BPF=y. cgroup/connect4andcgroup/connect6BPF program types.- The capability surface
bpf_enforcement.available.v1is the wire signal that the binary supports the feature; combined withbpf_enforcement_activeon the health snapshot, operators can detect both feature presence and runtime state.
On older kernels or default builds without the BPF tag,
detection.connection_tracker_backend: auto falls back to the legacy
/proc/net/tcp[6] poller. In that state direct SMTP findings still
work when detection.direct_smtp_egress.backend is auto or
legacy, but BPF enforcement is inactive.
When CSM attempts BPF and cannot start it, it emits a
bpf_unavailable finding. The message reports whether the daemon is
running on a fallback backend or has no live fallback active.
Metrics
csm_bpf_enforcement_decisions_total{decision="allow|dry_run|deny"}csm_bpf_enforcement_uid_map_refresh_total– successful periodic refreshes of the safe-UID BPF map.csm_bpf_enforcement_uid_map_refresh_failures_total– failed refreshes (e.g. /etc/passwd unreadable).
Dry-run precedence
Three independent dry_run knobs interact:
auto_response.dry_run: keeps automatic network response in observe-only mode.detection.direct_smtp_egress.dry_run: detector-scoped action knob.bpf_enforcement.dry_run: kernel-side denial knob.
Rule: any dry_run=true wins. Live denial requires all three to be false at the layer they apply, plus a BPF runtime backend. Defaults are dry_run=true everywhere on first install.
Rollout recipe
- Enable the direct SMTP detector without BPF enforcement. Watch
csm_direct_smtp_egress_findings_totalfor a week. - Enable BPF enforcement with
dry_run: true. Watchcsm_bpf_enforcement_decisions_total{decision="dry_run"}and confirm dry-run denials track expected hosted-account egress. - Set BPF
dry_run: falseon a single canary host. Audit incidents for false positives. - Roll out to fleet.
CVE Mitigations
CSM treats CVEs as a three-layer problem:
- Operator-driven hardening via
csm harden ...- applies the right preventive control for the host (modprobe blacklist, seccomp drop-ins, sysctl tweaks). - Continuous enforcement by the daemon - re-asserts the control on every startup and as a periodic check, so a package upgrade or manual
modprobedoes not silently undo the mitigation. - Live detection - auditd/BPF listeners flag exploit signatures the moment they fire, even on hosts where the kernel itself cannot be patched.
The hardening audit detects what the host actually has (kernel build, KernelCare livepatches, seccomp coverage) and refuses to claim protection it cannot deliver. Run csm harden with no arguments for the full list of available mitigations on the current host.
Active mitigations
CVE-2026-31431 - “Copy Fail” (Linux kernel AF_ALG)
Two operator paths depending on whether AF_ALG is loadable on the kernel:
csm harden --copy-fail- writes/etc/modprobe.d/csm-copy-fail-mitigation.confblacklistingalgif_aeadandaf_alg, then unloads them. Refuses on kernels where AF_ALG is built in (typical cPanel / CloudLinux 8), since the blacklist would have no effect there.csm harden --copy-fail-seccomp- writes systemdRestrictAddressFamilies=~AF_ALGdrop-ins for the units that spawn untrusted code: LiteSpeed, Apache/Nginx, every PHP-FPM pool, cron, mail. The right path on built-in-AF_ALG kernels.
The audit recognises KernelCare CVE-2026-31431 livepatches via kcarectl --patch-info and reports pass only when the host is genuinely mitigated (module blacklisted, seccomp drop-ins applied, or KernelCare livepatch active).
Live detection and BPF blocking
BPF-tagged builds (make BPF=1 or go build -tags bpf) load an LSM socket_create program on kernels with BPF LSM and ringbuf support. That program refuses socket(AF_ALG, ...) from non-root UIDs before the AF_ALG socket is allocated, returns EPERM, emits a ringbuf event, and feeds the existing Critical af_alg_socket_use finding path. UID 0 keeps AF_ALG access for system crypto users. There are no UID-range or alert-only BPF tunables today; use detection.af_alg_backend to select auditd or none if a host needs to avoid kernel-side refusal.
Default builds, BPF-tagged builds on unsupported kernels, and hosts forced to detection.af_alg_backend: auditd keep the audit-log listener. The audit path catches non-system AF_ALG socket attempts at Critical within about 500 ms but cannot stop the syscall before it reaches the kernel. Hosts that are not exploitable skip the live listener.
If CSM attempts the AF_ALG BPF path and cannot start it, it emits a
bpf_unavailable finding. The finding says whether the audit fallback
is active or no live fallback is available.
Auto-response
auto_response.copy_fail_kill_process: true- SIGKILL the offending process when the live listener fires. Default off (alert-only).auto_response.disable_enforce_af_alg: true- suspend the periodic re-assertion of the module blacklist without removing the hardening marker. For triage only.
The daemon self-heals its auditd rule file on startup if it has drifted from the embedded copy, closing the upgrade gap where a new binary shipped without re-running postinstall would leave detection silently inactive.
Configuration knob
detection.af_alg_backend-auto(default) |bpf|auditd|none.auditdis the kill switch for a misbehaving BPF-tagged release.bpfis strict mode (no fallback).nonedisables the live listener entirely.
The csm_af_alg_backend{kind="bpf-lsm"|"auditd-tail"|"none"} Prometheus gauge surfaces which backend the coordinator selected at startup, so dashboards can see the active path without parsing logs.
BPF validation
On a BPF LSM host with a BPF-tagged CSM build:
-
Set
detection.af_alg_backend: bpffor strict validation, or leave it asautoand confirm BPF was selected. -
Start the daemon and check metrics:
curl -k -H "Authorization: Bearer $CSM_TOKEN" https://127.0.0.1:9443/metrics \ | grep -E 'csm_af_alg_backend|csm_bpf_backend'Expected selected series:
csm_af_alg_backend{kind="bpf-lsm"} 1 csm_bpf_backend{feature="af_alg",kind="bpf"} 1 -
As a non-root user, attempt an AF_ALG socket:
sudo -u nobody python3 - <<'PY' import errno import socket import sys try: socket.socket(socket.AF_ALG, socket.SOCK_SEQPACKET, 0) except OSError as exc: print(exc.errno) sys.exit(0 if exc.errno == errno.EPERM else 1) raise SystemExit("AF_ALG socket unexpectedly succeeded") PY -
Confirm the command prints
1(EPERM) and CSM emits a Criticalaf_alg_socket_usefinding. The finding details should say the call was refused by the BPF LSM program.
CVE-2026-41940 - cPanel/WHM auth-bypass
Detection in the access-log path:
- Non-infra WHM login attempts surface at Warning (suppressible alongside other cPanel-login alerts) so an operator can see brute force traffic against the bypass surface.
- The tokenless WHM-script request the published exploit uses for cache promotion surfaces at Critical, always-on, and feeds auto-block.
No operator hardening command is required. The host fix is to apply the cPanel patched build. CSM provides the detection layer for windows where patching has not yet rolled out.
Future CVEs
New mitigations land here as they ship. The bar is:
- The host can be measurably hardened (modprobe / seccomp / sysctl / config), and/or
- An exploit-signature detector can fire reliably without false positives.
CVEs that are purely “patch the package”, with no preventive control we can apply and no signature we can detect, do not get a CSM mitigation; the right answer is the vendor patch. The daemon’s package-integrity check (rpm -V / debsums) covers the “did the operator actually apply the patch” question.
Firewall (nftables)
CSM includes a native nftables firewall engine that replaces LFD and fail2ban. It writes rules through the kernel netlink API directly via google/nftables. Integrity monitoring reads the installed structure with the nft command.
Features
- Atomic ruleset - single netlink transaction, no partial application
- Named IP sets with per-element timeouts (blocked, allowed, infra, country)
- Rate limiting - SYN flood, UDP flood, and per-IP connection rate are dual-stack (IPv6 metered per /64); per-port flood meters are dual-stack per source address; the concurrent connection limit is IPv4-only
- Country blocking via MaxMind GeoIP CIDR ranges
- Outbound SMTP restriction by UID (prevent spam from compromised accounts)
- Subnet/CIDR blocking with auto-escalation from individual IPs and safety guards for infra, local, and allowed addresses
- Permanent block escalation after repeated temp blocks
- Dynamic DNS hostname resolution (updated every 5 min) with grace-period guard against transient resolver failures
- IPv6 dual-stack with separate sets
- Commit-confirmed safety - Juniper-style auto-rollback timer
- Infra IP protection - refuses to block infrastructure IPs
- Auto-response dry-run - safety default that records intended blocks without touching nftables
- Verdict callback - optional advisory hook to the panel before each auto-block (allow / block / attach metadata)
- cphulk integration - unblock flushes cphulk too
- Audit trail - JSONL log with 10MB rotation
- State persistence with atomic writes
Storage contract preparation
The durable firewall action service is implemented and tested through engine injection. Production activation, reader cutover, migration, restore, downgrade and operator recovery interfaces remain open.
Firewall actions record the actor, source, linked finding or incident, and complete before and after state before changing the kernel. Pending intent is separate from committed state. Recovery verifies target identity and expiry before recording an outcome; it does not blindly replay a mutation. Repeated request IDs reuse the original action and admission accounting. Audit delivery retries use the same action ID. Typed undo checks that the affected targets still match the recorded result.
Proven outcomes are retained for undo and review, bounded two ways. The retention sweep drops delivered outcomes older than the findings-history setting, and a hard cap on retained outcomes and on their total size applies even when sweeps stay off. The size cap includes retained audit evidence. Existing journals are indexed on their first retention operation. Pending actions and outcomes whose audit has not been delivered are never dropped. The newest outcome also survives the hard cap even if it alone exceeds it. Hourly scan counters keep only the newest windows. Scan admission refuses pruned windows after a backward clock correction, preserving budget safety. Deleting an outcome ends the undo window for that action.
When an action cannot be proven, for example because the kernel did not answer
in time, its outcome stays uncertain and the engine refuses further firewall
changes rather than guess. The daemon retries recovery at startup and on its
maintenance tick, which settles the outcome as soon as the kernel can answer.
csm firewall actions shows what is waiting and why. If the kernel can never
prove it, an operator inspects the host and records the answer with
csm firewall actions resolve, which takes applied or rejected and an
optional note. Kernel evidence wins over the operator: if recovery can prove
the outcome at that moment, the proven one is recorded and the operator’s
answer is not used. The decision and its note reach the action log.
Startup recovery runs before applying the firewall. If startup fails, pending actions remain available for automatic recovery and operator resolution. After resolving them, restart the daemon to retry firewall setup. A storage error while saving kernel evidence stops resolution; it never permits an operator assertion to replace that evidence. Failure to refresh committed state is also reported, even when the outcome was saved.
Applying an action writes only the entries that change, so the kernel write for one block does not grow with the size of the blocked set. Large removals use bounded messages within the same atomic update. Ranged sets are written whole because their start and end markers move together. If the kernel already expired an entry the action meant to remove, the set is rewritten instead, and both paths end at the same state.
Firewall state storage provides complete snapshot reads and revision-checked replacement using the existing database. Reads preserve expired entries, original timestamps, explicit provenance, duplicate allow entries and collection order. An uninitialized store, corrupt data or a failed read returns an error without usable state. Rejected or rolled-back writes preserve the previous snapshot and revision. An error after the transaction body succeeds reports an uncertain commit; callers must reconcile stored state before retrying and must not assume rollback or confirmed durability. Legacy bucket edits invalidate a committed snapshot instead of silently changing its revision. The contract does not activate runtime cutover.
Firewall storage metrics separate write wait from transaction duration and expose read time, pending writes, failures and snapshot batch size using fixed labels. Repeatable benchmarks exercise competing writers and local snapshot copying. Backup restore and downgrade compatibility require the later migration stage.
Mutation failures
Subnet blocks refuse ranges overlapping loopback or link-local scopes,
infrastructure, host interface addresses, or full-IP and port-specific allows,
plus unspecified individual addresses and default routes. Other ranges beginning
at zero, such as 0.0.0.0/8, remain blockable. All these safety refusals are
recorded as refused in the action log, like single-address refusals. Invalid
targets and storage or kernel errors remain failures. Refusing a permanent
promotion leaves the prior temporary block and its expiry unchanged.
WAF attacker reports for link-local addresses stay visible but advise investigating the traffic source instead of a block CSM would refuse.
Integrity checks compare live rule structure with the snapshot captured after
CSM last applied the firewall. Editing or rehashing configuration alone never
approves a changed ruleset. A failed snapshot capture reports a monitoring gap;
after correcting nft availability, re-apply the firewall to restore its
baseline. Dynamic set membership does not affect this comparison.
Firewall changes persist their intent before changing kernel rules. A failed atomic kernel transaction restores the previous state. Failed writes and rollbacks return errors without success audit records; expiry cleanup logs the failure and retries on a later pass. DNS refreshes also retry failed removals. Removing one allow source preserves any other active source for that address.
An unconfirmed-durability error means the replacement is visible on disk but
storage did not confirm its survival across power loss. For kernel changes,
CSM attempts to restore the previous state before returning the error. A
failed rollback is reported as well. Correct the storage problem, inspect the
saved state and live rules, and repeat the intended operation or run
csm firewall restart to apply the saved state. Do not treat an error as a
successful rule change.
Port-specific allow additions and removals update saved state only. Run
csm firewall restart to apply them to the kernel; the command acknowledgement
states that a reload is required.
Clearing a stale local threat score
Run csm firewall forget <ip> as root against the running daemon to clear one
address’s accumulated local threat score. It accepts exactly one IPv4 or IPv6
address, including equivalent IPv6 spellings; CIDRs and extra arguments are
rejected. There is no dry-run option.
The command clears equivalent stored spellings together and reports their highest score and total event count. Blocks, allow lists, whitelists and raw event history remain intact. New findings immediately start a fresh scoring record, even if they arrive while the command runs. This does not suppress findings about the host’s own address.
Persistence is attempted immediately but remains best effort: a success reply confirms removal from memory, not durability across a restart. Check the daemon logs for attack database write failures if an old score returns after restart.
Attack statistics and event queries use the daemon’s state database or its configured attack database directory. If neither is available, they do not read event files from the working directory.
Startup failures
Overlapping, nested, duplicate, and adjacent ranges are merged for the kernel, including IPv4 and IPv6 ranges ending at the last address. This applies to infrastructure, country, Cloudflare, DoS exemption, and blocked subnet sets. Stored subnet entries keep their own source and expiry; removing or expiring one entry rebuilds the remaining coverage in an atomic transaction. Default routes remain forbidden as subnet blocks to prevent operator lockout.
CSM tries to initialize and apply an enabled firewall up to three times, with one-second and two-second delays between attempts. Each attempt builds a fresh atomic transaction. A failed apply keeps the previous kernel rules in place. The retry delays stop when the daemon shuts down.
If all attempts fail, CSM continues monitoring but reports degraded health.
/api/v1/status and csm status --json expose
automation.firewall_enabled: true, firewall_managed: false, and
firewall_startup_error. The error remains available for the lifetime of that
process. csm doctor reports a failed firewall check and a recovery step.
Inspect journalctl -u csm.service, correct the reported configuration or
nftables permissions problem, and restart csm.service. A successful startup
clears the error and enables the firewall-dependent services. Disabling the
firewall deliberately does not degrade health.
CLI Commands
# Status
csm firewall status # Show status and statistics
csm firewall ports # Show configured port rules
# Block / Allow
csm firewall deny <ip> [reason] # Block IP permanently
csm firewall allow <ip> [reason] # Allow IP (all ports)
csm firewall allow-port <ip> <port> [reason] # Allow IP on specific port
csm firewall remove <ip> # Remove from blocked and allowed
csm firewall remove-port <ip> <port> # Remove port-specific allow
# Temporary
csm firewall tempban <ip> <dur> [reason] # Temporary block
csm firewall tempallow <ip> <dur> [reason] # Temporary allow
# Subnets
csm firewall deny-subnet <cidr> [reason] # Block subnet
csm firewall remove-subnet <cidr> # Remove subnet block
# Search
csm firewall grep <pattern> # Search blocked/allowed IPs
csm firewall lookup <ip> # GeoIP + block status lookup
# Bulk operations
csm firewall deny-file <path> # Bulk block from file
csm firewall allow-file <path> # Bulk allow from file
csm firewall flush # Clear all blocked IPs (subnet blocks kept)
# Safety
csm firewall apply-confirmed <minutes> # Apply the firewall block from csm.yaml with auto-rollback timer
csm firewall confirm # Confirm applied changes
csm firewall rollback status|confirm|revert # Manage pending config rollback
csm firewall restart # Reapply full ruleset
# Profiles
csm firewall profile save|list|restore <name> # Profile management
# Audit
csm firewall audit [limit] # View audit log
# GeoIP
csm firewall update-geoip # Download country IP blocks
# Cloudflare
csm firewall cf-status # Show Cloudflare IP whitelist status
Configuration
Firewall defaults can be edited in two places:
- Web UI: Settings -> Firewall section. Port lists, rate limits, flood protection, deny caps, country block, and outbound SMTP restriction are all editable. Changes are restart-class. The save endpoint warns if the WebUI listen port is missing from
tcp_in. Theport_floodper-port rule list is YAML-only for now. - YAML: edit
/etc/csm/csm.yamldirectly. Runcsm rehashthensystemctl restart csm.
Tentative apply (rollback timer)
The Firewall section in the Web UI offers two save buttons. Save writes
the new config and prompts you to restart. Apply with rollback timer
writes the new config, restarts the daemon, and starts a timer (default 5
minutes, range 1-30). If you do not click Confirm before the timer
expires, the daemon restores the previous config and restarts again. This
protects against locking yourself out by, for example, removing the WebUI
port from tcp_in.
When the Web UI is unreachable (firewall mistuned, daemon broken), use the CLI escape hatch:
csm firewall rollback status
csm firewall rollback confirm
csm firewall rollback revert
Rollback state survives daemon restarts (the snapshot and its firewall configuration are persisted in the state directory). On startup the daemon checks for a pending rollback: if the deadline has already passed it restores the previous config and restarts; otherwise it restores the running firewall configuration and rearms the timer for the remaining window. Backup and store export omit this transient state, so restoring an archive cannot re-arm an old confirmation window.
firewall:
enabled: true
ipv6: false # false = ALL IPv6 traffic bypasses the firewall; CSM raises a finding on dual-stack hosts
conn_rate_limit: 200 # new connections per minute per source (IPv6 per /64; 0 = disabled; null = default)
syn_flood_protection: true # per-source SYN flood meter (IPv6 per /64)
conn_limit: 400 # max concurrent connections per IPv4 source (0 = disabled)
smtp_block: false # restrict outbound SMTP
log_dropped: true
dyndns_hosts: # resolved every 5 min and whitelisted
- "monitoring.example.com"
Full firewall reference: Configuration - Firewall.
Auto-response interaction
Auto-block calls require firewall.enabled: true because they go through the firewall engine. The engine consults two policy hooks first:
-
auto_response.verdict_callback- when enabled, the engine POSTs a signed JSON request to the panel after local validation and infra-IP safety checks. When a secret is configured, CSM rejects unsigned callback replies by default. The panel can downgrade toallow(audit-only), attachtenant_idfor downstream correlation, or add a note. CSM fails open on hook errors. Wire contract:docs/verdict-callback-contract.md. -
auto_response.dry_run- when true (or absent; safety default),BlockIP()records the intended block to bbolt and returns success without touching nftables. Manualcsm firewall ...operator commands bypass viaBlockIPForceand always apply. Verify withcsm firewall statusafter policy changes; “Recently Blocked” timestamps newer than the last restart confirm live mode. See Auto-response - Dry-run safety default.
Subnet blocks refuse the default route and any range that contains an infrastructure IP, a resolved infra hostname, a local host address, a full-IP allow, or a port-specific allow. Remove the allow or narrow the CIDR before applying the block.
Allowlist precedence
The nftables input chain accepts infra_ips first, then drops
blocked_ips, then accepts allowed_ips. Because the drop is evaluated
before the allowed_ips accept, an allowlisted IP that lands in
blocked_ips would still be dropped. The same applies to port-specific
allows, because those rules are evaluated after blocked_ips too. To keep
operator allows effective, the auto-block path refuses to add an IP to
blocked_ips when it is on allowed_ips (set by csm firewall allow), has
a port-specific allow (csm firewall allow-port), or is in a verified-bot
range (built-in or reputation.verified_bots). Precedence:
infra_ips- hard protect. Never blocked by anything, auto or manual; subnet blocks containing one are refused.allowed_ips, port-specific allows, and verified-bot ranges - soft allow. The auto-block path skips them, but an explicit operator deny (csm firewall deny, Web UI manual block) still applies, because operator commands go throughBlockIPForceand bypass the soft-allow gate.
An operator block carries the lifetime the operator chose. csm firewall deny and the Web UI permanent block never expire; the Web UI 24 hour block
expires in the firewall and takes its threat-database evidence with it, so
the address does not keep scoring as malicious after the block is gone.
The Web UI 24 hour action refuses to shorten a permanent or longer block; unblock explicitly before changing its lifetime. Undo of a bulk block or unblock restores each original deadline, including permanent blocks, and never extends an expired block. A later operator decision supersedes the older undo action.
Lockout warnings
Config validation warns when an enabled firewall would cut off the management
plane. csm doctor, daemon startup, and the Web UI save path all run the same
checks, and the Web UI returns each warning once:
- The enabled Web UI port is missing from
tcp_in, or from an explicittcp6_inoverride when IPv6 filtering is enabled. An emptytcp6_ininheritstcp_in. - A
restricted_tcpentry also appears in an effective public TCP allow list, but noinfra_ipsare configured. The restricted list only filters public accepts; it does not open ports itself. Matching ports are therefore reachable only through the port-agnostic infrastructure-IP accept rule, and with no infrastructure addresses they are reachable from nowhere.
These stay warnings and never block a save or a start: fronting the Web UI with
a reverse proxy or reaching it over a VPN are legitimate reasons to leave the
port out of tcp_in.
One more warning compares the policy against the host instead of against the
config, so it runs in csm doctor, csm validate --deep, and the Web UI
firewall save rather than on every load:
- sshd listens on a port that
tcp_in(or an explicittcp6_in) does not allow. The shippedtcp_inleaves 22 out, because many hosts move sshd, so a host that never moved it loses SSH on the first apply. EveryPortdirective counts, including ones inIncluded drop-ins under/etc/ssh/sshd_config.d/, and a port named inrestricted_tcpwithinfra_ipsset is treated as a deliberate infra-only listener.AddressFamilyandListenAddresslimit the check to the IP families sshd exposes, and loopback-only listeners do not count because the inbound firewall always accepts loopback traffic. Hosts with no sshd config get no warning.
Egress
tcp_out is default-drop as well and ends in a TCP reset, so a host whose
policy omits a port it dials does not lock an operator out; it goes silent.
Every heartbeat, finding delivery or intel lookup fails at once with
“connection refused”, which reads like the far end being down, while the host
looks healthy locally. The same validation pass therefore warns when an
enabled firewall’s outbound policy would refuse a connection the daemon
itself needs:
- The port of every enabled outbound endpoint in the config is checked
against
tcp_out:alerts.email.smtp,alerts.webhook.url,alerts.heartbeat.url,alerts.audit_log.syslog.address(tcp and tls transports),auto_response.verdict_callback.url,reputation.rspamd.url,reputation.upstream.url,reputation.report.targets[].url,reputation.central.set_url,signatures.update_url,signatures.yara_forge.download_url,sentry.dsnandupdates.github_api_url. HTTP endpoints use their explicit numeric port, or the scheme default when the port is omitted. SMTP and TCP/TLS syslog addresses also accept TCP service names, matching their dialers. Disabled features are skipped, and so are loopback destinations, which the output chain accepts ahead of any port rule. - Port 443 is checked once for the built-in HTTPS endpoints (threat feeds, AbuseIPDB, MaxMind, YARA Forge, AI-crawler range feeds, release check), because dropping it silences all of them at once.
- Every port under
firewall.required_tcp_outis checked. That list is a declaration, never added to the policy: a conf.d fragment owned by an integration can state the ports its service needs, andcsm doctorreports when the effective policy drops one instead of the operator discovering it from a silent node. The check runs against the mergedtcp_out, so any config layer that permits the port satisfies the declaration.
On a restricted output chain, smtp_block installs per-user accepts for the
mail ports ahead of the port rules, and those ports never get a port rule of
their own. The daemon runs as root, so its alert mail is not warned about when
tcp_out omits a port that smtp_block still lets root reach. A port
declared for another service still warns under smtp_block, because a port
declaration cannot prove that service’s user is allowed. When IPv6 is managed,
an explicit tcp6_out is checked separately; an empty one inherits tcp_out,
and the single warning covers both families. When only the IPv6 lists are set,
IPv4 egress is accepted wholesale and the warning names tcp6_out. A literal
IPv4 or IPv6 destination is checked only against its own family.
This catches the daemon’s own egress and whatever has been declared. It does
not see what an arbitrary third-party process on the host dials; an agent
that ships its own conf.d fragment should declare its ports under
required_tcp_out there.
Value validation
Unlike the warnings above, these are errors, because the value cannot do what the operator meant:
- Ports outside 1-65535 in any port list, including
drop_nologand the passive FTP range, and a passive FTP range that ends before it starts. country_blockentries that are not two-letter ISO codes.port_floodentries with an out-of-range port, a protocol other thantcporudp, or a non-positive hit count or window.
Two enums elsewhere get the same treatment because their consumers fall back to
a default branch rather than failing: reputation.central.action (an
unrecognised value became a challenge policy) and
incidents.*.block_at_severity, which accepts only high or critical and
silently disabled incident blocking on anything else.
Incident auto-block escalation
An incident-driven block used to be requested once per incident and never
again. The block it applied expired after auto_response.block_expiry, but the
marker saying “already blocked” did not, so an attacker who kept going past the
expiry was never blocked a second time while the incident stayed open and kept
collecting evidence.
A block is now re-requested whenever the previous one has lapsed and the incident is still active and still receiving qualifying findings, and each request lasts longer than the last:
| block | lifetime |
|---|---|
| first | auto_response.block_expiry (24h by default) |
| second | 7 days |
| third and later | permanent |
A permanent block is never re-requested. Closing an incident, by an operator or by the stale-incident sweep, resets the ladder, so a later recurrence starts at the bottom rather than inheriting a months-old episode. Concurrent findings still collapse into one firewall call, and a declined or dry-run request is not recorded, so it can retry.
The ladder survives restarts and quiet intervals while the incident remains active. Closing the incident, manually or automatically, resets it; a pending block callback cannot restore the old ladder after that close.
The incident view carries a Block button whenever the incident has one
unambiguous source address. Mixed-source or truncated timelines without an
address in the correlation key do not offer a block target. The button asks
for confirmation, blocks permanently, notes the block on the incident timeline as
operator_block, and settles the ladder so the automatic hand-off does not
re-request a block for an address the operator just blocked.
The block API accepts an optional incident_id. An invalid, unknown, or
address-mismatched incident ID does not prevent the firewall block, but does
not change the incident. Refreshing an existing temporary block does not
advance the escalation rung. Blocks recorded on closed incidents remain audit
actions without restarting the ladder.
Infrastructure IP DNS guard
Hostnames listed in top-level infra_ips or firewall.infra_ips are resolved every 5 minutes and their current addresses feed the infra auto-block guard. If a hostname stops resolving, the daemon emits an infra_ips_unresolvable Warning finding and keeps the last known addresses protected during the grace period (default 10 min). This prevents a transient DNS outage from deprotecting the management plane. The finding auto-clears when resolution recovers.
DoS-exempt ranges
Operators can declare IP ranges that bypass the new-connection rate meter for their IP family, the IPv4 concurrent connection-limit, and mail-port flood meters, preventing false-positive throttling and subnet auto-blocks for carrier CGNAT pools or mail-provider egress. Configure under firewall.dos_exempt_ranges (your own CIDRs) and firewall.dos_exempt_known_mail_providers (adds Google and Microsoft mail ranges, on by default). See Configuration - firewall.dos_exempt_ranges.
What exempt sources bypass
Sources in the exempt set skip three categories of metering:
- Connection rate-limit - the new-connection rate meter (configured via
conn_rate_limit) does not apply for the source’s IP family. - Concurrent connection-limit - the IPv4 concurrent connection cap (
conn_limit) does not apply. - Mail-port flood meters - the
port_floodrules on TCP 25, 465, and 587 do not apply for the source’s IP family.
Subnet auto-block (spray, ASN-crawl, and netblock escalation) also skips any subnet block whose CIDR intersects an exempt range, and exempt IPs are excluded from the per-subnet threshold count so they cannot push a subnet over the netblock limit. Auto-response subnet blocks whose range falls inside an exempt range are removed automatically at daemon startup and at the start of each auto-block cycle. Manually created IP and subnet blocks are never pruned, even if they fall inside an exempt range.
What exempt sources do not bypass
The following protections remain in force regardless of exempt status:
- Manual blocks -
csm firewall deny <ip>andcsm firewall deny-subnet <cidr>go throughBlockIPForce, which bypasses the exempt check. An IP or range that is both exempt and manually blocked is still dropped. - SYN flood protection - the SYN flood meter is not affected by the exempt set.
- UDP flood protection - the UDP flood meter is independent of the exempt set.
- Country blocking - country CIDR blocks apply unconditionally.
- Port policy -
tcp_in,tcp_out, andrestricted_tcpport rules are not modified.
The rule ordering that makes this work: the nftables input chain evaluates blocked_ips (and subnet blocks) before the DoS-meter rules. So a manual block inside an exempt range still drops the traffic – the block is hit before the meter that exempt sources bypass.
Dynamic mail-provider ranges
When dos_exempt_known_mail_providers is true (the default), the daemon resolves Google and Microsoft outbound mail ranges at startup and pushes them into the firewall exempt sets before the first rule application. The ranges are discovered from the providers’ published SPF records (the Google and Microsoft mail SPF roots), so they track provider changes without a CSM update. They are cached on disk so the previous set is available immediately on subsequent starts. A built-in snapshot is used if the cache is missing or the first live refresh has not completed. The cache is refreshed every 12 hours; if a refresh fails or the nftables reapply fails, the previous overlay is preserved unchanged.
ModSecurity Integration
CSM detects ModSecurity (WAF) on Apache, Nginx, and LiteSpeed across cPanel and plain Linux hosts. Custom rule deployment, override writes, and reload management are currently cPanel-only; other platforms still receive status, staleness, and event detection.
Supported Web Servers
| Web server | Config candidates | Status check | Custom rule deployment |
|---|---|---|---|
| Apache on cPanel EA4 | /usr/local/apache/conf/*, /etc/apache2/conf.d/modsec*, whmapi1 modsec_is_installed | Yes | Yes (via cPanel modsec user conf) |
| Apache on Debian/Ubuntu | /etc/apache2/mods-enabled/security2.conf, /etc/apache2/conf-enabled/*, /etc/apache2/conf.d/modsec2.conf | Yes | No |
| Apache on RHEL/Alma/Rocky | /etc/httpd/conf.d/mod_security.conf, /etc/httpd/conf.modules.d/* | Yes | No |
| Nginx on any distro | /etc/nginx/nginx.conf, /etc/nginx/modules-enabled/50-mod-http-modsecurity.conf, /etc/nginx/modsec/main.conf | Yes | No |
| LiteSpeed | /usr/local/lsws/conf/httpd_config.xml, /usr/local/lsws/conf/modsec2.conf | Yes | cPanel only |
When ModSecurity is not installed, the waf_status check emits a platform-specific install hint:
# On Ubuntu + Nginx:
Install: apt install libnginx-mod-http-modsecurity modsecurity-crs
# On Ubuntu + Apache:
Install: apt install libapache2-mod-security2 modsecurity-crs && a2enmod security2
# On AlmaLinux + Apache:
Install (requires EPEL): dnf install -y epel-release && dnf install -y mod_security
# On AlmaLinux + Nginx:
Install (requires EPEL): dnf install -y epel-release && dnf install -y nginx-mod-http-modsecurity
# On cPanel:
Install: WHM > Security Center > ModSecurity
Rule-staleness alerts scan both the flat CRS layout (/usr/share/modsecurity-crs/rules/*.conf) used by distro packages and cPanel vendor trees, including nested layouts such as modsec_vendor_configs/VENDOR/rules/*.conf. On cPanel, CSM maps WHM’s active configuration files to their vendor trees and checks the newest artifact in each loaded tree. This honors individual configuration overrides, keeps retired trees out of the result, and prevents a fresh unloaded vendor from hiding a stale loaded one. LiteSpeed also keeps the on-disk check as a backstop while cPanel rebuilds the active configuration list and rule tree. If WHM cannot provide the mapping, and on other platforms, the check keeps the conservative oldest-artifact behavior. Update instructions are platform-specific (apt update && apt upgrade modsecurity-crs, dnf upgrade modsecurity-crs, or WHM on cPanel).
Features
- Custom CSM rules - IDs 900000-900999 in
configs/csm_modsec_custom.conf(cPanel only today) - Rule override management -
SecRuleRemoveByIddirectives for false positive suppression - Escalation control - change rule severity or action per-rule
- Live deny escalation - repeated ModSecurity deny events from one IP emit an escalation finding that feeds auto-response blocking. CSM-owned rules keep their existing per-rule escalation controls.
- Disabled-scope detection - reports domains and accounts with the engine switched off, covering both the userdata flag and the per-account and per-domain config includes used by Apache and LiteSpeed in the std and ssl trees
- WAF event log parsing - correlates events by IP, URI, and rule ID
- Hot-reload - apply changes without Apache restart (cPanel only)
- Rule activation - ModSecurity reads rules only when the web server starts or reloads. When CSM’s installed rule sections change, for example after an upgrade or
csm install, the daemon runsmodsec.reload_commandat startup or during its WAF check. Each rule change reloads once. Standalonecsm checkruns never reload. A failed reload raises awaf_statuswarning and is retried. Without a command, CSM only warns at startup and the rules wait for the next web server restart.
CSM checks the rule-action registry every five minutes and rebuilds it when rule file contents change. Read failures leave the build uncached so the next check retries. Files that were read in full but exceeded the parser’s line limit are reported without forcing unchanged files to be parsed again. Rules appended after parsing reaches the end of a file are picked up on the next check.
The reload command runs inside CSM’s systemd sandbox. Use a service-manager
command such as systemctl reload lsws for LiteSpeed, so the web server’s service
performs the reload. Direct reload scripts inherit CSM’s filesystem restrictions
and may fail. A LiteSpeed reload restarts workers and can briefly raise load.
The LiteSpeed Cache role-simulation filter covers privileged routes and writes, including WordPress REST method overrides, when requests carry simulation cookies with a weak hash. Ordinary public GET/HEAD crawling remains allowed. This is a request-scoped mitigation: public reads still run under the simulated identity, so upgrading the vulnerable plugin remains necessary.
The usual hash discriminator is 1-16 alphanumeric characters versus the fixed plugin’s 32-character hashes. Numeric equivalents are also filtered because the vulnerable plugin compares hashes loosely; padding or exponent notation must not turn a weak hash into an exempt one. These checks also keep public crawler reads allowed.
The WordPress user enumeration filter blocks anonymous requests for the REST
users route, whether the route follows wp-json/ in the path or starts at
wp/v2/users in the rest_route query parameter, in any letter case. Other
page paths and REST namespaces are left alone. Query values are matched as
already decoded by the query parser. Requests that carry an Authorization
header or a WordPress logged-in cookie pass, so admin screens, the editor and
Application Password clients keep working. Both are presence checks: the filter
turns away anonymous scanners, and WordPress still decides who is signed in.
The filter runs in phase 1 and does not inspect request bodies. A route supplied
only in a POST body, including with a REST method override, is outside its scope.
Sites that need to restrict the public users endpoint must enforce that policy
in WordPress. Disabling rule 900112 disables this filter for both route forms;
its helper rules only set transaction-local flags and do not block requests.
For Apache ModSecurity v2 regression validation, run
python3 scripts/test-litespeed-modsec.py in a disposable Debian Linux environment
with apache2, libapache2-mod-security2, libapache2-mod-php, and python3
installed. The test loads the complete shipped configuration and exercises HTTP
requests through the actual engine and PHP parser. It does not establish
LiteSpeed runtime compatibility; verify changed rules on the supported LiteSpeed
engine before deployment.
Web UI Pages
ModSecurity (/modsec) - WAF status overview, event log, active block list, filterable by time range, minimum severity, and source country
ModSec Rules (/modsec/rules) - per-rule management:
- View the CSM rules with descriptions and hits in the last 24 hours
- Enable or disable individual rules; changes are staged and applied with one reload
- Turn firewall escalation off for a rule: the rule still denies the request, but CSM does not block the IP in the firewall
- Escalation exclusions are listed and edited here even when rule management (
modsec.rules_file,modsec.overrides_file,modsec.reload_command) is not configured, since the daemon applies them either way
API Endpoints
GET /api/v1/modsec/stats WAF statistics
GET /api/v1/modsec/blocks Blocked request log
GET /api/v1/modsec/events WAF event details
GET /api/v1/modsec/rules Loaded rules list
POST /api/v1/modsec/rules/apply Apply the set of disabled rules and reload
GET /api/v1/modsec/rules/escalation Rule IDs excluded from firewall escalation
POST /api/v1/modsec/rules/escalation Exclude one rule from escalation or turn it back on
Signature Rules
CSM uses YAML rules for real-time scanning and finding re-checks. Optional
YARA-X rules also run during deep scans and email attachment scanning. Rules
are stored in /opt/csm/rules/. An engine that loads zero rules (a mistyped
signatures.rules_dir, an empty rule sync) raises a realtime_rules_missing
finding at startup and after every reload, because both engines otherwise
treat an empty directory as a successful load and scan every write against
nothing.
Deep scans are rolling: each scheduled run resumes from a persisted cursor and scans as much as fits in its time budget, so the whole content set is covered across runs even when a single run cannot finish it. A warning finding is raised if no full pass has completed within 30 days.
The same rolling walk also feeds the JavaScript keystroke taint analyzer (js_keylogger_dataflow) and the PHP remote-source taint analyzer (php_remote_taint); see the deep checks reference for both. Each consumer keeps its own cursor and completion record, so none of them stalls the others: a missing or failed YARA backend does not hold up either taint analyzer, and the PHP analyzer being unavailable does not affect YARA or JavaScript coverage.
A rule declaring file_types: [".php"] is applied to every extension a stock PHP handler executes (.php2 through .php8, .phtml, .pht) and to .phps, because that source-view extension still contains PHP source. .phps stays outside the set of extensions CSM treats as executable, so this does not change which files the real-time dropper tracker considers runnable.
Both engines skip ZIP, gzip, bzip2, xz, 7z, and RAR containers, but only when the file name carries an archive extension as well as the archive signature. Matching compressed bytes or stored filenames produces false positives without inspecting the archived file, so real containers are scanned when they are extracted onto monitored storage instead. A file whose name a web server would execute is always scanned whatever its first bytes look like, because PHP echoes any leading bytes and runs the rest. Uncompressed tar files and executable PHP archives (PHAR) remain scannable, and filename-based phishing-kit archive detection is unchanged.
YAML Rules
The WordPress REST API exploit signature is a YAML-only heuristic. It requires a users-endpoint URL literal and a password query parameter or a nearby PHP, JSON or form-encoded password field, including payloads prepared before the URL. Endpoint prose alone does not qualify. This is bounded textual evidence, not PHP data-flow analysis or proof that a request is unauthorized.
rules:
- name: webshell_c99
severity: critical
category: webshell
file_types: [".php"]
patterns: ["c99shell", "c99_buff_prepare"]
min_match: 1
- name: phishing_login
severity: high
category: phishing
file_types: [".html", ".php"]
patterns: ["password.*submit", "credit.*card.*number"]
exclude_patterns: ["legitimate_form_handler"]
min_match: 2
Fields:
name- unique rule identifierseverity- critical, high, or warningcategory- webshell, backdoor, phishing, dropper, exploitfile_types- file extensions to match (or["*"]for all)patterns- case-insensitive literal stringsregexes- regex patterns, compiled case-insensitivelyexclude_patterns- literal patterns that suppress a match (false positive reduction)exclude_regexes- regex patterns that suppress a matchmin_match- minimum total number of matching literal and regex entries; each entry counts oncerequire_regex- require at least one regex among the matches counted towardmin_matchmax_file_bytes- skip this rule when the complete scanned file is larger than the byte limit; omitted or0is unboundedmax_file_bytes_exempt_regexes- high-confidence regexes that let the rule continue normal evaluation abovemax_file_bytes
A regex only runs on a file that contains the fixed text every match of it
must contain, compared without regard to case: for eval\s*\(\s*base64_decode
that is both eval and base64_decode. Write regexes around distinctive words
such as function names: a regex with no fixed text of two or more characters
runs over every file of its types, and realtime pays that cost on each write.
Required text can include escaped bytes such as \x00; these remain part of
the literal when checking whether a file could match.
When a regex includes a literal listed in patterns, the same content can
satisfy both entries. Use independent entries when a rule needs multiple pieces
of evidence. The bundled HTTP tunnel rule requires both socket creation and a
CONNECT request. The legacy PHP callback rule uses the same narrow signature in
YAML and YARA-X: a direct function call with a quoted parameter list, a variable,
null, a simple array lookup or a short helper call as its first argument.
The second argument is a decoder or request lookup, optionally preceded by one
concatenated literal. Quoted lists can contain commas, semicolons and escaped
quotes. Array indices accept a single quoted key or an unquoted scalar. Helper
arguments accept at most one literal among unquoted scalar operands, including
implode(',', $args). These bounded forms consume quoted operands whole and
exclude comments, interpolation and nested expressions, so delimiters inside
data cannot supply the body-source evidence. Double-quoted array keys, helper
literals and body prefixes must escape dollar signs; unescaped dollars require
interpolation analysis.
Shared positive and benign fixtures check both engines. Generated socket and
funchand wrappers and ordinary legacy callbacks stay silent under these rules.
The PHP goto-obfuscation rule requires three independent signals in both engines: a PHP opening tag, at least nine jumps to digit-bearing generated labels or eleven to alphabetic labels, and a decode call, execution call, dynamic include, or request input. Fixed-path bootstrap includes do not supply this evidence. Variable and array callback calls count even when the function name is constructed and the argument is a literal. Comments between a callable and its opening parenthesis do not hide the call. Line breaks and keyword case do not change the label counts. Long encoded strings and data URIs alone are not execution evidence. These are source-text heuristics, not PHP dataflow analysis; they cannot resolve arbitrary dynamically generated code or distinguish every benign use of these operations.
The PHP content heuristic that runs during scans applies the same evidence
rule: generated goto labels and descriptive goto labels both need a decode
call, execution call, dynamic include, or request input before they count as
an obfuscation indicator. call_user_func is deliberately not evidence in
either place, because plugin loaders dispatch their own callables through it.
The content heuristic uses the signature’s evidence expression, including
comment-separated calls, grouped and array callbacks, and case-sensitive PHP
superglobal names. A regression check guards against expression drift.
Legacy callback parser follow-up
The callback signature does not inspect quoted function bodies. Doing so needs
PHP string decoding, tokenization and expression analysis: for example,
assert($x > 0) is an ordinary boolean check, and "eval($x)" can be data.
The rules scan source text, so they do not promise general PHP comment or string
awareness, nor complete coverage of dynamically generated code.
The following cases were covered by the expanded regex and are deliberately outside the narrowed signature. They remain acceptance cases for a parser follow-up, not claims of current detection by this rule:
| Deferred case | Examples to restore |
|---|---|
| Constructed parameter lists | Concatenation, nested array indices, helper calls with multiple quoted operands, and chr(100/(1+1)) as the first argument; only the bounded simple forms above are covered |
| Comments inside constructed parameters | Comments in array indices or helper arguments, especially those containing closing delimiters or body-source names |
| Interpolated parameter operands | Double-quoted array keys or helper operands containing unescaped dollars, including interpolation with nested quoted keys |
| Comments between arguments | Block comments with embedded commas, and line comments before the body source |
| Constructed body expressions | Grouped concatenation, parenthesized decoders and trim(base64_decode($payload)); a single literal concatenated onto request input is covered |
| Interpolated bodies | A double-quoted body such as "return {$_POST['code']};", or a concatenated double-quoted prefix containing unescaped dollars |
| Literal executable bodies | eval($x) or string-capable assert($x), with statements, strings or comments before them; both outer quote styles and escaped quotes |
| Literal expression contexts | return, or, do, case, include, include_once, require, require_once, clone, yield from, comparisons, shifts and inequality before an execution sink |
| Literal lexical edges | Global assert, comment backslashes before * or */, and quoted operands before a comparison |
| Outer call contexts | Calls immediately following a ternary colon, case-label colon or comparison operator |
This work belongs in internal/phptaint, which already uses VKCOM/php-parser
and records the second argument of create_function as a sink. Extend that
analysis to decode and parse statically known callback bodies under its existing
budgets, distinguish boolean assertions from string execution, and report
unresolved dynamic bodies as analysis gaps. It currently feeds a separate
scheduled check; it is not a post-filter for YAML or YARA. Realtime coverage
would need explicit integration and latency tests, and standalone YARA would
still have the narrower coverage.
Acceptance requires restoring the deferred cases as positive parser fixtures, retaining the shared benign fixtures, checking both quote styles and comment forms, and passing the clean-corpus gate without new baseline entries. Do not expand another regex into a PHP tokenizer to recover these cases.
YARA-X Rules (Optional)
Build CSM with YARA-X support:
CGO_LDFLAGS="$(pkg-config --libs --static yara_x_capi)" go build -tags yara ./cmd/csm/
The rules directory and rule files must be owned by root or the scanner user and
must not be group- or world-writable. The standard service and its YARA worker
run as root; the installer uses 0750 for the directory and 0640 for shipped
rules. Custom rules must also remain readable by the scanner.
Place .yar or .yara files alongside YAML rules in /opt/csm/rules/. CSM compiles them at startup and uses them for:
- Real-time fanotify file scanning
- Scheduled deep-scan filesystem sweeps
- Email attachment scanning
Scans hand every file to YARA-X regardless of its name, so a rule may decide on the leading bytes rather than the extension. The bundled rule for PHP carried inside an image does exactly that: it requires a genuine PNG, JPEG, GIF, WebP, ICO, BMP or TIFF container, a PHP opening tag, and an execution or remote-fetch construct beside it. A picture is a working payload store, and nothing about its name says so.
Its other half is the loader: a PHP file that includes or requires a
non-executable file while reading request input. The extension list covers
images, archives and opaque data files; source partials such as .html,
.tpl, .txt and .svg are deliberately absent, because including those
can be ordinary templating. Literal targets must end the include statement;
an image named later in a rendered HTML expression is not an include target.
Encoded targets remain suspicious regardless of the hidden extension.
Findings include up to five distinct non-executable path literals near an
include or require. These are candidate payload references: proximity does
not prove variable identity, and concatenated paths may be partial. Extraction
advances through the content once and stops when the output cap is reached.
Scheduled YARA sweeps scan regular, non-empty files under configured
account_roots, or cPanel /home/*/public_html roots when no override is
configured. Files larger than thresholds.full_scan_max_file_mb, symlinks,
and special files are skipped. An unreadable or oversized file, or a scanner
backend error, emits yara_scan_incomplete and preserves findings from the
previous complete sweep. Scheduled findings and real-time fanotify findings
have separate ownership, so one path cannot purge the other’s results.
A real-time scan that cannot inspect a changed file emits
yara_realtime_scan_error instead, so a scanning outage stays separable from
the scheduled coverage report. The dashboard’s Components matrix shows this
finding as Fanotify’s last event; scheduled coverage reports do not advance
that event time. This error finding is suppressed while the file monitor is
stopping: the YARA backend is stopped before the file monitor has finished
draining, and a clean restart is not an outage.
Without the yara build tag, YARA rules are not loaded or evaluated.
Updating Rules
Before merging bundled YARA changes, run the rules against an unpacked corpus of clean WordPress core and plugin files:
YARA_FP_CORPUS=/path/to/corpus go test -tags yara ./internal/yara/ -run TestRepositoryRulesAgainstCleanCorpus -v
The same corpus measures the YAML rules that only realtime and finding re-check run:
YARA_FP_CORPUS=/path/to/corpus go test ./internal/signatures/ -run TestRepositoryYAMLRulesAgainstCleanCorpus -v
The same corpus also checks that no YAML regex is skipped on a file it matches:
YARA_FP_CORPUS=/path/to/corpus go test ./internal/signatures/ -run TestYAMLGatesSoundOnCleanCorpus -v
Both gates require at least 5,000 non-empty files within the default scheduled scan size limit. Rule-load, traversal, and read failures fail the relevant run instead of counting as clean; YARA backend errors do too.
The measured YARA baseline is empty. The YAML baseline records rules that already fire on clean plugin and core code and are named in the realtime-rule porting backlog. Tighten a noisy rule rather than excluding paths or filenames.
csm update-rules # download latest rules and reload the running daemon
csm update-rules now asks the daemon to reload through the control socket once the download completes. If the daemon is not running, the next start picks the files up automatically. kill -HUP $(pidof csm) still works.
Or from the web UI: Rules page > Reload Rules button.
Remote rule updates are now signature-verified. Any configuration that enables signatures.update_url or signatures.yara_forge.enabled must also set signatures.signing_key to the 64-character hex-encoded Ed25519 public key that verifies the downloaded .sig files. Both URLs must use https; a plain-http URL fails validation.
A valid signature proves who published a rules file, not that it is the current one. The YAML updater therefore also refuses a download whose version is lower than the installed file’s, or that carries fewer than half as many rules, and keeps the installed rules. The daemon emits a Critical finding for either refusal so a stale mirror or a replayed old release is visible instead of silently stripping detection. A missing or unparsable installed file is not compared, so a signed update remains the way to recover from a corrupt rules file.
For an intentional reduction below half the installed YAML rule count, set signatures.allow_rule_count_decrease: true, restart CSM, and run the signed update. Then return the setting to false and restart CSM again; this exception never permits a version downgrade.
Remote update URLs must use https and must not point at localhost, loopback, link-local, unspecified, or RFC1918 / ULA private addresses.
YARA Forge Integration
CSM can automatically fetch curated YARA rules from YARA Forge, which aggregates and quality-tests rules from 40+ public sources including signature-base, Elastic, Malpedia, and ESET.
Configuration
signatures:
signing_key: "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
yara_forge:
enabled: true
tier: "core" # core (5K rules, low FP), extended (10K), full (12K)
update_interval: "168h" # weekly
download_url: "https://mirrors.pidginhost.com/csm/yara-forge/{version}/yara-forge-rules-{tier}.zip"
disabled_rules: # rule names to switch off, in Forge and in the shipped rules
- SUSP_Example_Rule
The project operates the signed mirror shown above. A ready-to-use drop-in is shipped at /usr/lib/csm/profiles/yara-forge.example.yaml; copy or include it under /etc/csm/conf.d/ to enable Forge without editing the main csm.yaml. The matching signing_key is the project Ed25519 public key (hex), published on the release signing page.
signing_key must be a hex string for the Ed25519 public key that matches the private key used to sign the remote Forge artifact. It is not a PEM block and not a file path.
YARA Forge’s upstream GitHub releases publish ZIP files, but not CSM detached signatures. CSM therefore requires yara_forge.download_url to point at a mirror you operate. The URL may contain {tier} and {version} placeholders. The detached signature must be available at the resolved ZIP URL plus .sig.
When the URL has a standalone {version} path segment, CSM first reads a plain-text latest pointer from that version directory’s parent, for example https://mirrors.pidginhost.com/csm/yara-forge/latest. If the pointer is not published, CSM falls back to the upstream GitHub release tag. Mirror network errors and server errors fail the update instead of falling back, because GitHub may name a release the mirror has not signed yet. CSM accepts only short tag tokens from the pointer; the ZIP signature still gates installation.
If you do not have a signed update source yet, disable remote updates instead:
signatures:
signing_key: ""
update_url: ""
yara_forge:
enabled: false
Tiers
| Tier | Rules | Size | False Positive Risk |
|---|---|---|---|
| core | ~5,000 | 1.6 MB | Low (quality >= 70, score >= 65) |
| extended | ~10,500 | 3.3 MB | Medium |
| full | ~11,600 | 3.7 MB | Higher (includes score >= 40) |
Update Flow
- CSM resolves the latest YARA Forge version from the mirror pointer, falling back to GitHub only when no pointer is published
- If different from the installed version, downloads the ZIP for the configured tier from
yara_forge.download_urland its detached signature - Verifies the download against
signatures.signing_key - Filters out any rules listed in
disabled_rules - Compile-tests the rules with YARA-X before installing
- Atomically replaces the previous Forge rules file
- Reloads the YARA scanner
Custom rules in malware.yar are never overwritten by the Forge fetcher.
Disabling Rules
If a rule produces false positives, add its name to disabled_rules in the config and restart the daemon:
signatures:
disabled_rules:
- SUSP_XOR_Encoded_URL
- php_goto_obfuscation
The list covers every rule CSM loads, not only YARA Forge: Forge rules are
stripped from the download, rules shipped in malware.yml are skipped when
the real-time engine loads them, and rules shipped in malware.yar are
stripped before the scheduled engine compiles them. The self-test measures
the ruleset that is left, so disabling a rule shows up as the coverage it
costs.
Names are matched in full, ignoring surrounding whitespace and letter case; prefixes and substrings do not match. Removing a YARA rule preserves neighboring rules, including when declarations share a line or literals contain braces. The YARA worker receives the daemon’s effective disabled list and configuration paths; rule reloads and worker crash recovery retain that list.
A valid ruleset whose rules are all disabled loads as an empty set, including on reload; stale rules are not retained. The self-test reports the resulting misses. An empty or invalid replacement still reports a load error.
csm validate warns about a name that matches no rule, because a typo here
otherwise reads as “that rule is off” while the rule keeps firing. Disabling
a rule is a last resort and a standing gap in coverage; prefer fixing the
rule.
Validation recognizes a disabled YAML rule even if its regular expression is invalid, and counts repeated names only once.
Signature settings require a daemon restart. SIGHUP does not apply a changed disabled list; rule-file reloads keep the current list.
How Rules Avoid False Positives
Signature rules require structural nesting, not co-presence of strings. Two dangerous function calls appearing in the same file but in unrelated code paths won’t trigger a rule. The call must directly wrap or chain with the other for a match.
YARA-X rules cannot express nesting, so multi-string rules state how their evidence is related. A bounded distance is used when closeness is part of the malicious shape. When valid malware can carry padding between its signals, PHP diagnostic records are kept as separate contexts so one record cannot borrow evidence from another. Anchoring works the same way: a CGI webshell rule requires its shebang at offset 0, because that is where the web server needs it. None of these controls refers to a file’s name or path.
Realtime signature auto-quarantine adds a safety gate: only webshell and dropper matches are eligible, and the file must be at least 512 bytes and either have Shannon entropy >= 5.5 or hex density > 20% plus an obfuscated-execution signal. Legitimate plugin code (well below 5.5 entropy) passes through; obfuscated malware (5.8+) is caught.
Alert Rate Limiting
Default: 30 operator alert dispatches/hour (configurable via max_per_hour). CRITICAL findings and threat-intel reputation sightings always get through by email or generic webhook regardless of the rate limit, and they do not count against it. Other lower-severity alerts are rate-limited, including when they are batched with a CRITICAL finding: once the budget is spent, only the urgent findings in that batch are sent.
A dispatch uses one slot only when email or a generic webhook successfully delivers routine findings. If only urgent findings reach a channel and routine delivery fails, the slot remains available for a later dispatch. Email-disabled findings do not reserve a slot unless a generic webhook will carry them. The phpanel webhook stream bypasses this budget.
Suppressions
Create suppression rules to silence known false positives:
- From the Findings page: click the suppress button on any finding. The path pattern is a glob; the dialog pre-fills the finding’s own file with
[,],*,?and\escaped, so the rule matches that file only - From the Findings page with several findings selected: bulk Suppress creates one rule per selected file
- From the Rules page: manage suppression rules directly
- Via API:
POST /api/v1/suppressions
A suppression rule hides matching findings from the Findings page, stops their email and webhook alerts, and stops file, process and account remediation for them. It does not stop IP blocking, challenge routing or attack scoring: a rule that mutes a whole check would otherwise leave every attacker that check reports unblocked. To exempt an address that was blocked by mistake, whitelist it on the Firewall page under Allow Rules or with csm firewall allow.
Incident auto-blocking, credential-spray containment and central threat intelligence also use suppressed findings. Database response may still block suspicious session IPs when enabled, but a suppressed database finding cannot trigger cleanup or session revocation. IP action notifications are separate findings; suppressing their check type mutes those notifications without stopping the action. These rules apply to startup, scheduled and real-time scans, control-socket runs with alerts enabled, and replay after restart.
To suppress email alerts for specific checks while keeping them visible in the web UI, use disabled_checks in your config:
alerts:
email:
disabled_checks:
- "email_spam_outbreak"
- "perf_memory"
Email AV
CSM scans email attachments in real-time using ClamAV and YARA-X on the Exim mail spool.
How It Works
- fanotify watches the Exim spool directory for new messages, including every
split_spool_directoryhash subdirectory (the cPanel default layout); hash directories Exim creates later are picked up within a minute - Attachments are extracted and scanned by ClamAV (socket) and YARA-X (if available) Base64 and quoted-printable decoding tolerates stray characters, missing padding and malformed escapes. Encoded multipart bodies are decoded at every level. Ambiguous base64 is scanned under quartet and alphabet-only interpretations, and a valid padded prefix is scanned separately when data follows it. At most 16 additional interpretations are scanned per message. If decoding or multipart parsing reports an error, recovered content and later attachments are still scanned and the message is reported as incompletely scanned. Extraction size and MIME nesting limits still apply
- Zip and tar.gz attachments are unpacked within configured size and file limits
- Extracted parts are staged under
state_path/emailav-tmp, which must stay daemon-owned and private - Attachment names written to logs and the UI use sanitized base names
- Infected messages are quarantined with full metadata
- Sender, recipient, and message ID are logged
Quarantine and release preserve each queue file’s ownership, permissions, and modification time, including when the move crosses filesystems.
Quarantine persists recovery metadata before moving spool files and syncs file contents and directory changes. A failed rollback retains the remaining quarantine files. Release restores the body before publishing the queue header; storage failures retain metadata and report the affected paths. Inspect both the spool and quarantine locations before retrying a partially completed operation.
Web UI
The Email page (/email) shows:
- AV watcher status (active, engine health)
- Scan statistics (scanned, infected, quarantined)
- Quarantined email list with release/delete actions
API Endpoints
GET /api/v1/email/stats Scan statistics
GET /api/v1/email/quarantine Quarantined email list
GET /api/v1/email/av/status AV watcher status
POST /api/v1/email/quarantine/ Release or delete quarantined email
Related Checks
email_content- scans outbound email body for credentials and suspicious URLs; every base64 body and part, attachments included, is decoded with the same tolerance as attachment scanning, using live MIME headers and the declared multipart boundariesemail_weak_password- detects email accounts with weak passwordsemail_forwarder_audit- audits forwarders for exfiltration redirectsmail_queue- alerts on queue buildup (spam outbreak indicator)mail_per_account- per-account sending volume spikes
Email password audit
The deep scan checks mailbox passwords against account-derived candidates and the bundled weak-password list. Verification runs inside CSM; passwords and stored hashes are never passed to subprocesses or included in findings. Confirmed matches can still use the HIBP range API for breach counts, sending only a SHA1 prefix.
Supported stored formats follow Dovecot password schemes:
| Scheme | Accepted format and audit limit |
|---|---|
| CRYPT, SHA512-CRYPT, SHA256-CRYPT | SHA crypt with a nonempty salt of up to 16 characters; 1,000 to 1,000,000 rounds, or the standard 5,000 when omitted |
| CRYPT, MD5-CRYPT | MD5 crypt with a nonempty salt of up to 8 characters |
| CRYPT, BLF-CRYPT | bcrypt 2a, 2b, or 2y; cost 4 through 14 |
| ARGON2I, ARGON2ID | Version 19; at most 64 MiB memory, 4 passes, and 4 lanes; salt 8 to 64 bytes, digest 16 to 64 bytes |
| PLAIN | Plaintext |
| PLAIN-MD5, LDAP-MD5, SMD5 | MD5 digests; salted variants allow 1 to 64 salt bytes |
| SHA, SHA1, SSHA, SHA256, SSHA256, SHA512, SSHA512 | SHA digests; salted variants allow 1 to 64 salt bytes |
Unprefixed hashes use CRYPT. Scheme names and .hex, .b64, and .base64
encoding suffixes are case-insensitive. Unsalted digests also accept Dovecot’s
hex/base64 autodetection. DES crypt, bcrypt 2x, yescrypt, PBKDF2, and
mechanism-specific formats such as SCRAM are not audited.
Stored values are limited to 4 KiB and candidates to 256 bytes without embedded NULs. At most three hash calculations run concurrently. Cancellation returns promptly; a calculation already running keeps its worker slot until its bounded work finishes.
An unsupported, malformed, or over-budget hash produces
email_password_audit_incomplete. CSM keeps earlier password findings and does
not record that mailbox as successfully audited. Other mailboxes continue to be
checked. When these are the only failures, the audit waits for
password_check_interval_min before trying again; temporary verification failures
remain eligible for retry on the next deep scan. Cycles skipped by the interval
preserve both earlier password findings and the audit warning. A completed audit
replaces them with its current results. Upgrading
starts a fresh audit even when an older CSM version recorded the same hash.
Threat Intelligence
CSM tracks, scores, and correlates attacks using a local attack database enriched with external feeds and GeoIP data.
Attack Database
- Per-IP event tracking (brute force, webshell upload, phishing, C2, WAF block)
- Local scoring from attack volume, types and targeted accounts
- Auto-block on reputation threshold
- Top attackers leaderboard
Successful cPanel, FTP, webmail and PAM login audit events, and authenticated
File Manager writes, remain in event history and account counts as
auth_success (Authenticated Activity). They add no event-volume or
multi-account score. Multi-IP login and authentication-failure signals keep
their existing scoring.
file_upload remains readable for historical records; no current check
produces it. Existing scores recorded under older classifications are not
rewritten because aggregated login counts cannot distinguish ordinary logins
from multi-IP alerts. For a confirmed false positive, follow
Clearing a stale local threat score.
That operation preserves historical events and does not remove firewall or
permanent-blocklist entries; review those separately.
IP Intelligence
Combines multiple sources into a unified verdict:
| Source | Data |
|---|---|
| Local attack DB | Event count, types, score |
| AbuseIPDB | External reputation (if API key configured) |
| Rspamd | Per-IP rolling history (if controller access configured) |
| Upstream HTTP cache | Panel-side shared score (if reputation.upstream configured) |
| Permanent blocklist | Operator-managed persistent blocks |
| Firewall state | Currently blocked/allowed status |
| GeoIP | Country, city, ASN, ISP |
| RDAP | Network name, organization (cached 24h) |
Verdicts: clean, suspicious, malicious, blocked
Pluggable sources
Threat-intel sources implement a small Source interface (lookup-by-IP returning a score + reason). The aggregator queries every enabled source in parallel, applies per-source weighting, and produces the unified verdict above. Adding a new source means implementing the interface and registering it; no existing source code changes.
Currently shipped:
- AbuseIPDB (
reputation.abuseipdb_key) - external IP reputation feed. CSM caps uncached lookups per cycle and reserves store-backed daily quota before sending requests. While the quota is exhausted (daily budget or an API 429/402 backoff) CSM emits areputation_quota_exhaustedWarning so the degraded coverage is visible; athreat_feed_staleWarning fires when previously downloaded free threat feeds have not refreshed in over 7 days. Persistent conditions remain visible in Findings but send at most one reminder per day, including across daemon restarts; these coverage warnings never trigger reputation scoring or blocks. - Rspamd (
reputation.rspamd.*) - per-IP rolling-history signals from the local rspamd controller. Delivered ham dilutes the score, temporary deferrals are neutral, and definitive spam actions count against the sender. Token resolution readstoken_envfrom the process environment at query time. Changing the external environment requires a daemon restart; see credential rotation. - Upstream HTTP cache (
reputation.upstream.*) - shared panel-side cache of AbuseIPDB or proprietary scores. Useful in fleets: agents pay a bounded local cache hit (cache_ttl_min, default 15 m) instead of hammering the upstream once per agent. CSM temporarily opens a fail-open circuit breaker after repeated upstream failures and lets only one cooldown probe through at a time. Use HTTPS for remote panels; plain HTTP is accepted only for loopback. Wire contract:docs/upstream-threat-intel-contract.md.
Verified crawlers
reputation.bot_verify_enabled verifies claimed crawler User-Agents
with static IP ranges first, then strict forward-confirmed reverse DNS.
reputation.verified_bots adds operator-defined crawler identities with
name, ua_substrings, and one verification method: rdns_suffixes or
ip_ranges. With rdns_suffixes the source IP must forward-confirm under
a registrable domain (public suffixes and shared-hosting suffixes are
rejected; a PTR-only match is not trusted). With ip_ranges the source IP
must fall in one of the published CIDRs – this is for crawlers such as
GPTBot and PerplexityBot that publish address ranges instead of crawler
reverse DNS. Over-broad or non-public ranges are rejected. All
checks run at config load and on reload.
Built-in rDNS verification covers Googlebot, Bingbot, Applebot, DuckDuckBot, Amazonbot, the Facebook and Meta crawlers, Brave, and the SERanking backlink bot. Googlebot, Bingbot, and Applebot also match a shipped IP-range snapshot first and fall back to reverse DNS; DuckDuckBot, Amazonbot, Facebook/Meta, Brave, and SERanking are rDNS-only.
Reverse-DNS verification is asynchronous. A newly admitted verification job receives a short, bounded pending window, including time spent waiting in the queue. High-volume traffic in that window can route to the proof-of-work challenge when it is enabled. A full queue, stopped worker, unsupported identity, or expired pending window uses the ordinary flood and scanner controls instead.
DNS failures, missing reverse DNS, and failed cache writes do not prove a
spoofed identity. They leave verification unresolved, delay retries, and do
not renew pending treatment on each retry. Attempt history is bounded. When
it is full, new sources can replace completed entries after their initial
cooldown; retries cannot extend that reservation. Live jobs and newly granted
pending windows keep their history. If no entry can be replaced, admission is
refused until capacity becomes available instead of running untracked lookups.
Evicted or expired sources without a persisted missing-PTR record can receive
another pending window, but the initial cooldown prevents continuous renewal
under churn. Expiry preserves live jobs and their retry delay after completion.
A missing PTR suppresses further DNS lookups for one fixed hour in the state
database, including across restarts. Scans do not extend that hour. After it
lapses, DNS retries resume, but the record remains as attempt history for a
day, so a restart or in-memory eviction cannot grant fresh pending treatment.
Records older than that are removed as new ones are written. Cleanup is
attempted at most once an hour, including after a failed cleanup. A definitive
verification result replaces that history. The record is not a verdict: it
grants neither the verified-crawler exemption nor pending treatment, and it
never counts as a spoof. csm store reset-bot-verify and a verified_bots
change clear these records together with the cached results.
A confirmed negative remains eligible for spoof detection. A cached positive
receives the normal verified-crawler exemption.
GPTBot, ChatGPT-User, OAI-SearchBot, PerplexityBot and ClaudeBot are recognized
out of the box: their published IP ranges ship as an embedded snapshot and are
refreshed from the vendor endpoints by an auto-updater (reputation.bot_ranges,
default on, outbound HTTPS, configurable interval; restart required for setting
changes). Fetched ranges are validated with the same over-broad and non-public
guards as operator entries, and the embedded snapshot is the trusted fallback
when a fetch fails. Anthropic publishes one combined feed for ClaudeBot,
Claude-User and Claude-SearchBot and documents IP-list verification rather than
reverse DNS, so CSM verifies ClaudeBot by address from that feed; the legacy
anthropic.com reverse-DNS suffix is kept only as a fallback.
csm update-bot-ranges refreshes these ranges on demand (mirroring
csm update-geoip): it fetches the vendor feeds, writes the on-disk snapshot,
and asks the running daemon to apply them without a restart. The auto-updater
and the manual command both export metrics – refresh success/failure, prefix
count per crawler, and the last successful refresh time – under the
csm_botranges_* names.
Abuse Reporting
reputation.report can send minimized confirmed-abuse reports to a central
database or private collector. It is off by default. Remote targets must use
HTTPS; plain HTTP is accepted only for loopback collectors. Keys and target
wiring are read at daemon startup, so changes to this block require a restart.
Web UI
The Threat Intel page (/threat) provides:
- IP lookup with composite scoring
- Two separate block actions per IP: a 24 hour block and a confirmed permanent block, singly or over a selection of attackers
- Top attackers with GeoIP enrichment
- Attack type breakdown chart
- Hourly trend chart
- Whitelist management (permanent and temporary)
A 24 hour block records threat evidence that expires with the firewall block, so a mistaken block of a customer address stops counting against it once the block lapses. A permanent block records evidence that stays until an operator clears it. When an address is no longer blocked but still holds permanent evidence, the lookup says so, because that address is flagged again the next time it is seen.
The 24 hour action refuses to shorten an existing permanent or longer firewall block. Unblock explicitly before changing that lifetime. Bulk requests skip those addresses and report warnings. Existing permanent threat evidence remains until cleared; feed updates do not remove a timed operator record before its expiry.
Bulk undo restores each prior firewall lifetime, using the original deadline for timed blocks. Expired blocks stay expired, and a later block, clear, or whitelist decision invalidates the older undo action. Only successful firewall reversals restore the corresponding threat evidence.
API Endpoints
GET /api/v1/threat/stats Attack stats and type breakdown
GET /api/v1/threat/top-attackers Top attacking IPs with GeoIP
GET /api/v1/threat/ip IP threat lookup
GET /api/v1/threat/events IP event history
GET /api/v1/threat/whitelist Whitelisted IPs
GET /api/v1/threat/db-stats Attack database statistics
POST /api/v1/threat/block-ip Block IP for 24 hours
POST /api/v1/threat/block-ip-permanent Block IP with no expiry
POST /api/v1/threat/whitelist-ip Permanent whitelist
POST /api/v1/threat/temp-whitelist-ip Temporary whitelist
POST /api/v1/threat/clear-ip Clear from attack DB
POST /api/v1/threat/unwhitelist-ip Remove from whitelist
GeoIP
MaxMind GeoLite2 integration for IP geolocation and ASN enrichment.
Features
- City database - country, city, latitude/longitude
- ASN database - ISP, organization, autonomous system number
- Auto-download on first use
- Auto-update every 24 hours (configurable)
- RDAP fallback for detailed ISP/org info (cached 24h)
Where It’s Used
- Threat intel page (top attackers, IP lookup)
- Firewall audit log (country flags)
- Login alerts (geographic context)
- Country-based login suppression (
trusted_countries) - Country blocking (firewall CIDR ranges)
Configuration
geoip:
account_id: "YOUR_MAXMIND_ACCOUNT_ID"
license_key: "YOUR_MAXMIND_LICENSE_KEY"
editions:
- GeoLite2-City
- GeoLite2-ASN
auto_update: true
update_interval: 24h
Free account: maxmind.com/en/geolite2/signup
CLI
csm update-geoip # Manual database update
csm firewall update-geoip # Download country CIDR blocks
csm firewall lookup <ip> # GeoIP + block status lookup
Address lookups report country, city, ASN, organization, and network when the local databases supply them. Country-block matches retain all matching countries and do not prevent ASN or city enrichment. Missing fields are omitted.
API
GET /api/v1/geoip IP geolocation (?ip=&detail=1)
POST /api/v1/geoip/batch Batch lookup (array of IPs)
Challenge Pages
JavaScript proof-of-work challenge pages - a CAPTCHA alternative for suspicious IPs.
How It Works
- Suspicious IP hits a protected resource
- CSM serves a challenge page requiring client-side SHA-256 proof-of-work
- Browser computes the proof (shows progress bar)
- On valid solution, CSM issues an HMAC-verified token
- Subsequent requests pass through automatically
Features
- SHA-256 based difficulty - configurable 0-5 levels
- Client-side computation - no server load
- HMAC token verification - prevents replay attacks
- Nonce-based anti-replay
- User-friendly - progress bar, instant feedback
- Bot filtering - headless browsers and scripts fail the challenge
Use Cases
- Gray-listing alternative to hard IP blocks
- Protecting WordPress login pages
- Rate limiting without blocking legitimate users
- DDoS mitigation layer
Routing Behavior
When challenge.enabled: true, CSM routes eligible IPs to the challenge page instead of hard-blocking them. This works independently of auto_response settings.
Challenge-Eligible Checks
Pre-auth, browser-visible attack signals only: wp_login_bruteforce,
xmlrpc_abuse, wp_user_enumeration, webmail_bruteforce,
http_scanner_profile, http_claimed_bot_unverified, ip_reputation,
local_threat_score. Post-auth audit events (cPanel, webmail, file upload, WHM
logins), WAF high-volume attacker findings, and non-browser protocols (SSH,
FTP, DNS recursion, outbound traffic, API auth) are excluded - their IPs have
no useful challenge step or no browser session to render the PoW page.
http_scanner_profile routing is operator-selectable:
auto_response.http_scanner_action: "challenge" (default) routes the IP here;
"block" skips the challenge and hard-blocks directly. With the challenge
subsystem disabled, both values block.
Always Hard-Blocked
Confirmed malware (webshells, YARA/signature matches), WAF high-volume attackers, C2 connections, backdoor ports, phishing pages, database injections, and spam outbreaks are hard-block candidates immediately, even when challenge is enabled.
Timeout Escalation
If an IP doesn’t solve the PoW challenge within 30 minutes, it is
automatically escalated to a hard firewall block. The pending claimed-bot path
(http_claimed_bot_unverified) is the exception: it expires without timeout
escalation because the reverse-DNS verifier decides the next action. A confirmed
spoof is hard-blocked by the later http_ua_spoof finding after it reaches the
UA threshold, while a real crawler is removed from the challenge list once
verification succeeds.
Bind address
The listener binds to 127.0.0.1 by default, so enabling the challenge
server alone does not expose a new public port. The webserver integration
uses direct redirects to challenge.public_url; installed direct mode
therefore needs a non-loopback listener and a public URL ending in
/challenge.
challenge:
enabled: true
listen_addr: 0.0.0.0
listen_port: 8439
public_url: https://cpanel.example.com:8439/challenge
tls_cert: /var/cpanel/ssl/cpanel/mycpanel.pem
tls_key: /var/cpanel/ssl/cpanel/mycpanel.pem
When CSM’s firewall is enabled and challenge.port_gate.enabled is true,
the daemon also opens challenge.listen_port in the main firewall rules.
The port-gate chain still drops traffic to that port unless the source is
loopback, an infra_ips entry, or an IP currently on the challenge list.
Port-gate rules follow the configured listener address family. An IPv6-only
listener gates only IPv6 clients; IPv4 challenge entries stay in the
webserver map but are ignored by the IPv6 nftables set.
Run csm doctor challenge after changing these fields. The command checks
the public URL shape, TLS files, port-gate setting, installed webserver
snippet version, webserver configtest, and the live /challenge/gate
endpoint. Add --json for automation.
TLS
The challenge listener serves HTTPS when challenge-specific TLS material is configured. Loopback listeners stay on plain HTTP by default. Direct/public listeners can reuse the Web UI cert.
Resolution order:
challenge.tls_cert+challenge.tls_key(explicit per-service).webui.tls_cert+webui.tls_key(shared cert; cPanelmycpanel.pemcovers both webui and the challenge port without extra config) only whenchallenge.listen_addris not loopback.- Plain HTTP. This is expected for the default loopback-only path.
Public listeners without TLS log a startup warning.
HSTS-pinned parent domains (cPanel, phpanel, customer apex) will
fail with
ERR_SSL_PROTOCOL_ERRORbecause the browser auto- upgrades the URL to https; ship TLS material in production.
challenge:
tls_cert: /var/cpanel/ssl/cpanel/mycpanel.pem
tls_key: /var/cpanel/ssl/cpanel/mycpanel.pem
Trusted Proxies
By default, the challenge server uses RemoteAddr to identify clients.
The shipped webserver integration redirects browsers directly to
challenge.public_url, so it does not need trusted_proxies. Configure
trusted proxies only for a custom proxy deployment where CSM receives
traffic from a proxy and must trust X-Forwarded-For from that proxy.
challenge:
enabled: true
trusted_proxies:
- "127.0.0.1"
- "::1"
Without trusted_proxies, X-Forwarded-For is ignored to prevent IP spoofing.
Successful Verification
When a client passes the challenge:
- The IP is removed from the challenge list so the webserver stops sending that visitor to the challenge flow
- A verification cookie is set, so a visitor who is challenged again within its lifetime passes without seeing the puzzle
Passing the challenge changes nothing in the firewall. The visitor is treated like any other client from then on: normal rate limits and blocks still apply, and a later attack signal can list the IP again. The challenge page, the verify endpoint and the CAPTCHA fallback answer only for IPs currently on the challenge list; anyone else gets 404.
Webserver Integration
The webserver integration redirects challenge-listed IPs to
challenge.public_url. The installer refuses to run until that URL is
an absolute http or https URL ending in /challenge, and the
configured challenge listener is non-loopback.
csm webserver-integration install # initial wire-up
csm webserver-integration upgrade # re-apply after a CSM upgrade
csm webserver-integration status # show detected stack + version drift
csm webserver-integration validate # run the webserver's configtest
csm webserver-integration remove # uninstall the snippet
The installer auto-detects the active webserver via
internal/platform. Supported stacks and snippet paths:
| Stack | Snippet path |
|---|---|
| cPanel + Apache (EasyApache) | /etc/apache2/conf.d/csm-challenge.conf |
| Debian/Ubuntu Apache | /etc/apache2/conf-enabled/csm-challenge.conf |
| RHEL family Apache (httpd) | /etc/httpd/conf.d/csm-challenge.conf |
| LiteSpeed (LSWS) | /usr/local/lsws/conf/templates/csm-challenge.conf |
| Nginx (plain + Engintron + phpanel) | /etc/nginx/conf.d/csm-challenge.conf |
The snippets are rendered from the effective CSM config. Apache and LSWS
read their RewriteMap from /var/cache/csm/challenge_ips.txt; Nginx reads
a native map include from /var/cache/csm/challenge_ips.nginx.map. The
directory is the service’s systemd cache directory: it is world-readable so
the webserver user can read the maps, and unlike the runtime directory it
survives systemctl stop csm, package upgrades and reboots. That matters
because Apache and LSWS validate the map at config-parse time and Nginx
fails on a missing include, so a map that vanished with the daemon would
take every site down at the next webserver reload. CSM rewrites the Nginx
include on challenge-list changes and reloads Nginx only when the file
content changes.
At startup the daemon creates the default maps if they are absent,
including when challenge mode is disabled. It also checks the legacy
installer-deployed snippet (/etc/apache2/conf.d/csm_challenge.conf). If
its RewriteMap points at a map file the daemon does not maintain – which
makes webserver config validation fail host-wide once that file goes
missing – the daemon re-deploys the shipped template. A CSM-managed
integration snippet from an older template version is refreshed the same
way, through the configtest-then-reload flow below. Files without the
directive, operator-edited snippets, and snippets already pointing at the
daemon maps are never touched. While such a snippet still references a map
under the old runtime directory, the daemon keeps that file present until
the snippet is refreshed with csm webserver-integration upgrade.
csm uninstall and the package post-removal hook remove the snippets
before the maps, and keep the maps whenever a snippet cannot be removed
(for example an Nginx server {} block that still uses $csm_challenged).
On every run, the installer:
- Writes the new snippet to a sibling temp file and renames it into place atomically.
- Runs the webserver’s own configtest (
apachectl configtest,nginx -t,lswsctrl conftest). - On pass: reloads the webserver gracefully and exits 0.
- On fail: restores the previous snippet bytes (or removes the file if it did not exist before) and exits non-zero with the captured configtest output. The webserver is never reloaded with a broken config.
The snippet header carries a version marker; upgrade is a no-op when
the on-disk version matches the shipped version. Hand-edited files
(missing or mismatched marker) trip a “modified” status and the
installer refuses to overwrite them - remove or rename first.
Hosts with no detectable webserver exit with status=skipped so
package post-install hooks succeed cleanly on, e.g., a plain phpanel
worker that doesn’t run nginx locally.
Bypass Paths
Three opt-in bypass mechanisms let legitimate traffic skip the PoW page entirely. All default off; an upgraded csm.yaml with no new blocks behaves exactly as before.
CAPTCHA Fallback (JS-Disabled Visitors)
The PoW solver requires JavaScript. Visitors with JS off (older mobile browsers, accessibility tooling, text browsers, scripted integrations) would otherwise be locked out. When configured, CSM renders a Cloudflare Turnstile or hCaptcha widget inside a <noscript> block; on completion the form posts to /challenge/captcha-verify and CSM validates the token server-side against the provider’s siteverify endpoint.
Provider rejections do not spend the page nonce, so a visitor can retry the
same challenge page after a mistyped, expired, or failed widget response.
challenge:
captcha_fallback:
provider: turnstile # turnstile | hcaptcha | "" (off)
site_key: "0xAAAA..." # public key embedded in the widget
secret_key: "0xBBBB..." # verified server-side; never sent to client
timeout: 10s
Verified Operator Sessions
Operators who repeatedly hit the challenge during normal admin work can mint a signed cookie that bypasses PoW for the cookie’s TTL. The signing key is generated at daemon startup and rotates on every restart – old cookies stop working automatically.
challenge:
verified_session:
enabled: true
cookie_name: csm_admin_session # default
ttl: 4h # default
admin_secret: "long-shared-secret" # required
To issue a cookie, POST the secret to the challenge server:
curl -i -X POST -d "secret=long-shared-secret" \
https://your-host:8439/challenge/admin-token
# 204 No Content
# Set-Cookie: csm_admin_session=...; Path=/; HttpOnly; Secure; SameSite=Lax
The cookie binds to the requester’s IP, so a stolen cookie does not work from a different network.
Verified Search Crawlers
Googlebot and Bingbot can be allow-passed by reverse-DNS forward-confirm. CSM looks up the visitor’s PTR, checks it ends in a known crawler suffix (e.g. .googlebot.com), then forward-resolves that name to confirm the original IP appears in the result. A spoofed User-Agent: Googlebot from a residential IP fails forward-confirm and falls through to PoW.
challenge:
verified_crawlers:
enabled: true
providers: [googlebot, bingbot]
cache_ttl: 15m
Positive results cache for cache_ttl; negative results cache for one-fifth that long so a transiently-broken resolver does not lock out a real crawler for the full window.
Operational
Backups
csm store export and csm store import capture the bbolt store, state JSON files (baseline file hashes), and signature-rules cache into a single tar+zstd archive. Use these for re-provisioning, cluster cloning, and disaster recovery rather than re-baselining a 200k-file account tree from scratch. The daemon writes the archive under its state directory (the only tree its systemd sandbox can write) and the CLI moves it to the path you give, so any destination the CLI can write works, including /var/backups. The destination directory and every directory above it must be owned by root or by the account running the command, and must not be writable by other accounts unless they carry the sticky bit as /tmp does; the destination itself must not already exist as a symlink or as another account’s file. Otherwise another account could swap the archive out after the CLI has verified it, and the export is refused before the daemon writes anything. In-progress staging files and unconfirmed firewall rollback artifacts are excluded from backups, so restoring an archive cannot re-arm an old confirmation window.
csm store export /var/backups/csm-$(date +%F).csmbak
sha256sum -c /var/backups/csm-$(date +%F).csmbak.sha256
# transfer the .csmbak + .sha256 to the target host
systemctl stop csm
csm store import /var/backups/csm-2026-04-27.csmbak
systemctl start csm
Partial restore: --only=baseline restores only the file-hash state (useful after a full re-install where firewall and history should stay fresh); --only=firewall merges the firewall buckets into an existing daemon (useful for cloning blocklists across a cluster).
Performance Monitor
CSM monitors server performance metrics and generates findings when thresholds are exceeded.
Critical Checks (every 10 min)
| Check | What it monitors |
|---|---|
perf_load | CPU load average vs core count (critical/high/warning thresholds) |
perf_php_processes | PHP process count and total memory usage |
perf_memory | Swap usage percentage and OOM killer activity |
Host-wide OOM events are Critical; memory-cgroup limit events are Warning. Both are reported when present in the last hour, with separate deduplication identities for each scope and victim process.
Deep Checks (default every 60 min, thresholds.deep_scan_interval_min)
| Check | What it monitors |
|---|---|
perf_php_handler | PHP handler type (DSO vs CGI vs FPM) and configuration |
perf_mysql_config | MySQL my.cnf settings (buffer pool, connections, query cache) |
perf_redis_config | Redis memory limits, persistence, eviction policy |
perf_error_logs | Error log file sizes (bloat detection) |
perf_wp_config | WordPress wp-config.php hardening and debug settings |
perf_wp_transients | WordPress database transient bloat |
perf_wp_cron | WordPress cron scheduling (missed crons, excessive events) |
Web UI
The Performance page (/performance) shows real-time metrics:
- Server load and CPU usage
- PHP process and memory charts
- MySQL and Redis health
- WordPress performance indicators
PHP worker counts include LiteSpeed and PHP-FPM pool workers. The PHP-FPM master is excluded from these request-worker counts; security process checks continue to inspect it. Database memory comes from the MySQL or MariaDB server process, independent of its PID-file location or wrapper-related arguments.
The findings list also exposes admin-only fixes, per-row and as a Bulk fix dropdown that applies one fix to every matching finding at once:
-
perf_error_logs: truncate a bloatederror_login place. The inode is preserved so running PHP processes keep writing to the same file. -
perf_wp_config: disabledisplay_errorsin.user.ini,php.ini, or.htaccessby commenting the matched line and appending an Off override. -
perf_wp_cron: adddefine('DISABLE_WP_CRON', true)towp-config.phpand install a per-user system cron that runswp-cron.phpon a fixed interval. Disabling WP-Cron alone would stop scheduled WordPress tasks, so the cron is installed in the account owner’s own crontab (visible and editable by the customer). The cron is installed before the define is written, so a crontab failure leaves WordPress scheduling unchanged. The define is inserted before the “stop editing” marker (or thewp-settings.phprequire); insertion points inside multiline comments or heredocs are ignored, and the fix refuses awp-config.phpwith no safe insertion point rather than corrupt it.The installed schedule is staggered per account and docroot (for example
7-59/15) instead of a wall-clock-aligned*/15, so many managed sites do not all fire in the same second and spike the host load. Non-divisor intervals use a shifted minute list so the gap stays within the configured interval. The command also runs underflock -nwith a per-docroot lock file in the account home, so a slow pass skips the next run instead of overlapping it. On daemon start, managed crontab lines installed by older releases are upgraded to this format automatically (only lines under the# CSM WP-Cronmarker are touched; customer-authored cron entries are never rewritten).
These actions are limited to configured account roots, reject symlinks and unsupported file types, and remove the fixed row from the active findings view after a successful edit.
WP-Cron fix settings
Tune the WP-Cron remediation under Settings -> Performance:
performance.wp_cron_fix.interval_minutes(default15, range 1-60): how often the installed system cron runswp-cron.php. The interval only bounds task latency – WordPress keeps its own event schedule and a cron pass with nothing due is a wasted full bootstrap – so 15 minutes is right for most sites. Lower it per host only when busy stores need tighter Action Scheduler latency.performance.wp_cron_fix.php_bin(default empty): overrides the PHP interpreter for the cron line. Leave it empty and each site runs under the PHP version its own vhost is set to when cPanel’s domain map provides an unambiguous version. CSM falls back to the detected CLI interpreter when the map has no usable match; an existing managed job keeps its interpreter while the map is unavailable. Setting a value pins that one interpreter for every managed site. CLI php is used instead of an HTTP request so the job never ties up a web worker.
To let the daemon apply this fix automatically on every WP-Cron finding, set
auto_response.fix_wp_cron: true (default false; requires
auto_response.enabled: true). It is opt-in because it edits customer
wp-config.php files and crontabs.
MySQL telemetry auth
The MySQL panel runs mysql -e "SHOW STATUS LIKE 'Threads_connected'" from
the csm process. The client needs to authenticate against the local server,
and csm supports two setups out of the box:
-
A
~/.my.cnffor the csm runtime user with credentials for a MySQL account that holds at least thePROCESSprivilege. cPanel and CloudLinux ship/root/.my.cnffor the root user; csm running as root picks it up automatically. -
A unix-socket grant for the csm OS user, e.g. on Debian/Ubuntu MariaDB:
CREATE USER 'root'@'localhost' IDENTIFIED VIA unix_socket; GRANT PROCESS ON *.* TO 'root'@'localhost';
If neither is configured, the MYSQL card renders n/a / n/a instead of a
misleading 0 conn. csm makes no attempt to connect over TCP or store
credentials on its own.
Redis telemetry auth
The Redis panel connects to local Redis at 127.0.0.1:6379. If Redis
requires a password, set REDISCLI_AUTH in the csm daemon environment.
The dashboard uses that password for its in-process Redis client.
API
GET /api/v1/performance Current performance metrics snapshot
POST /api/v1/perf/fix-error-log
POST /api/v1/perf/fix-display-errors
POST /api/v1/perf/fix-wp-cron
Web UI
HTTPS dashboard that updates live from the finding event stream (see Live updates). Dark/light theme toggle.
Static assets use content-versioned URLs. Only a URL matching the served file receives immutable caching; older or unversioned URLs must revalidate. Replacing a file changes its version even when its size and modification time are preserved. Text assets use gzip when accepted by the client, except for byte-range requests.
Navigation
The sidebar groups pages by operator workflow. URLs are stable; the groups only reorder visibility:
- Overview - Dashboard
- Triage - Incidents, Findings (Active and History tabs)
- Response - Firewall, Quarantine, Cleanup History, Email Security, ModSecurity, Threat Intelligence
- Operations - Performance, Server Hardening, Audit Log
- Configuration - Rules, ModSecurity Rules, Verified Bots, Settings
A page has the same name in the sidebar, the browser tab and its heading.
The old /blocked address redirects to /firewall.
Sidebar group expand/collapse state is saved in the browser. On
viewports under 992px the sidebar collapses into a top-bar drawer
toggled from the hamburger button. Account detail (/account?name=<account>) is
not in the sidebar; the finding detail, account groups on Findings, incident
detail, the accounts targeted in a Threat Intel lookup and the dashboard’s
accounts-at-risk card link to it, and the command palette opens an account
typed by name. The palette also lists Sessions. Browser logins require administrator
scope. The header links to session management.
Pages
| Page | URL | Purpose |
|---|---|---|
| Dashboard | /dashboard | Triage queue, daemon status strip, Components matrix, system posture, 24h stats, accounts at risk, auto-response summary, brute-force summary, timeline charts. A queued finding opens its own detail, and each 24h severity count opens the History tab for the last 24 hours at that severity |
| Findings | /findings | Active findings with search, check/account filters, header grouping toggle, detail panel, fix/dismiss/suppress actions, a permanent Block for findings that report an attacker address, sticky bulk operations (fix, dismiss, suppress), modal account scan. The open finding is kept in the URL as ?key=<finding key>, so the link reopens it |
| Findings > History | /findings?tab=history | Paginated archive of all findings, newest first, with date range and severity filters, 25 to 200 rows per page (hperpage), CSV export; window=24h (1 to 720 hours) shows a rolling window instead of calendar days |
| Quarantine | /quarantine | Every file backup: quarantined files and cleaners’ pre-clean backups, with type, live state of the original path, content preview, restore and delete; filters by account, detector, type and date |
| Cleanup History | /cleanup-history | DB-object backups with preview and restore controls; file backups are on the Quarantine page |
| Firewall | /firewall | Subview-tabbed page (?view=overview/blocks/allow/config/audit/danger; ?ip=<address> opens the lookup for that address): blocked IPs/subnets with GeoIP, bulk unblock of selected rows (with undo), the whitelist and allow rules (Allow Rules tab), search, audit log; the lookup links to Threat Intel for the same address; destructive actions live under the Danger tab |
| ModSecurity | /modsec | WAF workbench: status strip, Active WAF pressure summary list (top attackers by hits), top rules / domains side panel, Blocked IPs / Events tabs, and a Manage Rules link to ModSecurity Rules. Block detail panels show first-seen, top URIs, sample events, and direct links to Threat Intel, Firewall lookup, and rule management |
| ModSecurity Rules | /modsec/rules | Enable or disable CSM rules (applied with one reload) and firewall escalation exclusions; the exclusion list works even when rule management is not configured |
| Email Security | /email | Mail queue and AV status, grouped account/auth/queue/malware findings, quarantine, senders, forwarders, provider deferrals, and PHP-relay abuse. Queue actions distinguish real mail from frozen null-sender backscatter; held external forward copies can be released or deleted without affecting the local delivery. |
| Verified Bots | /verified-bots | Editor for the verified-crawler allowlist (reputation.verified_bots): UA, reverse-DNS suffix, and IP-range identities, plus auto-update posture, with apply-and-reload. Admin scope |
| Threat Intelligence | /threat | IP lookup with scoring/GeoIP/ASN (?ip=<address> runs it on load), 24 hour and permanent block and whitelist actions (single and bulk), top attackers, attack type charts, trends; the lookup links to Firewall for the same address, and the whitelist itself is kept under Firewall > Allow Rules |
| Server Hardening | /hardening | On-demand hardening audit, stored report, score, and remediation guidance |
| Incidents | /incident | Correlated incident list and grouped view, both filterable by every status, with detail panel and bulk status changes (contained, resolved, dismissed) for the selected incidents on the page, plus forensic timeline search by IP or account |
| Rules | /rules | YAML/YARA rule management, suppressions, state export/import, test alerts |
| Account | /account | Per-account analysis: findings, quarantine and history. An on-demand account scan runs from the Findings page |
| Audit Log | /audit | Every operator action in the Web UI and API, including logins, logouts and session revocations, with the credential that acted, search, action and date filters, URL state, and export. Failed logins go to the daemon log instead |
| Performance | /performance | Server load, PHP processes, MySQL, Redis, WordPress metrics |
| Settings | /settings | Searchable config editor with grouped large sections, field-level validation errors, restart notices, redacted secret updates, and firewall tentative apply with rollback timer. Commands, file paths, sockets and environment variable names are shown read-only and change only in csm.yaml; changing the rspamd or upstream address requires entering its credential again |
| Sessions | /sessions | Active browser logins, individual revocation and logout of every session (confirmed first) |
Audit attribution is captured when the action is authorized and remains available if the browser session expires or is revoked while the action runs.
Account views use the recorded finding owner when available, with account paths as a fallback. They also recognize resolved paths under linked account roots, including files already moved into quarantine.
Dates and time zones
Every page shows dates in the time zone chosen under Preferences: the browser’s, the server’s, or a named zone. Date filters pick whole days in that zone. Changing the zone reloads the page so dates already on screen follow it, provided the browser can save and read back the preference. If browser storage is unavailable, new renders use the preference without forcing a reload. Days with a midnight clock change start at the first valid time of that day; a repeated midnight uses its first occurrence.
Live updates
The Web UI keeps an event stream open (/api/v1/events) and shows a Live
indicator while it is connected. Findings are batched for up to a couple
of seconds before pages update:
- Findings shows the “new findings” banner at once instead of on its next check
- Dashboard sends desktop notifications and refreshes the 24h counts and the triage queue
- Incidents reloads the current list, unless incidents are selected for a bulk change
While the stream is connected, the checks those pages run on a timer slow down to once a minute as a safety net; when it drops they return to their normal pace. Pausing auto-refresh also pauses live updates.
Desktop notifications handle findings that arrive out of order within a batch or share a timestamp, without repeating them on the next history poll.
Refresh
The header shows when the page’s data was last loaded (“Updated N ago”). It moves when the page loads its data, on each automatic refresh and on the Refresh button; an action or a detail lookup does not change it. The pause button appears only on pages that refresh on a timer, and pauses that refreshing in this browser.
Refresh reloads the page’s data in place, keeping filters, scroll and open panels. On a page with unsaved edits (Settings, Verified Bots, staged ModSec rule changes) it asks before discarding them. Hardening’s Refresh reloads the stored report and does not run a new audit. Pages rendered by the server, such as Sessions, reload.
A panel that fails to load says what failed and why, with a Retry button that loads only that panel again. The panel’s earlier content returns once a later load succeeds.
The visible History tab refreshes with its current filters and page size. Editors wait for an ongoing save or load before accepting another refresh; Verified Bots keeps edits made while a reload is pending. ModSecurity exclusion controls wait for the exclusion list to load, and a failed load stays visible instead of appearing as an empty list. Late responses cannot replace a newer account tab, settings section, history page, or finding detail, or reopen a finding detail that was closed.
Notifications
Success and information notices fade after five seconds. An error stays until you close it, and the same error is not shown twice while it is on screen.
Confirmations
A confirmation for an action that deletes data, blocks traffic, turns protection off or ends sessions names the action on a red button, such as Delete or Block, and starts with Cancel focused, so pressing Enter does not carry it out. Logging out every browser session asks first. Cancel keeps focus after the dialog finishes opening. A later text prompt uses its own normal OK button, and bulk actions keep one confirmation open at a time.
Bulk file actions
Select-all and every bulk action reach only the rows the table currently shows. Rows on other pages or hidden by a search or filter are never selected or acted on; set the page size to All to act on every row. Quarantine selection counts and buttons are refreshed whenever the visible rows change.
A failed file restore cleans up its own destination copy while retaining the quarantined evidence. A replacement created by another writer is preserved.
Quarantine deletes large file selections in sequential batches. If a request fails, later batches are not sent; the page reports the confirmed deletion count and refreshes the list. File restore and delete controls stay disabled until the operation and refresh finish, and Refresh waits for them.
Threat Intel bulk block and whitelist actions accept up to 100 selected IPs and retain one undo action. Larger selections must be narrowed before sending. Findings bulk fix also asks for a smaller selection when the request would exceed the API body-size limit, which includes finding details.
Findings bulk suppress creates one path rule per selected file, up to 100 files, with the file name matched literally. Selected findings that name no file are skipped; a rule for a whole check is made from one finding. Rules are saved one at a time, and a failure stops the rest and reports how many were saved.
Security
- Authentication - API bearer tokens in the header; opaque server-side browser sessions in HttpOnly/Secure/SameSite=Strict cookies
- CSRF - HMAC-derived token bound to the browser session on cookie-authenticated POST, PUT, PATCH, and DELETE requests; a form sends it in the body, never the query string
- Headers - X-Frame-Options DENY, Content-Security-Policy (scripts, styles and forms from the Web UI only; no plugins,
<base>or framing), HSTS, nosniff, and the legacy XSS auditor turned off - TLS - Auto-generated self-signed certificate, renewed automatically within 30 days of expiry and picked up without a restart; a certificate you install is never replaced, and replacing its files takes effect on the next connection. Renewal keeps the existing private key, so a failed certificate write leaves the working pair intact; explicitly configured certificate and key files must already exist
- Rate limiting - 5 login attempts/min, 600 API and
/metricsrequests/min per IPv4 address or IPv6 /64 - Token length - tokens shorter than 32 characters are reported as warnings at startup and by
csm validateandcsm doctor; they keep working - Bearer auth skips CSRF (for API-to-API calls)
Browser sessions
Log in with an administrator credential from webui.tokens (or the migrated
legacy webui.auth_token). The cookie contains a new random session secret;
the reusable API credential never appears in it. Old token-valued cookies are
rejected, so an upgrade requires a fresh login. Read-scope API tokens cannot
create browser sessions.
Use Sessions in the header (/sessions) to see login names, client address,
browser, creation time, last activity and absolute expiry. Revoke one session
or log out every browser, including your own. These operations do not rotate
API credentials. Logout uses a CSRF-protected POST. Logout and revocation
accept the same browser origins as API writes (webui.allowed_origins); the
login form does not check the origin.
webui:
session_lifetime: "24h"
session_idle_timeout: "30m"
Both durations require a restart. Lifetime must be between one second and 30 days; idle timeout must be at least one second and no longer than lifetime. Zero does not disable expiry. Idle time means time without operator activity: page loads and API requests made within a minute of keyboard, pointer or scroll input. Background polling by an open page does not count, including metrics scrapes, HTML page fetches and event-stream connections. Browser navigation to a page counts as a page load; fetching that page on a timer does not. The activity marker is recalculated when each API request is sent, so a dashboard left open still logs out after the idle timeout. Activity is committed at bounded intervals, so idle expiry can occur slightly early, never late. Passive event-stream heartbeats do not extend the session; streams check revocation and expiry before each event and heartbeat.
Every daemon restart invalidates all browser sessions. Token removal, rotation, name or scope changes take effect after the required restart; log in again with a current administrator credential. Reauthentication creates a new session and revokes the previous one. Operator preferences remain tied to the login credential, so a new session does not reset them.
The local transactional store keeps session verifiers, never raw cookie secrets.
Failed persistence cannot issue a login or claim successful revocation. If the
previous session cannot be read during reauthentication, login fails without
issuing a replacement cookie or changing that session. Concurrent requests
preserve the latest committed activity even when they arrive out of order.
If the session store is unavailable, browser authentication fails closed; API bearer
authentication remains independent. Session admission is bounded and refuses new
logins at capacity instead of evicting active sessions. The login response says
so; wait for idle sessions to expire, or revoke sessions from a logged-in
browser or with an admin API token (DELETE /api/v1/sessions). Expired sessions are removed during admission, and startup clears
the session records. Backup exports exclude live session records, and full
restores discard any session records from older archives.
MFA is a separate planned feature. See the session management guidance for the security principles behind opaque identifiers, expiry and revocation.
Keyboard Shortcuts
General
| Key | Action |
|---|---|
? | Show shortcut help |
/ | Focus search input |
Ctrl-K / Cmd-K | Open command palette |
Navigate
| Key | Action |
|---|---|
g d | Go to Dashboard |
g f | Go to Findings |
g h | Go to Findings > History tab |
g t | Go to Threat Intel |
g r | Go to Rules |
g b | Go to Blocked IPs (Firewall) |
Findings page
| Key | Action |
|---|---|
j / k | Move selection down/up (focus moves to the row) |
o / Enter | Open selected finding |
d | Dismiss selected finding |
f | Fix selected finding |
Finding and incident rows, finding group headers and sortable table headers also work from the keyboard: Tab to them and press Enter or Space. A sorted header reports its order to screen readers. Closing the detail panel or a dialog returns focus to what opened it, and a dialog opened from the detail panel keeps Tab and Escape to itself. If a refresh removes the opener, focus returns to the open detail panel or main content. Shortcuts act on the focused finding and stay inactive while a dialog or detail panel is open.
Each finding row offers up to four actions: Fix (apply the automated remediation, shown only when one exists), Re-check (re-evaluate the finding against the live filesystem and clear it if the condition is gone, useful after fixing something by hand instead of waiting for the next scan), Dismiss (stop alerts for it while it stays unchanged; a later scan that still finds it lists it again, and undo is offered for 30 seconds), and Suppress (create a rule to hide similar findings for good).
Dismissal undo preserves later dismissals, successful re-checks and baseline resets. Findings first received in real time stop alerting when dismissed, even before the next scheduled scan records them.
Re-check appears only when CSM can test a current condition again. Supported
targets include file permissions and content, phishing and .htaccess files,
selected accounts and system integrity checks, WordPress core/plugins, CMS
database rows, administrator accounts, and database objects. Re-check uses the
stored finding identity and current host state; the browser cannot substitute a
different path, row, account, or object.
The operation fails closed. Missing or unreadable evidence, package-manager errors, failed CMS discovery, and failed database queries leave the finding active. A file whose bytes changed since detection is never cleared either, because a partial clean and an evasion edit look alike. When the replacement can be proven inert, an empty file or a comment-only stub, the finding drops to Warning instead of clearing, and the original severity comes back if the file stops being inert. Historical events such as login attempts, WAF blocks, and IP reputation cannot be re-evaluated and therefore have no Re-check action. Broad aggregates and dependency findings require a new account or full scan.
WHM Plugin
CSM installs a WHM plugin (addon_csm.cgi) that redirects operators from WHM to the daemon Web UI. After the redirect, API calls are same-origin requests to the daemon.
API requests that carry a browser Origin header are accepted from https://<hostname>:<port> (the configured hostname and webui.listen port), from an https loopback origin such as https://localhost:9443 over an SSH tunnel on any local port, and from every origin listed in webui.allowed_origins (bare https://host[:port] entries, hot-reloadable). A loopback origin is accepted only when it is the origin the request was sent to (its Host): another local service in the same browser shares the Web UI’s cookies, which ignore the port, and is refused. Any other origin gets 403 Cross-origin request blocked, which shows up as a read-only UI: pages load but every action fails. The Host header is used only for this loopback comparison; it never admits a non-loopback origin.
API Reference
Machine-readable HTTPS API. All endpoints require token authentication. State-changing POST, PUT, PATCH, and DELETE requests require CSRF protection for browser cookie sessions.
Authentication
# Bearer token (header)
curl -H "Authorization: Bearer YOUR_TOKEN" https://server:9443/api/v1/status
# Cookie-based (after login)
curl -b "csm_auth=SESSION_COOKIE_FROM_LOGIN" https://server:9443/api/v1/status
Cookie-authenticated state-changing requests require the X-CSRF-Token header (obtained from the authenticated page meta tag). The token belongs to the browser session that loaded the page; another session’s token is refused, and a form field is read from the request body only. Admin-scope Bearer requests are CSRF-exempt because the Authorization header is the write credential.
Browser session management
A successful browser login exchanges an admin token for an opaque session cookie. API tokens are never valid cookies. Sessions expire on idle or absolute deadlines and all end on daemon restart. The header’s Sessions page provides the same management operations as these admin-only endpoints:
| Method | Path | Result |
|---|---|---|
| GET | /api/v1/sessions | items list of sessions with id, name, created, last_seen, expires, remote_ip, user_agent, current |
| DELETE | /api/v1/sessions/<id> | Revoke that session; an unknown ID answers 404 |
| DELETE | /api/v1/sessions | Revoke all browser sessions, including the caller’s |
Revocation returns {"ok":true} only after committing to the store. Failure
returns 503; invalid session IDs return 404. Lists contain no credential hashes
or session verifiers. Cookie-authenticated DELETE requests require CSRF; admin
bearer callers are exempt. Read-scope bearers cannot list or revoke sessions.
POST /logout requires authentication and CSRF for browser callers; it revokes
the current session and clears the cookie. GET /logout cannot log users out.
Token scopes
Configure tokens under webui.tokens: with a scope of admin or read:
webui:
tokens:
- name: "operator"
token: "..."
scope: admin # full read+write
- name: "panel-readonly"
token: "..."
scope: read # status, findings, history, stats, challenge stats, blocked IPs, scan jobs, health, components, capabilities, SSE
The legacy single-token webui.auth_token: is migrated automatically to a legacy-auth-token admin entry on first start. Read-scope tokens are intended for orchestrators and dashboards that consume status, findings, history, stats, challenge stats, blocked-IP summaries, scan jobs, health, components, capabilities, and SSE events. Admin scope is still required for write routes and for sensitive reads such as quarantine, settings, firewall internals, threat-intel detail, rules, account detail, exports, incident timelines, and audit history. (ModSecurity stats/blocks/events are read scope; only the ModSecurity rules and escalation routes need admin.) metrics_token: is a separate, read-only credential for /metrics only.
Errors
Every failure answers a non-2xx status with a JSON body that has an error
message. That includes CSRF, origin, rate-limit and wrong-method refusals.
Some failures add detail next to it; a settings save that fails validation
adds errors, a list of fields and messages. An unknown path under /api/
answers 404. No 2xx response reports a failure.
{"error": "invalid IP address"}
Actions
A request that changes state answers "ok": true with the action’s own
fields, such as undo_token or warning. Work that continues after the
answer, such as a daemon restart, a firewall rollback or a scan job, answers
202. A batch where some items failed answers 200 and lists the failures; a
batch where nothing changed answers an error status. A read-only route
answers 405 to any method but GET.
An empty fix batch is rejected. Threat actions report firewall failures, and bulk actions list invalid addresses as well as failed changes. A failed action can have applied some steps before the error; inspect the current state before retrying it. A response that cannot be encoded answers 500 with an error body.
/api/v1/firewall/check and /api/v1/firewall/unban also send
"success": true for callers written against the older API. It will be
removed; use the status code and ok.
Lists
Every GET route that returns a list answers a JSON object, never a bare
array. The list is under items. An empty list anywhere in a response is
[] and an empty map {}, never null; null means a value is not known.
totalis the number of matches the server counted. It is larger than the length ofitemswhen the route pages or cuts the list.offsetandlimitcome with routes that page or cap the list. Capped routes that do not support paging reportoffset: 0.truncatedis true when matches were left out: past the page, past the limit, or past a scan cap. When a scan cap stopped the count,totalcounts only what was scanned. Incident groups also sendscan_truncatedto distinguish a scan cap from a page limit; theirtotalis exact whenscan_truncatedis false.- Other keys next to
itemsdescribe the whole list, such ascheck_typeson/findings/enrichedorsummaryon/email/forwarders.
/audit, /threat/top-attackers and /threat/events do not count every
match. They send limit and truncated without total.
{"items": [{"ip": "203.0.113.9", "reason": "wp_login_bruteforce"}], "total": 1}
Times and durations
Every time is an RFC 3339 instant in UTC with sub-second precision, such as
2026-09-22T10:04:05.123456Z. The same applies to the event stream. A time
that is not set is left out, never sent as 0001-01-01T00:00:00Z.
The API does not send times it formatted for reading, such as “5m ago”,
“1h2m” or a clock time without a date. Clients format instants in their
own time zone and count down to an expires_at themselves.
Durations are numbers of seconds in keys that end in _seconds, such as
uptime_seconds, elapsed_seconds, duration_seconds,
oldest_age_seconds and update_interval_seconds.
Temporary whitelist responses use duration_seconds, rollback status uses
remaining_seconds, and incident timelines use window_seconds. Request
parameters such as hours and editable configuration values retain their
documented units.
Two values keep a text form. started_at_token is an opaque token for
restart polling. The temporary reason on /api/v1/firewall/check keeps its
“(expires in …)” text for existing callers; expires_at carries the
instant.
Severity
A severity is its label: CRITICAL, HIGH or WARNING. That covers
severity, severity_max, demoted_from and the attack event sev.
Clients derive colours and sort order from the label. A severity filter
takes the label in any case, or the older 0, 1 or 2 level. The findings
inside an email quarantine entry keep the scanning engine’s own rating.
Status & Data
GET /api/v1/status Full health snapshot: version, uptime, watchers, severity counts,
store health, blocklist size, capabilities[], config_hash, binary_hash,
automation rollout state, challenge pending count, rollback state.
`started_at_token` changes after a daemon restart and is suitable
for restart polling.
`security_posture` is `healthy`, `warning`, or `critical` after
combining daemon faults with active high/critical incidents.
`incidents_open_by_severity` breaks open and contained incidents
down by `critical`, `high`, and `warning`.
`correlation_attribution` (present after the first active-set
merge) lists per check the findings correlation could not
attribute to an account: `current` for the active set now,
`cumulative` since daemon start.
`queues` reports named protection queues, including depth,
capacity, running work, recent and cumulative drops, and lag.
A degraded queue changes `status` and `security_posture`.
`latest_scan` is the canonical last-scan timestamp; `last_scan_time`
is a legacy alias kept for older clients and will be removed.
`uptime_seconds` is the time since the daemon started.
GET /api/v1/challenge/stats Challenge-routing activity for the UI: `pending`, `escalated`
(timeouts that became hard blocks), `routed_by_check` (per source
check, since restart), and `recent` routes. Read scope.
GET /api/v1/capabilities Static feature list (e.g. `confd.dropins.v1`, `events.sse.v1`,
`webhook.phpanel.v1`, `webui.prefs.v1`, `webui.undo.v1`,
`mail.queue.composition.v1`,
`detect.http_scanner_profile.v1`, `challenge.stats.v1`,
`firewall.rollback.v1`, `firewall.dos_exempt.v1` on Linux builds).
Use for orchestrator feature-detect.
GET /api/v1/components Watcher/component matrix with attachment, event, and upstream freshness state.
GET /api/v1/events Server-Sent Events stream of findings as they dispatch.
Read-scope token sufficient. One JSON event per `data:` line.
Writes and flushes have a three-second deadline. A failed
write or flush closes the stream and frees its subscriber slot.
GET /api/v1/health Daemon health (fanotify, watchers, engines)
GET /api/v1/findings Current active findings
GET /api/v1/findings/enriched Enriched findings with GeoIP, accounts, fix info, and a list version.
?limit=N orders by severity, then newest, even when all rows fit;
returns at most N rows with limit and truncated, while total and
the severity counts cover all.
?fields=version returns only {version, total}, for change polling.
block_ip is set only for checks that report an attacker address,
the same evidence auto-block acts on
GET /api/v1/finding-detail Finding detail with action history (?check=&message=)
GET /api/v1/history Paginated history (?limit=&offset=&from=&to=&severity=&search=&checks=).
checks is a comma-separated list of check names to include.
total counts every match; truncated is true when matches exist past the page
GET /api/v1/history/csv CSV export of the newest 5,000 entries matching the /history filters
GET /api/v1/stats 24h severity counts, accounts at risk, auto-response summary
GET /api/v1/stats/trend 30-day daily severity counts; each `date` is a calendar day in the server's time zone
GET /api/v1/stats/timeline Hourly severity counts for the last 24 hours; each bucket names its `start` instant
GET /api/v1/quarantine Quarantined files with metadata (incl. htaccess pre_clean backups)
GET /api/v1/quarantine-preview Preview quarantined file content (?id=)
GET /api/v1/db-object-backups db_object_backups bucket (MySQL trigger/event/procedure/function drops)
GET /api/v1/db-object-backup-preview Preview captured CREATE SQL (?key=)
GET /api/v1/blocked-ips Blocked IPs with reason and expiry
GET /api/v1/accounts Accounts a server-wide scan covers
GET /api/v1/account Per-account findings, quarantine, history (?name=)
GET /api/v1/audit Newest 200 UI audit log entries; truncated marks older ones. Each entry names the credential that acted (actor) and whether it came as an API token or a browser login (via)
GET /api/v1/export Export state (suppressions, whitelist)
GET /api/v1/incident Incident timeline (?ip=&account=&hours=)
GET /api/v1/performance Performance metrics snapshot (admin scope)
POST /api/v1/perf/fix-error-log Truncate a fixed-row error_log finding (admin scope, CSRF)
POST /api/v1/perf/fix-display-errors
Disable display_errors for a fixed-row config finding (admin scope, CSRF)
POST /api/v1/perf/fix-wp-cron Disable WP-Cron and install a system cron for a perf_wp_cron finding (admin scope, CSRF)
GET /api/v1/hardening Last stored hardening audit report (admin scope)
Date ranges
/api/v1/history, /api/v1/email/groups and /api/v1/email/relay-abuse take
from and to as a calendar date (YYYY-MM-DD) or an RFC 3339 time. A date is
a day in the server’s time zone, and to includes the whole day. An RFC 3339
to is exclusive. The web UI sends RFC 3339 times so a day follows the
operator’s time zone preference. A value in neither form is rejected with 400.
An empty or whitespace-only bound is treated as absent. Calendar days start at
their first valid local time, including when midnight is skipped or repeated;
a skipped calendar date covers an empty range. Fractional seconds are preserved
when filtering RFC 3339 bounds.
WordPress verification coverage
wordpress_verification contains core and plugins coverage when installation
history is available. The same fields appear in csm status --json; human status
shows the counts. Each entry reports verified, modified, unverified, unknown,
not_wordpress and last_attempt. unknown means discovery found a directory
but no completed attempt is recorded. modified means core verification
completed with integrity differences, including missing core files. A verified
plugin inventory can still contain vulnerable versions. The timestamp is the
latest recorded attempt in that group.
Counts retain the last observed results across restarts and cached scan cycles.
Overlapping scans retain the latest attempt and its consecutive failure history
in scan order. Discovery without an attempt does not replace verification evidence.
An absent group means there is no retained installation history, not proof of
coverage. An error field reports unreadable history instead of clean counts.
The first failed attempt increases unverified; repeated failures also produce
an installation-specific Warning. In host scans, when one cause stops more
installations than the per-cause limit, the warnings collapse into a single
finding that names the cause, the total and a sample, so a host-wide fault cannot spend the whole
alerts.max_per_hour budget. The limit applies across accounts for each kind of
verification and cause. Account scans keep installation-specific warnings.
A summary has no installation path and names an account only when all affected
installations share it. These findings are separate from queue health:
a command that completed with a refusal did not lose queue work. Existing queue
loss accounting remains responsible for interrupted or missing command work.
Protection queue health
status.queue_health.v1 advertises the queues map on status responses.
The same measurements appear under snapshot.queues in csm status --json
and as named checks in csm doctor.
findings.ingest, fanotify.analyzer, fanotify.staged_packages,
fanotify.dropper, fanotify.dropper_findings and spool.scanner report
waiting work, capacity, running work, cumulative losses,
losses during the last minute, the oldest waiting item’s age and the oldest
running item’s processing time.
Waiting work includes producers blocked on admission. Ingest work remains
running while the dispatcher holds or processes its batch, including startup.
Ages pause while the startup hold is in place and resume from its release, so
a long baseline scan is not reported as a stall; work lost during the hold
still counts. Releasing a hold preserves any future eligibility time for
deliberately deferred work.
Loss totals include the undelivered tail of a batch canceled during shutdown;
scan warnings intentionally excluded from alerts do not count as lost work.
Recovered file and spool scanner panics count as lost scan work. The workers
continue processing later events, but repeated failures still degrade health.
fanotify.reconcile reports the recovery queue in directory tasks, with
1,024 waiting slots. Repeated drops in a directory retain its original queue
age. A detached scan batch remains running while new drops queue separately,
including new work for a directory that is already being scanned.
Eviction, work older than the recovery scan window, unreadable directories or
candidate files, interrupted batches and unfinished shutdown work count as
failed recovery tasks. An incomplete task may still have scanned some files;
each directory task counts at most once. Directories and files that no longer
exist when recovery runs hold nothing left to scan and complete the task,
which keeps bulk extraction and package restores out of the loss counters.
The recovery window includes its cutoff; work completed exactly at that
boundary does not count as expired.
Staged package verification reserves capacity for its whole running batch.
Its full-queue timer starts when admission fills the queue and continues while
the verifier retains those slots, including between retries.
Files awaiting another attempt retain their original waiting age; retrying
does not reset lag. Shutdown counts files still awaiting verification after
the analyzer workers have stopped. Package metadata I/O cannot block health
polling for this queue.
Dropper candidates and held findings each have 16,384 waiting slots. Their
lag starts after the configured unlink TTL or the finding’s 45-second grace
period, respectively; intentional waiting does not count as overdue work.
Detached probes and emission batches remain visible as running work.
Processing time starts when each batch leaves its waiting queue, independently
of the timestamp used to decide which work is eligible.
Retries and refreshed observations retain earlier eligibility, and exhausted
probes, capacity refusals and unfinished shutdown work count as losses.
mail.delivery reports the selected mail-log reader’s 64-slot delivery queue.
Its running time includes parsing and delivery of resulting findings. Losses
include oversized complete file records, unreadable journal entries, canceled
admission, consumer failure and buffered records abandoned on shutdown.
Counters survive reader replacement and changes between file and journal
sources. They count records already read, including complete records abandoned
in file read-ahead, not unread source history or partial file records that have
not reached a newline.
mail.file_source reports sampled untransferred bytes after the first file
attachment, with depth_unit: bytes and capacity_unavailable: true. This
includes disk backlog, buffered read-ahead and raw partial-line bytes, even when
the reader retains only a bounded prefix of an oversized line. Transfer to
mail.delivery removes the source bytes before output admission can block.
These byte and record measurements are separate stages and must not be added.
A reader-owned sampler checks the current descriptor every two seconds, so
appends remain visible while delivery is blocked. Health reads memory only.
lag_basis: consumer_progress measures time without consumption while unread
work remains. Sampling or new arrivals do not reset this clock. A partial line
at EOF waits for more input and does not report a stalled consumer. Reader and
sampler operations also report processing_lag after one minute without
progress, including blocked reads, metadata calls and cleanup.
File I/O failures report source_io. A successful read cannot clear a failed
cursor check; each operation must recover. Failed metadata sampling makes depth
unavailable until a successful sample. Rotation and observed truncation retire
the old generation, and late samples cannot overwrite its replacement. The
initial attachment still starts at EOF; replacements start at the beginning.
Copytruncate detection retains its existing limit: truncation followed by growth
past the old descriptor offset between checks may be indistinguishable from an
append.
Complete records already held in read-ahead count once in delivery loss when
discarded. Unread disk data and unterminated fragments have no measured record
count; generation changes, source failures and shutdown preserve
dropped_lower_bound in the source row. A working file or journal replacement
clears the retired source’s current error while retaining that history. Source
path disappearance still follows the watcher’s existing grace period.
mail.journal_source reports the journal reader after its first successful
attachment. The cursor exposes no exact unread-record count or queue capacity,
so depth_unavailable and capacity_unavailable are always set, and
lag_basis: unavailable marks the missing backlog age rather than reporting a
waiting time of zero. The row measures the current operation instead. A selected
entry remains in flight until delivery acquires it; known unreadable or abandoned
selected entries count once in mail.delivery. The two stages must not be added
together.
Journal progress follows cursor advancement, entry reads, bounded idle waits,
output admission and close. One minute without operation progress reports
processing_lag, including a stuck cursor with no known selected entry. Cursor,
entry, wait and close failures report source_io; known abnormal exits remain
visible through actual cleanup. An unknown cursor outcome or unread shutdown
sets dropped_lower_bound without inventing a record count. Successful reads
restore current health; historical uncertainty survives reader replacement.
Successful attachment of a file source clears the retired journal’s current
error. Failed attachment leaves it visible. Health reads memory only and does
not call the journal or change its tail positioning, retry or delivery policy.
When their BPF backends are active, bpf.af_alg.output,
bpf.connection.output, bpf.execution.output and
bpf.sensitive_files.output report the 256-slot userspace delivery queues.
Their losses include decoding failures, admission refusals, consumer panics
and buffered output left after the reader stops. A received event remains
running until evaluation and delivery to the finding queue finish, including
events intentionally filtered during evaluation. Kernel ring occupancy and
reservation failures are separate from these userspace measurements.
bpf.connection.verdict reports the advisory annotation pool’s 256 waiting
slots and up to four running callbacks. Repeated findings for a pending
destination, reason and severity share one request and retain its original age.
Capacity refusal, callback failure and abandoned shutdown work count as lost
annotations, separately from lost findings. Callback errors release the pending
key and allow a later finding to retry; a successful retry preserves earlier
loss totals. Shutdown closes admission, waits for running callbacks and counts
the queued requests left behind. Cached answers remain available without new
work. Findings continue immediately when an annotation is unavailable.
processctx.enrichment reports 1,024 waiting requests and up to two running
process-context workers. It appears only after a BPF consumer initializes the
shared pool; reading health does not start workers. Overflow replaces the oldest
queued request and preserves the waiting age of the remaining requests.
Running time includes the process read, latency observer, account resolution and
cache write. Read errors and interrupted processing count as lost enrichment;
vanished processes and stale or unverified identities are completed filtering.
The queue dropped metric counts refused, evicted and abandoned requests;
queue health also includes failures after a worker receives the request.
The daemon stops the pool after BPF producers have joined, lets running reads
finish and counts buffered requests discarded at shutdown. Final loss evidence
remains available. This row measures enrichment requests, not each individual
deadline-based file read inside a request.
processctx.proc_reads reports the 64 shared slots for deadline-limited process
file and symlink reads, including process-start captures on the BPF path.
Waiting and running reads share this capacity. A slot remains occupied until
both the underlying read and its caller have finished; an undelivered result
still occupies its slot. A timeout counts the lost result immediately and keeps
a blocked syscall visible as running work until it returns. Timeout and a later
read failure count as one loss. Refusals at capacity and read failures count as
losses; missing process files are expected churn. Synchronous reads without a
deadline consume no slots. The initial process-directory check and the cached
boot-time read are synchronous and are not measured by this row.
smtp_rdns.resolves reports the 64 reverse-DNS lookup slots used by direct SMTP
egress detection. A slot remains occupied until its resolver and caller finish.
The one-second caller deadline also bounds healthy processing age; timing out
counts a lost result once while keeping a resolver still running visible.
Capacity refusals and resolver failures count as losses, while NXDOMAIN and
successful empty responses are normal negative results. Cache hits perform no
queued work. The bounded result cache is retained data, not waiting lookups.
Status can initialize the empty cache but never performs DNS or waits for its
cache lock. Synchronous lookups without a deadline use no slots or queue row.
email_password.mailboxes reports waiting mailboxes and the five concurrent
audits per scan, retaining each audit through finding collection and cache writes.
Concurrent scans have no fixed combined waiting capacity. A busy pool uses each
audit’s five-minute budget or shorter scan deadline; an idle slot with no dispatch
progress for one minute reports backlog lag. Admission and post-audit work have
their own one-minute budgets. Cancellation removes waiting demand without adding
losses, while actual audits remain visible until they return. Expired work, failed
verification, cache write failures and abnormal exits count once per mailbox.
Unsupported, malformed and over-budget hashes remain normal incomplete-audit
outcomes. Health retains loss counts after the batch drains and reads only memory.
Deadline losses are recorded where an audit actually stops unfinished. A deadline
arriving after a successful cache write does not turn that mailbox into a loss,
and context evaluation during drain never holds the health lock.
email_password.hashes reports the three shared password-verification slots.
Each slot remains occupied until its KDF and caller have both finished, including
after scan cancellation. email_password.waiting reports callers awaiting a
slot; it sets capacity_unavailable because concurrent scans have no fixed
global waiting limit. Both rows use the password audit’s five-minute check budget
for lag. The occupied pool also reports sustained full capacity. A deadline while
waiting counts one admission loss; a deadline after admission counts one lost
verification result. A later KDF error or abnormal exit cannot count it twice.
Explicit cancellation withdraws demand without loss; actual verification errors
still count. Successful matches and nonmatches are normal results. Rejected
input and already canceled callers start no queued work. These rows contain no
password, hash, mailbox or account data, and status never runs a verification.
The outer mailbox scan’s discovery, network enrichment and persistence are
separate work from these hash slots.
checks.executions reports dispatched check functions across host and account
scans. An execution remains present until both its function and caller finish,
including while its result awaits consumption or its function outlives a timeout.
Lag is evaluated against each call’s original deadline, including a shorter parent
deadline; a delayed function start does not reset it. Timeouts, panics and
abandoned results count once per execution. Explicit cancellation withdraws
demand without counting a loss, but a later function failure still counts.
This aggregate sets capacity_unavailable: separate scans have their own wrapper
limits, and timed-out functions can outlive those slots. It does not measure
checks awaiting dispatch, scan-job persistence or later automatic actions.
Status reads memory without waiting for a check or accessing the state database.
checks.plugin_inventory reports sites waiting for the five inventory workers
per refresh, retaining each site through command execution, result collection and
cache storage or failed-result cleanup. Concurrent refreshes have no fixed combined
capacity. Inventory uses a four-minute budget for its two bounded commands, or
the shorter check deadline; admission and result handling each have one minute.
A full pool stays healthy within those budgets. A free slot with no dispatch
progress for one minute reports backlog lag. Command, decoding and storage
failures count once per site, including when cleanup also fails. A command that
ran and returned output (on stdout or stderr) with an error, such as a tree
wp-cli will not inventory, answered the check and counts no loss. Missing command output still counts as
lost work, including when the command exited with an error. Cancellation
adds no losses; deadline withdrawal counts unfinished sites. Buffered sites stay
visible until the original workers stop consuming and the refresh joins them.
Actual commands remain in flight until they return. Health reads memory only and
retains loss evidence after recovery. Shared-refresh waiters remain owned by
checks.executions; they do not create another set of site jobs. Optional domain
lookup fallback and plugin metadata enrichment keep their existing behavior.
checks.wordpress_core reports installations waiting for the five checksum
workers per scan, retaining each through its command, integrity findings and
verified-file caching. The unit is installations; concurrent scans have no
fixed combined capacity. Command lag uses the two-minute command budget or a
shorter parent deadline. Result and cache work use one minute without progress.
A full finite batch stays healthy within those budgets; a free worker with no
dispatch progress for one minute reports backlog lag. Returned operational
failures and abandoned work count once per installation. A command killed by a
signal counts as failed work while retaining any partial integrity findings.
Recognized integrity results from a completed command, including deliberately
filtered output, complete without a queue loss. So does a command that ran and
refused the tree, such as a directory that is not a WordPress installation or
one whose configuration fails to load.
Missing command output still counts as lost work.
Cancellation withdraws unfinished demand without loss, while deadline withdrawal
counts unfinished installations. Commands ignoring cancellation remain in flight
until they return. A deadline during caching cannot undo completed verification,
and an ordinary cache read failure does not turn a verified site into lost work.
Health reads only metadata, including while result or cache locks are occupied,
and retains confirmed losses after recovery. Discovery is separate check work.
auto_block.waiting reports scan, direct-block, firewall-flush and startup
observation calls waiting for shared state, with no fixed waiting capacity.
auto_block.active reports the single state owner through firewall operations,
state writes and cleanup.
lag_basis: operation_progress times the current operation within a batch;
advancing batches do not degrade solely because their total duration exceeds a
minute. One minute without progress degrades the active row and any waiting
callers. A free state slot with no admission for one minute also reports lag.
Returned direct-block or flush errors, failed state reads and writes, and abnormal
exits count once per call; protected-address refusals do not. Known write failures
are recorded before readback and diagnostic output. Known errors remain visible
during later cleanup, and the common loss threshold and recovery policy apply.
These rows count state-lock callers, independently of persisted per-IP retry records.
Health snapshots use memory only and cannot wait for the state lock or I/O.
auto_block.pending reports durable retry records, including records loaded at
startup when automatic blocking is disabled. Its capacity is the retry admission
limit of 1,000 records. Records selected for a cycle remain in flight until their
state-file outcome is known; newly requeued records also remain in flight until
persistence finishes. auto_block.candidates reports distinct IPs admitted to the
current cycle, with no fixed capacity. Candidate and record counts describe
different stages and must not be added together. A candidate remains visible
through its firewall callback, subsequent bookkeeping and durable requeue.
Pending age uses its original queue timestamp, or first observation when that
timestamp is absent. An eligible record receives a timestamp on its first requeue;
records without check identity are withdrawn under the current block policy.
Refreshing the reason, check or severity does not reset the original queue age.
Normal quota waits stay healthy within the existing two-hour
retry lifetime; older waiting records report backlog_lag. Active record and
candidate work use one minute without operation progress, so advancing batches
can run longer without a false warning. These measurements add no retry scheduler
and do not change the hourly quota, expiry or overflow policy.
An unsuccessful attempt whose retry survives reports retry_failed without a
loss. Successful blocks, dry runs and expected refusals remain completed even
if later bookkeeping fails. Withdrawal under the current check policy or
login-blocking setting is an expected refusal and adds no loss. Persisted record
identities include check and severity, so a refused record cannot acknowledge a
different eligible retry after a failed write. Confirmed removal of expired,
invalid or overflowed records counts as pending loss; a fresh candidate that neither completes nor
survives on disk counts as candidate loss. Existing duplicate coalescence adds no
loss while a retry survives. The common loss threshold and recovery policy apply.
State read or write failures report state_io. Failed writes are read back
because an error after rename can leave the new state in place. If that read also
fails, depth_unavailable and dropped_lower_bound expose the uncertainty. A later
successful read restores measured depth; lifetime loss totals remain lower bounds
when earlier outcomes could not be established. New work with known outcomes
still contributes exact losses. Completed candidate history is released after
settlement, and health snapshots never read files or wait for firewall callbacks.
auto_block.cleanup reports distinct IPs awaiting bookkeeping cleanup after a
successful firewall flush. It observes saved cleanup retries at startup and admits
the flushed engine entries before reading the tracker. Cleanup also includes
tracked blocks, with duplicate IPs counted once. Work remains in flight through
store removal, threat-record cleanup and the final state-file outcome. This row
has no fixed capacity and is independent of the block-candidate stages.
Cleanup waiting age starts at first observation with lag_basis: deferred_checkpoint. Age alone does not degrade health: cleanup retries run on
the next explicit flush, without a background retry deadline. One minute without
active operation progress reports processing_lag. Failed cleanup whose retry
survives reports retry_failed; an old tracker block can retain a retry even
when saving its cleanup marker failed. Completed cleanup stays acknowledged
across failed saves, while a newly blocked IP creates fresh cleanup demand.
Recovery first establishes the known block baseline. Cleanup of a later block
has an exact outcome even when older cleanup history remains uncertain.
Cleanup state and snapshot failures report state_io. An unreadable tracker
sets depth_unavailable and dropped_lower_bound; speculative waiting records
are not reported as measured depth. The last bounded batch retains its identity
for later read recovery. A subsequent unreadable flush replaces that history
with its own batch. A readable state restores measured depth and counts confirmed
unfinished work that no longer has a retry source. An incomplete pre-flush engine
snapshot marks lifetime losses as a lower bound without inventing a missing count.
Health reads memory only. These measurements preserve the existing flush policy,
returned errors and cleanup retry behavior.
attackdb.events reports events awaiting persistence, including historical imports.
The buffer has no fixed capacity. Waiting age starts at admission, independently
of an event’s timestamp. A detached batch remains in flight through writes, close
and error logging; newly arriving events stay in the waiting count. Processing
age uses lag_basis: operation_progress and resets when a write returns, so a
progressing batch does not appear stalled solely because of its total duration.
Waiting age or absent write progress of one minute degrades health.
Completed JSONL records accepted by the file writer and successful database
writes count as persisted. Buffered data alone does not. Confirmed unwritten
events count as losses, including interrupted work left buffered or unencoded, and
returned encoding failures. These losses are visible before cleanup; close errors
or abnormal I/O exits preserve uncertainty
with dropped_lower_bound and report persistence_uncertain for one minute,
ahead of a backlog or a stalled write, as the record queue already does.
The common loss threshold and recovery policy apply to confirmed losses.
These measurements preserve the existing write, retention and shutdown policy;
they do not add retries or promise storage durability beyond the writer’s result.
Memory-only databases have no persistence backlog. Health reads queue memory,
independently of database state locks and I/O.
attackdb.records reports distinct IP records awaiting a saved update or deletion.
Repeated changes to one IP coalesce before the snapshot; a mutation during an
active write remains separate pending work. A successful deletion can also
satisfy a repeated deletion. Loaded records needing normalization and expired
records needing removal enter the same queue. Retained scoring records are not
counted as pending persistence. There is no fixed capacity.
Waiting age starts at the first pending change. The active snapshot stays in
flight through writes, error logging and retry bookkeeping. Processing age uses
lag_basis: operation_progress and resets at each returned store write or delete,
or at the complete flat-file write result. One minute without progress or with
waiting work degrades health. Returned failures report retry_failed while the
retry remains queued, preserving its original age without counting it as lost.
Successful retries clear that condition. A failed shutdown flush retains the
in-memory retry; the stopped background saver does not schedule another attempt.
An interrupted operation counts confirmed abandoned demand only when no pending
update or deletion survives. Unreturned write outcomes report
persistence_uncertain for one minute and retain dropped_lower_bound afterward.
Known losses remain counted through recovery. Health reads queue metadata only,
without waiting for database state or storage locks. These measurements preserve
the existing persistence, retry and shutdown behavior; they do not make retained
in-memory retry state durable across process exit.
block_digest.records reports the actual buffer of watched-country block
records, bounded to 5,000 with drop-oldest overflow. Age starts when a record
enters the buffer, independently of the block timestamp. Normal batching and a
full buffer alone do not report a stall: waiting becomes overdue one minute
after the configured interval. Detached digest preparation remains in flight;
one minute without progress reports processing_lag. The row uses
lag_basis: operation_progress. Records arriving during preparation belong to
the next window. Expected per-IP coalescing and delivery filtering are successful
completion, while abandoned preparation counts its discarded records as losses.
block_digest.email and block_digest.webhook report configured destinations
in notifications, with capacity_unavailable. Each eligible destination owns
one notification before delivery starts, including live alerts and configured
empty heartbeats. Default delivery follows current alert settings at admission,
attempt and interrupted cleanup. Disabled destinations add no new waiting work
or loss; destinations enabled during an earlier send are still handled. Explicit
destination selection keeps its error when that alert channel is disabled.
One minute waiting or active reports lag. A returned sink error
counts one lost notification and stays in flight through error logging. If an
operation exits without returning, its attempted delivery reports
delivery_uncertain for one minute and retains dropped_lower_bound; a remaining
destination that was never attempted counts one confirmed loss. Known losses
use the common warning threshold and survive recovery in lifetime totals.
These observations preserve the existing digest schedule, live deduplication,
best-effort delivery and shutdown policy. Records retained after the ticker
stops report consumer_stopped and can still be explicitly flushed; no retry or
durability is added. Health reads metadata independently of collector state
locks, country lookups and delivery callbacks. Disabled collectors have no rows.
state.pending reports findings parked for the next startup, bounded to the
newest 10,000 findings. The store observes existing parked work when opened.
Depth is the last confirmed file contents; in-flight findings include incoming
appends and cleared batches still in replay. The bound applies to the stored
batch, not concurrent callers or an already detached replay. Overflow and new
findings that fail to reach storage count as losses. Duplicate occurrences
remain separate findings.
Waiting age starts at admission or startup observation, independently of the
finding timestamp. Surviving findings keep their age through later appends;
evicted findings no longer determine it. Waiting uses lag_basis: deferred_checkpoint: a parked or full batch awaiting restart does not itself
report a stall. state.pending_operations measures callers waiting for the
state lock and active persistence or replay, in operations with no fixed bound.
One minute waiting or without operation progress reports lag. Actual read,
write and clear results advance progress; replay stays active through dispatch.
Read, write and clear errors report state_io. Failed mutations are read back:
a returned error after replacement does not imply the new findings were lost.
Readback compares the recovered JSON payload, including its normal repair of
invalid UTF-8, so repaired log text does not invent uncertainty or hide overflow.
Unreadable outcomes set depth_unavailable and dropped_lower_bound. A later
read restores measured depth, while lifetime losses remain lower bounds. Known
encoding failures and overflow are counted even when other outcomes are unknown.
Indistinguishable retained payloads keep the earlier age after a failed write.
An interrupted write or clear marks disk state unknown before releasing the
state lock. Interrupted replay retains measured disk depth but cannot claim an
exact loss for a partially dispatched batch; uncertainty reports
persistence_uncertain for one minute. Confirmed losses use the common warning
threshold and remain in lifetime totals after recovery.
The file format, append order, returned errors and clear-before-replay policy are unchanged. These observations add no retry or extra replay. Health snapshots read metadata only and do not wait for state locks, files or dispatch callbacks.
incident.persist.waiting reports immutable incident snapshots waiting for the
ordered writer, with no fixed waiting capacity. incident.persist.active reports
the single occupied writer. A writer or free-slot admission stalled for one minute
reports lag. Failed writes and abnormal callback exits count once; abandoned bulk
snapshots count as waiting losses and release their ordering slots for later writes.
Callbacks remain owned through cleanup and error logging. Store failures retain
in-memory transitions and the existing warning log.
incident.persist.deferred reports coalesced bookkeeping waiting for a later
mutation or explicit flush, including shutdown. Its oldest age uses
lag_basis: deferred_checkpoint; age alone does not degrade health because there
is no periodic flush deadline. Full snapshots supersede earlier bookkeeping;
restoration and retention discard the affected markers. Memory-only correlators
have no persistence work. All three rows read memory independently of state locks
and database I/O. The common loss threshold and recovery policy apply to writes.
checks.file_index.waiting reports live scans waiting for the shared baseline
slot, with no fixed waiting capacity. Each wait keeps its shorter parent deadline
or the file-index check budget; a free slot with no admission for one minute also
reports backlog lag. checks.file_index.active reports the single occupied slot.
Walking and file analysis use the check’s original execution deadline (normally
15 minutes) or the shorter parent deadline. Setup, persistence and cleanup each
have a separate one-minute budget. A healthy long scan alone does not report a
full queue. Actual filesystem work remains visible after cancellation until it
returns. Deadline withdrawal, incomplete walks, unreadable PHP content, failed
executable metadata or state reads and writes, and abnormal exits count once per
live scan. Executable entries disappearing during enumeration and explicit
cancellation alone add no loss.
The common three-loss warning threshold applies, and recovery retains total losses.
A successful late baseline commit adds no loss. Findings and shrink protection
keep their existing behavior. Force-file-index audits bypass this stateful slot
and remain owned by the check execution row. Snapshots read memory only.
checks.reputation_queries publishes accepted lookups immediately after quota
reservation, including while refused lookups receive fallback scoring. It follows
the accepted work through HTTP,
response cleanup, buffered results, supplemental scoring and cache or quota-backoff
storage. Each scan has at most five requests; concurrent scans have no fixed
combined capacity. HTTP uses the client’s timeout, while cleanup, local result
handling and persistence have separate one-minute budgets. Supplemental scoring
uses the combined timeout of its enabled sources or the shorter parent deadline.
Buffered results follow their own active consumer’s deadline; another scan cannot
hide their delay. Query failures, failed storage and abnormal exits count once per
result. Quota responses and refusal before dispatch remain expected outcomes with
the existing quota health warning. Parent cancellation does not cancel the existing
HTTP requests; they remain visible until response handling and storage finish.
Successful late results add no loss. Health reads memory only and retains losses
after recovery. Local feed matches, cache hits and serial pre-query discovery do
not create query jobs. A failed cache write retains the findings already produced.
checks.dispatch reports pending checks and occupied runner wrappers across host
and account scan batches. Concurrent batches have no fixed global waiting limit,
so capacity is unavailable. Lag uses consumer_progress within each batch:
pending work degrades after one minute without progress when that batch has a
free worker slot. A busy pool may keep working within each check’s own deadline.
Setup and result handling have a separate one-minute budget; they cannot borrow
a heavy check’s longer deadline. A wrapper panic or abnormal exit counts a loss,
as does a parent deadline before a check can start. Explicit cancellation
withdraws queued demand without loss. Execution failures belong to
checks.executions; dispatch tracks the surrounding scheduling operation.
Reading status never takes a scan, context or database lock.
scans.jobs reports eight waiting full-scan jobs and one worker. Work remains
visible through enumeration, scanning, remediation, persistence and cleanup.
Its consumer_progress lag follows that worker and its own check batches:
healthy check deadlines allow long scans to proceed, while an overdue child
cannot be hidden by another child’s progress. Enumeration, result handling,
individual file actions and database operations each have a one-minute progress
budget. Thirty seconds at full waiting capacity also degrades health.
scans.admission separately reports callers waiting for admission and writing
their initial job record, with a one-minute lag budget and unavailable capacity.
Status reads only memory; operator polling does not advance the worker’s clock.
Queue refusal, failed job operations and abandoned work count once per job.
Configured finding-history truncation, explicit cancellation and successful
shutdown draining do not count as loss. A terminated worker refuses new jobs.
At startup, both queued and running records from the previous process become
errors with reason daemon_restarted; their lost requests count in job health.
email_av.scans reports antivirus engine work through execution and both result
handoffs. Timed-out engines remain in flight until they return; buffered results
remain queued until the message scan consumes them. Execution lag uses the
original per-part deadline, while result waiting and processing each have a
one-minute budget. Moving a result between buffers preserves its waiting age.
Concurrent messages and engines outliving timeouts have no fixed global limit,
so capacity is unavailable. Timeouts, errors and abandoned work count once per
engine scan; later failures cannot count the same scan twice. Detections and
unavailable engines are normal outcomes, and mail verdict behavior is unchanged.
Status reads only memory, and watcher restarts reuse the same engine health.
php_taint.requests follows callers waiting for the worker lock, active requests
and pipe operations that outlive their callers. Waiting, setup, reply handling
and cleanup each have a one-minute lag budget. Active worker communication uses
its configured timeout or the shorter caller deadline; a buffered reply cannot
borrow a long worker timeout. The request stays in flight through reply decoding
and cleanup, and until any outstanding pipe operation finishes. Capacity is
unavailable because callers have no fixed global waiting limit. Worker failures,
timeouts, breaker refusals and abnormal exits count once per request; known
failures are recorded before cleanup. Oversize input, caller cancellation and
refusal after an intentional stop do not add losses. Status reads memory without
the worker lock, and shutdown retains the supervisor’s health evidence.
Recovered analyzer panics also count as failed work when delivered in a valid
worker reply. The report and worker reuse policy remain unchanged.
central.actions reports 1,024 waiting central-intelligence actions and one
running action. Backlog remains visible while the signed feed refreshes;
processing time includes the action handler and its evidence delivery.
Overflow, action failure and abandoned shutdown work count as losses. A
protected-address refusal or absent firewall engine completes the queue task
without claiming that a block happened. Challenge tasks measure delivery to
the challenge list; they do not measure its later file or firewall writes.
Shutdown cancels feed refreshes, waits for an already-running action and
discards the remaining queue. Captured dispatch hooks refuse later work and
preserve its loss count after shutdown. Logging does not reset loss totals.
bot_verification.requests reports 256 waiting bot-identity requests and one
running verification. Duplicate requests share their original waiting age.
Running time includes DNS lookups and the cache write. Queue overflow, DNS
timeouts or transient failures, failed cache or missing-PTR record writes and
abandoned shutdown requests count as losses. Missing PTR records, unknown bot
identities and definitive positive or negative answers remain expected
outcomes. A source with a recorded missing PTR is not queued again until its
one-hour suppression lapses; suppressed requests do not count as losses.
For a day after the lookup, lapsed records prevent renewed pending grace
without blocking DNS retries.
Queue health does not turn a resolver failure into a spoof finding. Shutdown
cancels DNS, waits for any active cache write and discards waiting work. Late submissions
are refused, pending keys are released and loss totals remain available.
abuse_reporting.ingress measures reports awaiting durable storage, with a
capacity of 10,000 or the configured spool limit when smaller. Running time
includes persistence to every configured target. Memory overflow, incomplete
persistence and work abandoned by a worker failure count as losses; one source
report counts once even when several targets fail. Outbound delivery can delay
memory work, and that backlog remains visible. The queue stays open through
the daemon’s final finding flush. Reporter shutdown then closes admission and
persists every accepted report before closing the spool; a failed write does
not discard unrelated waiting reports. Captured hooks refuse later reports.
This row measures the memory queue, not reports already retained on disk for
retry; normal shutdown does not count those durable reports as lost.
abuse_reporting.spool measures durable reports per destination, with the
configured spool capacity. Reports already being sent remain in flight through
the database acknowledgment. Failed admission and discarded records count as
losses; a delivery retry, failed acknowledgment or normal shutdown retains the
report and does not count it as lost. An evicted record counts as lost only if
no send was acknowledged during this process. An active send settles that count
when it finishes; a failed retry cannot undo a prior acknowledgment.
lag_basis: observed_age means waiting age starts when this process first sees
the record. Existing records start at spool open; retries keep that age. Doctor
labels this as observed_lag, since time spent waiting before restart is unknown.
Waiting or running work degrades after two minutes, allowing for the normal
one-minute delivery interval. The common full-queue and recent-loss thresholds
also apply. spool_io names failed admission, read or acknowledgment operations;
delivery_failed names a failed or panicking sender. These states clear when
the affected operation recovers. Health reads remain independent of database
writes and outbound requests. Reopening with a smaller cap preserves existing
records until the next admission applies the configured overflow policy.
phpanel.spool measures the durable panel webhook queue, capped at 100,000
records per state directory. In-flight work includes writes waiting to commit
and deliveries awaiting database acknowledgment. Retried findings keep their
original waiting age; a failed attempt is not a lost finding while its record
remains durable. An overflowed record counts as lost only when no send was
acknowledged during this process; an active send settles that count when it
finishes. Malformed findings count once when removed from delivery;
their bounded diagnostic archive is retained history, not pending work.
Waiting age uses the persisted enqueue timestamp after restart. Boundary
records absent from the accounting rebuilt at open are adopted with the same
timestamp, so retrying them does not reset their age. Missing, damaged or future
timestamps are timed from queue open or adoption. A minute of waiting
or processing degrades the row; delivery and database errors remain visible
until the affected operation succeeds. Health reads use memory only, so a
stalled database or collector cannot block status. Disabling delivery preserves
the durable backlog and cumulative loss; enabling it resumes the stored work.
Stopped queue instances refuse late admissions.
actionlog.writes measures the 64 process-wide action-log write slots. A sink
write that outlives the caller’s 250ms wait budget stays in flight until the
sink and any panic reporting finish; a batch of actions serialising behind one
file lock is normal, and only a minute without the write returning degrades the
row. A caller deadline after
admission does not count as loss, since the sink may still record the action.
Refused admission, sink errors and panics count as lost records, once per record.
Changing or disabling the sink preserves outstanding work and cumulative loss.
Health reads do not wait for the sink. Normal writes still complete before the
caller returns; saturated or stalled recording keeps the existing caller budget.
events.deliveries aggregates live event-stream subscribers in one row without
client identities. Capacity sums their buffers (64 findings per daemon stream);
depth counts waiting deliveries and in-flight work includes encoding, writing
and flushing. One full subscriber can degrade this row even while others drain.
The common one-minute lag, 30-second fullness and recent-loss thresholds apply.
Overflow, encoding errors and failed streams count as losses; a failed stream
also counts its abandoned buffered events. Normal request cancellation and
server shutdown withdraw pending demand without adding losses. Cumulative loss
survives subscriber removal, and an outstanding write remains visible until it
returns. Closing the bus preserves buffered work for consumers still draining it.
HTTP writes retain their three-second deadline; successful delivery here means
the write and flush returned, not that the remote application acknowledged it.
Each active BPF backend also exposes a .kernel row. depth_unit: bytes
labels ring occupancy and capacity. lag_basis: consumer_progress means
lag_seconds measures time without observed consumption while data remains,
starting at the first sample that sees pending data. It is not an event age.
One minute without progress degrades the row; recently observed reservation
failures use the same loss threshold as userspace queues. Kernel loss counters
are separate from decode failures and userspace admission loss.
After a reader unmaps its ring, depth_unavailable: true and
lag_basis: unavailable prevent zero fields from claiming an empty live ring.
The last counter sample adds submitted records the reader never consumed.
Submission accounting precedes publication, so a fast reader cannot consume
an event before it has been counted.
dropped_lower_bound: true marks this final shutdown total: kernel detachment
can leave callbacks finishing after the sample. Doctor prints dropped>=...
for that bound. A counter lookup failure or an incoherent occupancy reading
marks the measurement unavailable and degrades the row as
measurement_unavailable once it lasts half a minute, so one artefact of
reading a live ring raises nothing; a failed final sample degrades at once and
remains degraded. A stopped reader is reported instead of the measurement
artefacts it causes. Invalid occupancy cannot inherit fullness or consumer
stall alarms from an earlier sample; independently counted losses remain visible.
The required kernel suite fills the shipped connection program’s ring with
real non-root connect calls and checks reservation loss and retained output.
fanotify.kernel and spool.kernel report pending notification records and
use the same consumer-progress lag measurement. Their group capacity is not
exposed by the kernel: capacity_unavailable: true marks it as unknown, and
doctor prints depth=N/unknown records. The current system queue limit may
differ from the limit captured when the watcher was created.
Their loss totals are lower bounds: each overflow record proves at least one
loss, and shutdown adds the records known to be unread before closing the
descriptor. Events can still arrive between that sample and close.
An unavailable pending-record measurement degrades the row once it lasts half
a minute; the reading taken at close degrades at once and remains degraded.
Records the kernel has already dropped are reported ahead of an unreadable
depth. Closing a descriptor leaves known zero occupancy.
fanotify.reader and spool.reader track batches after a kernel read, until
all records have been dispatched or filtered. A stalled batch remains visible
even when the kernel queue is empty. Reader losses count batches interrupted
by consumer failure, separately from kernel-record losses.
Spool replacements retain loss and running-batch evidence, while the new
descriptor starts its own occupancy and progress measurements. Health reads,
event reads and permission responses cannot use a descriptor after close.
The required kernel suite verifies pending records, stalls, drain recovery
and unread shutdown loss using real fanotify events.
forwarder.kernel and phprelay.kernel report inotify backlog in bytes,
because records include variable-length filenames. Capacity is unknown and
lag follows observed consumer progress. Overflow markers each prove at least
one lost event. Closing with unread bytes adds one further known loss; the
byte count does not reveal the number of discarded events. These totals use
dropped_lower_bound: true.
forwarder.reader and phprelay.reader track the running read batch, including
synchronous callbacks. Failed callbacks count as interrupted batches. PHP
relay replacements retain earlier losses with fresh occupancy measurements.
Descriptor reads, watch changes, polling and close share one lifetime guard.
phprelay.index.persistence reports the message attribution writer’s 4,096
waiting slots. Writes remain running from channel receipt through the pending
batch and its database transaction. Batches contain at most 256 writes, also
during an explicit flush. Each flush covers the backlog present when the
writer accepts it; concurrent arrivals cannot extend that flush indefinitely.
Refused submissions, encoding failures and writes in a failed transaction count
as losses. Failed writes are settled before emitting their error finding, so a
blocked reporter cannot hide the loss. A failed transaction counts each of its
writes once and does not prevent subsequent batches from committing.
Shutdown closes admission and drains accepted writes; later submissions count
as refused work. The existing persistence dropped metric counts refused
submissions, while queue health also includes writes that fail after admission.
Some rows carry advisory. Their work is best effort: a client that stops
reading its event stream, an unreachable panel asked for an optional
annotation, and expired process context reads all lose detail around findings
that are still detected, stored and delivered. An advisory row reports its own
degradation with the same evidence, but leaves the host status and security
posture unchanged, warns instead of failing csm doctor, and raises no
notification.
Each degraded row names its condition: backlog_lag for the oldest waiting
item past its budget, processing_lag for the oldest running item,
queue_full for capacity held continuously, dropped_work for recent losses,
consumer_stalled for a sampled queue whose consumer stopped advancing,
reader_stopped for a released kernel descriptor, measurement_unavailable
for a reading that stayed unavailable, and persistence_uncertain for a write
whose outcome is unknown. The owner-specific conditions retry_failed,
source_io, spool_io, state_io, delivery_failed, delivery_uncertain
and consumer_stopped are described with their rows above.
Doctor names each age by what it measures: observed_lag for an age taken at
observation, consumer_stall for consumer progress, operation_stall for the
current operation, deferred_age for work parked until a restart, and
lag=unavailable where no age exists.
A queue becomes degraded after three losses in a minute, thirty seconds
continuously full, or a minute waiting or processing. These are operational
alert budgets, not measured throughput guarantees. Health is computed directly
from the counters, independently of finding delivery. The daemon records and
dispatches protection_queue_degraded at most once per queue every five
minutes while pressure remains, then one protection_queue_recovered event.
The five-minute bound spans recoveries, so a queue that clears and degrades
again inside the window stays degraded in status without a second event, and
no recovery event follows a degradation the bound suppressed.
These are CSM health events and do not feed account-compromise correlation or
automatic response. Recovery preserves cumulative loss evidence; restarting a
spool watcher also preserves it. Restarting the daemon resets the counters.
Inspect worker errors and CPU, memory and I/O pressure when a queue degrades. Reduce competing bulk work and confirm the queue drains and recent losses stop. This surface currently covers finding delivery, file and spool kernel readers and scanners, recovery scans, staged package verification, dropper processing, BPF queues, process context enrichment and deadline reads, mail-log delivery, forwarder and PHP relay notification queues, and PHP relay index persistence; other bounded queues remain in the roadmap.
GeoIP
GET /api/v1/geoip IP geolocation (?ip=&detail=1)
POST /api/v1/geoip/batch Batch GeoIP lookup (body: {"ips":["192.0.2.1"]}, maximum 500)
Threat Intelligence
GET /api/v1/threat/stats Attack stats, type breakdown, hourly trend
GET /api/v1/threat/top-attackers Top attacking IPs with GeoIP (?limit=)
GET /api/v1/threat/ip IP threat lookup (?ip=)
GET /api/v1/threat/events IP event history (?ip=&limit=)
GET /api/v1/threat/whitelist Whitelisted IPs
GET /api/v1/threat/db-stats Attack database statistics
POST /api/v1/threat/block-ip Block IP for 24 hours
POST /api/v1/threat/block-ip-permanent Block IP with no expiry
POST /api/v1/threat/whitelist-ip Permanent whitelist
POST /api/v1/threat/temp-whitelist-ip Temporary whitelist (with expiry)
POST /api/v1/threat/clear-ip Clear IP from attack database
POST /api/v1/threat/unwhitelist-ip Remove from whitelist
POST /api/v1/threat/bulk-action Bulk block (24h or permanent) / whitelist across many IPs
The 24 hour endpoint returns 409 Conflict when the address already has a
permanent or longer firewall block. Unblock it explicitly before shortening
its lifetime. Bulk actions use action: "block", "block_permanent", or
"whitelist"; refused blocks are excluded from count and explained in
warnings. Repeated canonical IPs count once. Permanence cannot be selected
through an extra request field.
Bulk undo restores the original per-IP firewall deadlines. A later change to any target invalidates that undo action; it cannot restore dismissed evidence or replace the later decision.
Firewall
GET /api/v1/firewall/status Config, blocked/allowed counts
GET /api/v1/firewall/allowed Whitelisted IPs
GET /api/v1/firewall/subnets Blocked subnets
GET /api/v1/firewall/audit Firewall audit log
GET /api/v1/firewall/check Check if IP is blocked (?ip=)
POST /api/v1/block-ip Block an IP
POST /api/v1/unblock-ip Unblock an IP
POST /api/v1/unblock-bulk Bulk unblock IPs
POST /api/v1/firewall/allow-ip Allow an IP
POST /api/v1/firewall/remove-allow Remove IP from allow list
POST /api/v1/firewall/deny-subnet Block subnet
POST /api/v1/firewall/remove-subnet Remove subnet block
POST /api/v1/firewall/flush Clear all blocks
POST /api/v1/firewall/unban Unblock IP + flush cphulk
POST /api/v1/firewall/cphulk-clear Flush cphulk bans only
The audit log reports each timestamp as an RFC 3339 instant in UTC and
the block or allow lifetime as duration_seconds. Blocked addresses,
subnets and allow rules carry expires_at, left out for a permanent entry.
ModSecurity
GET /api/v1/modsec/stats WAF statistics (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET /api/v1/modsec/blocks Blocked requests log, aggregated per IP, with resolved source country (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET /api/v1/modsec/events WAF event details with resolved source country (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET /api/v1/modsec/rules Loaded rules list
POST /api/v1/modsec/rules/apply Apply the set of disabled rules and reload
GET /api/v1/modsec/rules/escalation Rule IDs excluded from firewall escalation, sorted
POST /api/v1/modsec/rules/escalation Exclude one rule from escalation or turn it back on
POST /api/v1/modsec/rules/escalation takes {"rule_id": 900112, "escalate": false}. The rule ID must be a CSM rule (900000-900999). escalate: false adds the exclusion and escalate: true removes it; other excluded rules are left as they are.
A failed rules reload reports rolled_back: true only when restoring the previous
overrides succeeds. A rollback failure is reported in the response and audit log;
inspect the overrides before trying another reload.
Rules & Suppressions
GET /api/v1/rules/status YAML/YARA rule counts, version
GET /api/v1/rules/list Rule files
GET /api/v1/suppressions Suppression rules
POST /api/v1/rules/reload Reload signature rules from disk
POST /api/v1/suppressions Add a suppression rule
DELETE /api/v1/suppressions Delete a suppression rule by id
POST /api/v1/suppressions takes {"check", "path_pattern", "reason"}. The path pattern is a glob and must be valid. A rule that covers every path of a check, hiding all its findings and stopping their remediation, needs "all_paths": true and no path_pattern; an empty pattern without it returns 400. check must be a check name, not a pattern: letters, digits, _, ., : and -. A name that is neither a known check nor the check of a current finding is saved, and the response carries a warning saying the rule matches nothing yet.
GET /api/v1/email/stats Email scanning statistics
GET /api/v1/email/forwarders Mail forwarder inventory with destination providers and local-copy flags (read scope)
GET /api/v1/email/deferrals Outbound deferral rollup by provider and sending IP with reason codes, parsed from exim_mainlog (read scope)
GET /api/v1/email/queue-composition Mail queue makeup: real vs null-sender bounce backscatter, frozen count, oldest age, top stuck recipients (read scope)
POST /api/v1/email/queue/flush-backscatter Request removal of frozen null-sender messages from the exim queue on cPanel hosts; returns the count of targeted messages no longer queued, reports incomplete verification as 500, or returns 503 when unavailable (admin scope, CSRF)
GET /api/v1/email/held Forward copies held by the forward guard (admin scope)
POST /api/v1/email/held/{id}/release Re-inject a held forward copy to its external recipient (admin scope, CSRF)
DELETE /api/v1/email/held/{id} Discard a held forward copy (admin scope, CSRF)
GET /api/v1/email/groups Server-grouped action rows (kind=compromised_account|spam_outbreak|auth_failure|queue_alert|malware) with from/to/limit (read scope)
GET /api/v1/email/relay-abuse Outbound PHP-mail abuse detections (spam outbreaks, high-volume scripts/accounts) with per-site script breakdown; from/to/limit (read scope)
GET /api/v1/email/quarantine Quarantined email list
GET /api/v1/email/av/status Email AV watcher status
GET /api/v1/email/quarantine/{id} One quarantined message
POST /api/v1/email/quarantine/{id}/release Release a quarantined message (admin scope, CSRF)
DELETE /api/v1/email/quarantine/{id} Delete a quarantined message (admin scope, CSRF)
Hardening
GET /api/v1/hardening Load last hardening audit report (admin scope)
POST /api/v1/hardening/run Run hardening audit and save report (admin scope, CSRF)
Scan Jobs
csm scan --full enqueues full-scan jobs that run inside the daemon and persist to the store. Jobs are report-only unless an account-scope request sets quarantine: true; server-wide jobs reject quarantine.
GET /api/v1/scan-jobs List full-scan jobs (read scope)
GET /api/v1/scan-jobs/{id} Job status and stored report (read scope)
GET /api/v1/scan-jobs/{id}/findings
Paginated findings for one job (?offset=&limit=, limit 500 by
default and at most 5000; truncated marks more pages) (read scope)
POST /api/v1/scan-jobs Enqueue a full-scan job (admin scope, CSRF)
POST /api/v1/scan-jobs/{id}/cancel Cancel a queued or running job (admin scope, CSRF)
Verified Bots
GET /api/v1/verified-bots Configured verified-crawler allowlist plus live verification state (admin scope)
POST /api/v1/verified-bots/apply Validate, apply, and reload an edited verified-bots list (admin scope, CSRF)
Actions
POST /api/v1/fix Apply fix for a finding
POST /api/v1/fix-bulk Bulk fix multiple findings
POST /api/v1/dismiss Dismiss one finding {key} or up to 500 {keys}; returns undo_token
POST /api/v1/scan-account On-demand account scan
POST /api/v1/verify-finding Re-check a single finding on demand (admin scope, CSRF)
POST /api/v1/quarantine-restore Restore quarantined file
POST /api/v1/quarantine/bulk-delete Bulk-delete quarantined files; returns count and the ids it could not delete
POST /api/v1/db-object-backup-restore Restore a dropped MySQL object from its db_object_backups record
POST /api/v1/test-alert Send test alert through all channels
POST /api/v1/import Import state bundle (suppressions, whitelist)
fix and fix-bulk act on the file the stored finding names. A request may
repeat that path in file_path, but a different path is refused, and a target
is never a remediation root itself (/home, /tmp, /var/tmp, /dev/shm)
or an account’s home directory.
verify-finding returns the verifier verdict in checked, resolved,
demote, and detail. When that verdict also changes the stored finding,
severity_change is demoted or restored. A demote verdict without
severity_change means no stored severity changed; for example, the finding
was already demoted or a scan replaced its snapshot while verification was
running. Callers must not report a state change from the verdict alone.
Settings
GET /api/v1/settings List editable config sections
GET /api/v1/settings/<section> Read a config section (secrets redacted)
POST /api/v1/settings/<section> Update a config section (safe fields reload, restart fields queue)
POST /api/v1/settings/restart Request a daemon restart. Returns 202 with `started_at_token`
for polling until the restarted daemon reports a new marker.
POST /api/v1/settings/firewall/tentative-apply Save firewall config, restart, and arm rollback timer
GET /api/v1/settings/firewall/rollback Read pending rollback state
POST /api/v1/settings/firewall/confirm Confirm tentative firewall changes
POST /api/v1/settings/firewall/revert Revert tentative firewall changes now
Sections map to top-level config keys: alerts, auto_response, challenge, reputation, performance, infra_ips, sentry, etc. Writes persist to csm.yaml, re-sign the integrity hash, and hot-reload where possible; restart-required changes are queued for /api/v1/settings/restart. Invalid field values return 422 and do not touch disk.
Fields marked file_only in the schema are shown but refused on write with a 422: anything that names a command, an executable, a file path, a socket or an environment variable the daemon acts on as root can only be changed in csm.yaml. Changing reputation.upstream.url or reputation.rspamd.url also returns 422 unless the same request enters the effective credential for that address again, and is refused outright when the credential comes from an environment variable. A credential entered in the request must match the resulting configuration after conf.d merging; an overridden replacement cannot authorize sending the stored credential to another address. Firewall tentative apply is restart-class by design: it snapshots the previous config, writes the new one, restarts the daemon, and auto-reverts unless the operator confirms before the timer expires.
Operator preferences
Per-operator state (UI density, timestamp display, default auto-refresh,
saved filter views) is keyed server-side by SHA-256 of the auth token,
so preferences follow the operator across browsers and devices without
the daemon ever storing the raw credential. Capability flag:
webui.prefs.v1. These endpoints require admin scope because they read
or mutate operator-private UI state.
GET /api/v1/prefs/user Read this operator's UI preferences
PUT /api/v1/prefs/user Replace the prefs blob (CSRF on cookie sessions)
GET /api/v1/prefs/views List saved views; `?page=findings` filters by page
PUT /api/v1/prefs/views Upsert one view {page, name, params} (CSRF on cookie sessions)
DELETE /api/v1/prefs/views Delete one view {page, name} (CSRF on cookie sessions)
Response shape for GET /api/v1/prefs/user:
{
"density": "comfortable",
"timezone": "local",
"auto_refresh": "on",
"table_columns": { "findings-table": ["check","severity","when"] }
}
density is comfortable or compact. timezone is server, local,
or an IANA-shaped zone string (e.g. Europe/Bucharest). auto_refresh
is on or off. Server-side sanitisation drops any other value. Unset
prefs encode as empty strings; the UI applies comfortable, local, and
on defaults.
Response shape for GET /api/v1/prefs/views:
{
"items": [
{
"name": "Critical SSH",
"page": "findings",
"params": { "severity": "critical", "check": "smtp_bruteforce" },
"updated": "2026-05-25T21:47:35Z"
}
],
"total": 1
}
Saved views are operator-scoped and capped at 200 per operator. The saved
view collection is stored as one 64 KiB preference blob. page and
params keys must be simple identifiers: ASCII letters, digits,
underscore, hyphen, or dot, up to 64 bytes. Each view has at most 32
params, and param string values are capped at 256 bytes. name must be
1-80 bytes with no control characters. PUT and DELETE return
"ok": true on success.
Bulk-action undo
Finding dismissals, bulk threat block / whitelist and bulk firewall unblock responses return
an undo_token when the daemon queues an inverse operation server-side
for 30 seconds. The UI surfaces a banner with the same TTL; CLI callers
can act on the token through the endpoints below. Each successful undo
writes an undo_<original_action> audit entry. Capability flag:
webui.undo.v1. These endpoints require admin scope because they read
or mutate operator-private action state.
GET /api/v1/undo/pending Latest pending undo entry for this operator (empty object if none)
POST /api/v1/undo/run Consume an entry and dispatch its inverse {id}; empty id uses latest
Non-empty response shape for GET /api/v1/undo/pending:
{
"id": "188d1f2a6c8b0000",
"action": "threat_bulk_block",
"inverse": "threat_bulk_unblock",
"summary": "Blocked 2 IPs",
"recorded_at": "2026-05-26T00:07:09Z",
"expires_at": "2026-05-26T00:07:39Z"
}
POST /api/v1/undo/run returns {ok, action, inverse, count} on
success, or 410 Gone when the entry is missing, already consumed, or
past its 30-second TTL. Recognised inverse action keys are
threat_bulk_unblock, threat_bulk_block, threat_bulk_unwhitelist,
threat_bulk_whitelist, firewall_bulk_reblock, and finding_undismiss. Other bulk actions
(quarantine delete, generic fix) do not surface an undo token because
they have no clean inverse.
Dismissals accept at most 500 keys per request and deduplicate repeated keys.
Undo skips findings changed by a later dismissal, successful re-check or baseline
reset, and count reports only the dismissals it actually reversed. Newer scan
copies remain in the current findings list.
Finding fields
Every finding in /api/v1/findings, /api/v1/events, and the JSONL audit log carries optional correlation fields when CSM can attribute them:
| Field | Meaning |
|---|---|
tenant_id | Tenant attribution from the verdict callback or panel-side webhook reply |
domain | Domain associated with the event (e.g. PHP-relay scriptKey host, mailbox domain) |
mailbox | Mailbox attribution (e.g. mail brute-force target, PHP-relay envelope-from) |
relay_total | PHP-relay trigger count for the path that fired |
relay_breakdown | PHP-relay script samples that contributed to the alert, with script key, hit count, last seen time, and a bounded sample subject when available |
Fields are omitted when the daemon could not attribute them. Orchestrators should treat absence as “unknown,” not “global.”
Cleanup fields
GET /api/v1/quarantine also powers the Cleanup page’s file-backup list. Entries include:
| Field | Meaning |
|---|---|
kind | quarantine or pre_clean |
live_state | original_missing, live_differs, original_not_file, archive_missing, archive_not_file, or unknown. Byte-identical restored entries are hidden. |
GET /api/v1/db-object-backups returns restored and restored_at when a captured MySQL trigger/event/procedure/function backup has already been replayed.
Incidents
GET /api/v1/incidents/groups is a read-scope rollup of active incidents by kind and source. It accepts status=active|all|open|contained|resolved|dismissed, kind, limit and offset, allowing a credential spray to render as one row per attacker rather than one row per target. total counts the groups and scanned_incidents the incidents they came from; truncated is true when the scan cap left incidents out.
GET /api/v1/incidents
Returns one page of incidents sorted by updated_at descending, with
items, total, offset, limit and status. limit defaults to 50
and is at most 500. status filters by open, contained, resolved,
dismissed, or active for open and contained; empty means all.
GET /api/v1/incidents/<id>
Returns one incident by id. 404 if not found.
POST /api/v1/incidents/<id>/status
Body:
{"status": "resolved", "details": "operator-marked"}
Status values: open, contained, resolved, dismissed. Closing an
incident (resolved/dismissed) means future findings for the same
correlation key start a fresh incident. Reopening an incident binds the
same key again. Incident JSON includes correlation_key when CSM has a
stored account, mailbox, domain, process, or remote-IP key.
Metrics (Prometheus)
CSM exposes a /metrics endpoint on its HTTPS web UI port
(default 9443). The endpoint serves the Prometheus text exposition
format (Content-Type: text/plain; version=0.0.4) and is safe to
scrape every 15 seconds.
“Available metrics” below is the shipped set. New call sites are
instrumented in ongoing releases; check CHANGELOG.md under
## [Unreleased] for the latest additions.
Enabling
Metrics are on whenever webui.enabled: true is set in csm.yaml.
The endpoint has its own auth knob:
webui:
enabled: true
auth_token: "<UI login token>"
metrics_token: "<long random string for Prometheus scraper>"
metrics_token is optional. When set, a Bearer header containing
this exact value grants access to /metrics. The UI auth_token or a valid
UI session cookie is also accepted so the dashboard can self-scrape,
but keeping the two tokens separate is recommended: rotating
auth_token does not then break Prometheus scraping, and giving
your monitoring stack the scrape token does not also give it UI
access.
Prometheus scrape config
scrape_configs:
- job_name: csm
scheme: https
tls_config:
# CSM serves a self-signed cert by default; either skip
# verification here or pin the CA you chose.
insecure_skip_verify: true
authorization:
type: Bearer
credentials: "<metrics_token from csm.yaml>"
static_configs:
- targets:
- csm-host-1.example.internal:9443
- csm-host-2.example.internal:9443
A complete, validated version of this snippet (with global: block)
ships as docs/src/examples/prometheus-scrape.yml. The CI pipeline
runs promtool check config against that file in the promtool-check
job; if the example ever stops validating, the pipeline fails.
Quick check
curl -sk -H "Authorization: Bearer $METRICS_TOKEN" \
https://localhost:9443/metrics | head
Available metrics
Build / process
csm_build_info{version}(gauge, always 1): build metadata. Scrape once to discover the running version. Join on it in queries viagroup_left(version).
Go runtime
Always exposed. Use these to watch daemon memory over time and tell live
growth from GC headroom (rising heap_alloc = real growth; large
heap_idle = retained-but-freed memory).
go_memstats_heap_alloc_bytes(gauge): heap bytes allocated and still in use (live).go_memstats_heap_inuse_bytes/go_memstats_heap_idle_bytes/go_memstats_heap_released_bytes(gauges): in-use spans, idle spans, and bytes returned to the OS.go_memstats_heap_sys_bytes/go_memstats_sys_bytes(gauges): heap and total bytes obtained from the OS.go_memstats_next_gc_bytes(gauge): heap-size target for the next GC.go_memstats_gc_cpu_fraction(gauge): fraction of CPU time spent in GC since start.go_goroutines(gauge): live goroutine count.
For a deeper leak hunt, enable the loopback pprof endpoint (debug.pprof_listen)
and run go tool pprof http://127.0.0.1:<port>/debug/pprof/heap over an SSH tunnel.
The same endpoint serves profile (CPU, 30s by default), goroutine, mutex
and block. Contention sampling for the last two starts with the listener and
stops when the last listener exits, including on a listener failure. Failed or
refused binds do not enable sampling.
While enabled, mutex profiling samples one in 100 contention events. Block
profiling samples waits of at least 10 microseconds and a proportion of shorter
waits. Timing and collecting these samples adds CPU overhead even when nobody
fetches a profile, so use this endpoint for diagnostics and leave
debug.pprof_listen empty when it is not needed. Profiles accumulate for the
process lifetime; disabling sampling does not erase previously recorded data.
For a CPU investigation, collect the CPU profile first and the goroutine dump alongside it:
curl -o cpu.pb.gz 'http://127.0.0.1:<port>/debug/pprof/profile?seconds=60'
curl -o goroutine.txt 'http://127.0.0.1:<port>/debug/pprof/goroutine?debug=2'
YARA-X worker (default-on; off only if signatures.yara_worker_enabled: false)
csm_yara_worker_restarts_total(counter): cumulative number of times the supervisor has restarted thecsm yara-workerchild. Alert on sustained growth: a single restart is routine (rule deploys), a steady climb means the worker is crash-looping and real-time YARA scans are degraded.
Findings
csm_findings_total{severity}(counter): every finding CSM records is counted here. Severities areCRITICAL,HIGH, andWARNING(matching thealert.Severityenum). Userate(...)for arrival velocity; watch for sudden CRITICAL spikes.
Alert delivery
-
csm_alert_dispatch_failures_total(counter): alert channel sends that failed after CSM detected findings. Counts email, webhook, and phpanel delivery failures. Sustained growth means findings are not reaching operators; check SMTP, webhook reachability, and credentials. -
csm_audit_sink_degraded{sink}(gauge): one when an enabled JSONL or syslog destination failed to open or write, zero when healthy or disabled. Values appear after the audit pipeline first runs and remain degraded during retry backoff. Syslog reports local write results, not receiver acknowledgement. -
csm_audit_events_dropped_total{sink}(counter): events whose destination was unavailable or whose write failed. Usecsm export --since <when>for backfill; failed live deliveries are not replayed automatically.
State
csm_store_size_bytes(gauge): on-disk size of the bbolt state database (/var/lib/csm/state/csm.dbby default). Retention sweeps bound logical growth. Startup compaction reclaims freelisted pages automatically when the file is large and mostly slack;csm store compactdoes the same immediately with the daemon stopped. The retention sweep logs a restart hint only under that same condition. It includes deleted pages still held by active readers, which become reclaimable when the daemon stops. Later writes can change whether compaction is due at startup. History and attack-event writes pack pages densely when their timestamps arrive mostly in order. A write uses balanced page splits when events older than the newest stored one are most of its records or at least half its inserted data. Space savings depend on record sizes and arrival order. Archive imports keep the snapshot’s page layout; compaction repacks an existing file.
Fanotify realtime monitor
csm_fanotify_queue_depth(gauge): current number of queued events waiting for the analyzer pool. The queue capacity is 4000; sustained values near that cap mean drops are imminent. Alert target:max_over_time(csm_fanotify_queue_depth[5m]) > 3500.csm_fanotify_events_total(counter): delivered fanotify file events whose admission decision has completed. Filesystem marks receive writes anywhere on their superblocks, including other bind mounts. Kernel queue overflow notifications are counted separately.csm_fanotify_events_admitted_total(counter): events queued for analysis, including dropper tracking. Subtract this counter andcsm_fanotify_events_dropped_totalfromcsm_fanotify_events_totalto count events rejected before queue admission. This includes path-resolution failures as well as filter rejections. Use rates over the same interval to compare these independently scraped counters.csm_fanotify_events_dropped_total(counter): cumulative events dropped because the analyzer queue was full. The reconcile pass still rescans drop-affected directories 60 s later, so dropped events do not disappear from detection – they arrive delayed. Alert target:rate(csm_fanotify_events_dropped_total[5m]) > 0paired with a short for-clause.csm_fanotify_kernel_queue_overflow_total(counter): kernel-side fanotify queue overflows (FAN_Q_OVERFLOW). Unlike analyzer-queue drops, the kernel gives no file descriptor for lost events, so only directories with concurrent analyzer drops can be reconciled; the rest are covered by the next scheduled deep scan. Any nonzero rate during a write storm means realtime coverage had a gap.csm_spool_fanotify_queue_overflow_total(counter): kernel queue overflows on the mail spool watcher. In permission mode an overflowed event means the message was delivered without being scanned; a Warning finding is raised when this happens.csm_fanotify_reconcile_latency_seconds(histogram): how long the post-overflow reconcile pass takes to walk drop-affected directories and rescan recent files. Buckets: 0.01 s .. 60 s. Watch p95: reconcile stealing tens of seconds means bulk events are piling up faster than the walker can keep up.csm_checks_domlog_discovery_dropped_total{reason}(counter): per-vhost access-log paths the WP brute-force domlog discovery helper dropped before scanning. Labels:reasonisevalsymlinks_error(broken symlink, attacker-removed log file) orstat_error(file vanished between glob and stat, permission regression on the log directory). Steady growth means a real chunk of vhosts is being silently skipped each cycle. Stale-mtime drops are intentional filtering and are NOT counted here.csm_realtime_content_scan_truncated_total{check}(counter): cumulative real-time content checks where the underlying file was larger than the main read window, so the full-rule pass saw only the leading window. The read cap protects RE2 cost on huge files; sustained growth on a label means full-rule coverage is capped on large files. Labels currently emitted:phpcontent_inline(known webshell filename),phpcontent_uploads(PHP in uploads),php_check(generic PHP content scan),crontab(per-user /var/spool/cron write),htaccess(per-vhost .htaccess write),user_ini(per-vhost .user.ini write),html_phishing(HTML phishing heuristic), andcgi_backdoor(CGI backdoor heuristic). Compare against finding history forwebshell_content_realtime(or the matching check name for non-PHP labels) to judge whether a raised cap would surface real findings.
Periodic check runner
-
csm_checks_crontab_base64_truncated_total(counter): crontab base64 candidates that exceeded the per-blob decode cap before decoded-content pattern matching ran. Sustained growth means encoded cron content is larger than the scanner currently inspects; review affected crontabs and tune the scanner before redeploying. -
csm_check_duration_seconds{name,tier}(histogram): wall-clock time each check takes to complete. Labelnameis one of the periodic check-runner names (fake_kernel_threads,webshells, …); labeltieriscritical,deep, orall. Buckets: 0.01 s .. 900 s. Most checks keep the 300 s timeout ceiling; heavy filesystem checks can run up to 900 s. Useful aggregations:# p95 of the slowest check in the critical tier: histogram_quantile(0.95, sum by (le, name) ( rate(csm_check_duration_seconds_bucket{tier="critical"}[10m]) ) ) # total time each cycle spends in deep-tier checks: sum by (tier) (rate(csm_check_duration_seconds_sum{tier="deep"}[1h]))
Threat intelligence
Registered when reputation.upstream.enabled: true.
csm_threatintel_cache_hits_total(counter): upstream threat-intel lookups served from CSM’s local per-IP cache.csm_threatintel_cache_misses_total(counter): upstream threat-intel lookups not served from the local cache. A miss may still fail open without an HTTP request when the breaker is open.csm_threatintel_backend_failures_total(counter): upstream backend failures from network errors, non-200 responses, malformed JSON, response IP mismatches, or out-of-range scores.csm_threatintel_breaker_open(gauge): 1 while the upstream circuit breaker is refusing calls, 0 when closed or allowing its single cooldown probe.
AI-crawler IP ranges
csm_botranges_refresh_total{result}(counter): AI-crawler IP-range refresh attempts, labelledsuccess(at least one vendor feed updated) orfailure. Driven by the daemon’s auto-updater and bycsm update-bot-ranges.csm_botranges_prefixes{bot}(gauge): current number of published IP prefixes per crawler identity in the active overlay.csm_botranges_last_success_timestamp_seconds(gauge): Unix time of the last successful refresh.
Firewall
csm_blocked_ips_total(gauge): number of IPs currently on the firewall block list. Excludes expired temp bans – the store’sLoadFirewallStatefilters those before the gauge reads.csm_firewall_rules_total(gauge): total firewall rules across all four categories (blocked IPs, allowed IPs, blocked subnets, port-specific allows). Excludes expired temp blocks and allow-list rows. Sudden drops are worth investigating; expected drops happen when temporary block or allow deadlines pass.
Config reloads
csm_config_reloads_total{result}(counter): SIGHUP reload attempts, by outcome. Labels:resultis one of:success– safe fields swapped in place, integrity hash re-signed, live config updated.restart_required– one or more fields that need a full restart changed; live config unchanged.error– YAML parse failure, validation failure, or re-sign failure; live config unchanged.noop– file edit produced no semantic change (identical values, whitespace edit, etc.). Alert target:rate(csm_config_reloads_total{result="error"}[5m]) > 0paired with a short for-clause.
Auto-response
csm_auto_response_actions_total{action}(counter): every auto-response action fired, by class. Labels:actioniskill,quarantine, orblock. Incremented once per finding the correspondingAuto*helper produces, so a batch blocking four IPs in one cycle adds 4 toaction=block. Useful for detecting response storms:rate(csm_auto_response_actions_total[5m]).
Challenge
csm_challenge_routed_total{check}(counter): IPs routed to the proof-of-work challenge, labelled by the sourcecheckthat flagged them (e.g.http_scanner_profile,wp_login_bruteforce,ip_reputation). Graph per-detector challenge volume withsum by (check) (rate(csm_challenge_routed_total[5m])).csm_challenge_escalated_total{outcome}(counter): challenge entries that timed out and were escalated to a hard block, by firewalloutcome:live(a new block landed),noop(the IP was already blocked), plusdry_run,allowed, andallowlisted.outcome=liveis how many challenges actually became blocks.
Firewall blocks
csm_firewall_block_outcome_total{outcome,source}(counter): every IP block attempt, labelled by firewalloutcome(live,dry_run,allowed,allowlisted,noop,protectedfor refused infrastructure IPs,errorfor engine failures or an unavailable engine) and bysource(scan,challenge,incident,central_intel,cli,web_ui, orunknownfor an invalid internal source). All combinations are exported at zero from startup, and unexpected outcome or source values collapse into the fixederrororunknownbuckets instead of creating new labels. This is the aliveness signal for auto-response: alert whensum(rate(csm_firewall_block_outcome_total{outcome="live",source!~"cli|web_ui"}[1h])) == 0while attack findings keep coming, and watchoutcome="error"for a dying engine.
Retention (when retention.enabled: true)
csm_retention_sweeps_total(counter): number of retention sweep cycles completed since daemon start. A flatline after a restart means the sweep goroutine is not scheduling; a healthy daemon increments this on everysweep_intervaltick.csm_retention_deleted_total(counter): cumulative entries deleted across thehistory,attacks:events, andreputationbuckets. Spikes on the first sweep after enabling retention (initial backlog), then settles to the steady-state churn. Useful for estimating when the file might benefit from a restart or acsm store compactmaintenance window.
PHP-relay (email abuse, cPanel only)
All series are prefixed csm_php_relay_. Registered when email_protection.php_relay.enabled: true and the host is cPanel; otherwise zero across the board. See Real-time detection.
csm_php_relay_findings_total{path}(counter): findings emitted per detection path. Labels:pathis one ofheader,volume,volume_account,fanout(and laterbaseline,reputationfor Stages 2-3). Userate(...)to spot detection storms; a sudden rate jump onheadertypically means a contact-form vulnerability is being exploited, onvolume_accounttypically means an account password was leaked.csm_php_relay_actions_total{action,result}(counter): auto-freeze invocations attempted. Labels:actionis currentlyfreeze;resultisokorfail. Pair withcsm_php_relay_findings_totalto confirm freeze keeps up with detection.csm_php_relay_action_gone_total(counter): messages already absent from the spool by the timeexim -Mfran. Normal queue churn; not a failure. Sustained growth means the spool is moving fast and the freezer is racing the queue runner.csm_php_relay_path_skipped_total{path,reason}(counter): path evaluation that bailed before producing a finding. Labels:pathmatches the finding labels above;reasonenumerates the gate that fired (e.g. ignore-list match, missing scriptKey).csm_php_relay_spool_scan_fallbacks_total{reason}(counter): AutoFreeze fell back to a full spool walk to find msgIDs. Labels:reasoniscapped(the in-memoryactiveMsgsper script hit its cap, so a fresh disk walk was needed) orreputation(a late reputation finding arrived for a script with no liveactiveMsgs). Sustained growth oncappedmeans a single script is firing faster than the in-memory window keeps state for; consider raisingheader_score_volume_minor adding an ignore.csm_php_relay_active_msgs_capped_total(counter): per-scriptactiveMsgsset hit its cap and dropped the oldest entry. Counts the eviction event itself; the next freeze for that script will land incsm_php_relay_spool_scan_fallbacks_total{reason="capped"}.csm_php_relay_windows_active{kind}(gauge): retained per-script / per-IP / per-account window state. Labels:kindisscript,ip, oraccount. Sized by Flow E sweep cadence (5 min for windows, 24 h retention for accounts); flat values across hours are normal.csm_php_relay_msgid_index_size{layer}(gauge): msgID dedup index size by storage layer. Labels:layerismemory(in-process map) orbbolt(persisted batch writer). Memory ceiling is 200k entries; bbolt grows freely until the 25 h Flow E sweep prunes it.csm_php_relay_msgindex_persist_dropped_total(counter): bbolt persist queue overflow drops (the 4096-deep buffered channel was full when the watcher tried to enqueue). Should be zero in steady state; a non-zero value means the bbolt writer is blocked on disk and the in-memory dedup is the only thing protecting against double-fire on a queue-runner re-write.csm_php_relay_msgindex_persist_errors_total(counter): bbolt commit failures from the async batch writer. Each bump also emits a Criticalemail_php_relay_msgindex_persist_failedfinding. Disk-full or permissions issue on/var/lib/csm/state/csm.db.csm_php_relay_inotify_overflows_total(counter): kernelIN_Q_OVERFLOWevents on the spool watcher. Each one triggers a bounded recovery scan (default cap 1000 files); if the cap fires, also emitsemail_php_relay_overflow_scan_truncatedCritical. Sustained growth means the spool is churning faster than inotify can keep up – usually a backup restore or a real attack.csm_php_relay_spool_read_errors_total(counter):emailspool.ParseHeaderserrors on-Hfiles the watcher tried to consume. Usually transient (file disappeared between inotify event and open) and self-correcting; sustained growth points at a permissions or filesystem problem.csm_php_relay_userdata_errors_total(counter):cpanelUserDomainsresolver errors reading/var/cpanel/userdata/. Used by the Path 1Frommismatch check; errors here mean Path 1 is potentially undercounting until the read recovers.
Signature retroactive rescans
csm_signature_rescans_total(counter): full deep-tier sweeps completed because a tracked signature file’s content changed. Re-installing identical rules, as a package upgrade or a repeated download does, does not count. Multiple updates before the next sweep can coalesce into one rescan.
The watcher compares content hashes when a file’s mtime or size changes; symbolic links use the target file’s metadata. Failed reads and files replaced during hashing are retried without discarding the last verified state. First observations and file removals do not queue a rescan.
State from older builds contains only mtimes. The first successful read adds a hash without queuing a rescan if the recorded mtime still matches. A moved mtime without a recorded hash conservatively queues one rescan.
Counter reset semantics
Prometheus counters in CSM live in process memory. They reset to zero
whenever the daemon restarts (config change, binary upgrade, crash
recovery). This is the standard behaviour for every
Prometheus-instrumented daemon; Prometheus’s scrape pipeline detects
counter resets on its own and rate(), increase(), and
rate_over_time() all handle them correctly.
Operators should not alert on “counter decreased across a scrape” as
a failure condition. Alert on rate() or increase() of a counter
over a window long enough to absorb expected restarts.
Persisting counters across restarts would require writing to bbolt on every increment, which would not pay for itself. If a specific metric needs restart-stable behaviour later, a gauge-over-the-bbolt-counter pattern can be added for that one case without affecting the rest.
Caveats
- Scrape the web UI’s HTTPS port, not a separate listener.
curl -k/insecure_skip_verifyis appropriate only when the cert is self-signed and the network path is trusted. Pin a CA for anything else.- Prometheus label cardinality: per-account and per-IP labels are deliberately not exposed. Shared-hosting deployments with 1000+ cPanel users would otherwise overwhelm a Prometheus server.
- Metric vectors cap label-value combinations at 1000 children per
metric, including the overflow bucket. Once a vector reaches that
cap, new combinations are aggregated under
_overflow_.
Not instrumented (yet)
- Per-account labels on any metric. Deliberately off: shared-hosting deployments with 1000+ cPanel users would blow out Prometheus cardinality.
- Fanotify inline auto-response actions (the quarantine-while-
seeing-the-write path in
fanotify.go). The periodiccsm_auto_response_actions_totaldoes not count those; a follow- up may split the metric or add asourcelabel. - bbolt per-bucket size breakdown,
csm_store_used_bytes, andcsm_store_last_compact_ts. Deferred to the online-compaction follow-up of the retention work (seeROADMAP.md).
Audit Log (SIEM)
Audit Log
CSM ships source observations and notification findings to one or more SIEM-friendly sinks, deduplicated by observation identity within each batch and against each destination’s recent successful deliveries. Audit records include sources suppressed by notification filtering or rate limits, so SIEM correlation can still identify the original observation.
Two sink types ship today, both opt-in via csm.yaml. They can be
enabled together or independently.
Schema
Every event, regardless of transport, has the same shape:
{
"v": 1,
"ts": "2026-04-28T10:32:14.512938Z",
"finding_id": "8e3f1c204c1d8b95",
"severity": "CRITICAL",
"check": "webshell_realtime",
"message": "PHP execution primitive in uploads/",
"details": "...",
"file_path": "/home/customer/public_html/uploads/x.php",
"hostname": "host.example.com"
}
The v field is the schema version. CSM bumps it on incompatible
changes and will not bump it for additive fields, so SIEM parsers
can pin on v: 1 and ignore unknown keys.
finding_id is a stable 16-hex-char hash of the canonical fields
(timestamp, check, severity, message, file path). Two emits of the
same finding produce the same ID, so downstream dedup works across
re-runs.
Firewall actions caused by a finding carry this same identity in the action log, including failed and refused attempts. The identity is captured before reason text is shortened and survives queued retries, challenge timeouts, central-intelligence decisions, and permanent-block promotion. Subnet and incident escalation link the latest known contributing finding. Older stored evidence without an identity and manual or maintenance operations remain unlinked; CSM does not reconstruct an identity from display text. Database session blocks link the original database finding, not a synthetic IP candidate.
The daemon audits source observations even when they repeat an earlier finding
or are filtered from operator notifications. Distinct observations keep their
own identities; recent replays of the same observation are suppressed per sink.
Receipts are kept in bounded memory and survive temporary sink failures. A
destination that missed a record can receive its replay without duplicating a
healthy destination’s record. Replays refresh their receipt’s position, keeping
recently replayed observations ahead of inactive receipts.
Receipt matching includes original details, scanner identity and source
attribution; distinct findings can share a legacy finding_id and still need
separate records.
Process enrichment alone does not create a new observation. Restarting or
reconfiguring sinks clears receipts; observations no longer in the receipt cache
can be emitted again. The legacy finding_id remains an action correlation key;
collectors must not treat it alone as a unique-record key.
Notification suppression and downstream finding observers keep their existing
behavior. Audit delivery still depends on the configured sink and scan findings
reaching the dispatcher.
The ts field records when CSM raised the finding, including process
monitoring, automatic actions, scanner health, and mail relay storage errors.
It is independent of when a sink delivers the event. A producer that builds
a finding without a time gets the moment the daemon received it, so ts is
never the zero time.
Alert dispatch fills missing times on its own copy of the findings. Reusing an unstamped input for a later occurrence gets a fresh time and finding id; replaying a finding with an existing time preserves its time and id.
Process context
Exec and outbound-connection findings on BPF-backed hosts carry an
optional process object with PID, PPID, UID, user, cPanel account
(when known), comm, exe, sanitized cmdline, and a parent chain. The
field is omitted when no context is available, so existing parsers
that ignore unknown keys see no schema change.
{
"severity": "HIGH",
"check": "outbound_connection",
"message": "Suspicious outbound connection",
"process": {
"pid": 4242,
"ppid": 4200,
"uid": 1001,
"user": "alice",
"account": "alice",
"comm": "ncat",
"exe": "/usr/bin/ncat",
"cmdline": ["ncat", "203.0.113.10", "587"],
"parent": {
"pid": 4200,
"ppid": 4100,
"uid": 1001,
"comm": "sh"
}
},
"timestamp": "2026-05-07T12:34:56Z"
}
The parent chain may be truncated at depth 5 and may stop early if an intermediate parent has been evicted from the cache.
File sink (JSONL)
alerts:
audit_log:
file:
enabled: true
path: /var/log/csm/audit.jsonl # default
The default path is created with mode 0640 and the parent dir
with 0750. The packaged logrotate fragment uses copytruncate
mode so the daemon’s open file descriptor stays valid across
rotation – no SIGHUP needed. It rotates daily and keeps 14 compressed
rotations. The 100 MB threshold permits early rotation when the host runs
logrotate more often than daily; it does not cap growth between runs. The sink
never drops records to stay under that size, because a detection flood would
otherwise blind the audit trail until the next rotation. On busy hosts, run
logrotate more often than daily so the size trigger can take effect.
Installation and upgrades refresh the fragment, including upgrades through
csm rehash.
If you move the audit log off the default path, add your own logrotate stanza for it: the packaged fragment names the default path only.
Tail it for an interactive view:
tail -F /var/log/csm/audit.jsonl | jq -c
Or hand it to a log shipper like Vector, Filebeat, or Fluentbit.
Syslog sink (RFC 5424)
alerts:
audit_log:
syslog:
enabled: true
network: udp # udp | tcp | unix | unixgram | tls
address: 127.0.0.1:514 # host:port for udp/tcp/tls, path for unix*
facility: local0 # default
tls_ca: "" # optional PEM file for tls transport
Wire-line is RFC 5424 with the JSON event embedded as the MSG body, so receivers that already understand the JSONL schema parse it the same way regardless of transport. UDP and unix-datagram emit one datagram per message; TCP, TLS, and unix-stream use LF framing.
Severity mapping onto the standard syslog level set:
| CSM severity | Syslog level | Numeric |
|---|---|---|
| CRITICAL | crit | 2 |
| HIGH | err | 3 |
| WARNING | warning | 4 |
Automated tests cover RFC 5424 output and UDP, TCP, TLS, Unix datagram, and Unix stream framing. Validate the chosen receiver configuration in a staging environment before production rollout.
Delivery failures and recovery
Each enabled destination retries independently after an open or write failure. The delay starts at one second and doubles up to one minute. A later dispatch retries a missing destination once its delay has elapsed; a quiet daemon waits until another finding arrives. Healthy destinations stay open during retries. Reloading audit configuration waits for in-flight writes before replacing sinks.
Failures and recovery are logged to journald. Monitor
csm_audit_sink_degraded{sink="jsonl"} and
csm_audit_sink_degraded{sink="syslog"} on /metrics: one means a configured
destination failed, zero means it is healthy or disabled. These values are
initialized when the audit pipeline first runs. Syslog health reflects local
connection/write results, not an acknowledgement from the receiving SIEM.
csm_audit_events_dropped_total{sink} counts events whose destination was
unavailable or whose write failed. Failed deliveries are not replayed
automatically; use the backfill command below to recover stored findings.
Backoff resets after a successful delivery, and configuration changes take
effect on the next dispatch.
Backfill
When you first turn on the audit log, the SIEM has no history. Use
csm export --since <when> to dump prior findings in the same JSONL
schema:
csm export --since 24h > recent.jsonl
csm export --since 2026-04-01T00:00:00Z > q2.jsonl
<when> is either an RFC 3339 timestamp or a duration relative to
now (24h, 7d). The output is one JSON event per line on stdout,
identical in shape to what the live sinks emit, so you can pipe it
straight into the same ingest pipeline.
Requires a running daemon.
What gets logged
Source observations and notification findings reaching the audit dispatcher, deduplicated by observation identity within each batch, before:
- the per-account rate limiter (so audit signal is not lost when email and webhook are throttled);
- the “blocked IP suppression” filter (so SIEM correlation sees events that operators were spared);
- the per-sink disabled-checks list (audit log is not subject to
email’s
disabled_checks).
This means audit-log volume is generally higher than the email or webhook stream. Plan SIEM retention accordingly.
What does not get logged
Before a record is written, its message and details replace recognized
password fields, API tokens, command-line secrets and cPanel session
identifiers with [REDACTED]. Session redaction covers cPanel, WHM,
Webmail, the shared server daemon, DAV and security purge log lines. It
keeps the account name beside a session identifier, leaves unrelated
lines of a multiline finding alone, and leaves already redacted text
unchanged when it runs again. Repeated and quoted values are all
covered, including the displayed form of NUL-separated arguments.
The same redaction runs on email digests, on new finding history in both the bbolt and the legacy JSONL backend, and therefore on the history the web UI serves and exports as CSV. Attack event messages are redacted before truncation, so a truncated line cannot hide a credential behind a cut service tag; attack events store no finding details. Account and IP attribution is read from the original finding, and finding IDs are computed from it too, so audit records still correlate with remediation records. Other structured fields are copied unchanged.
Two limits are worth knowing. Records written by earlier versions are not rewritten, so an existing log keeps whatever it already holds. The active-finding snapshot and the pending queues are process-local state that is read back and compared by the daemon itself, so they are not redacted.
The audit log is not a replacement for csm.history (the bbolt
history bucket). Only findings that pass through the audit dispatcher
are emitted. Internal state changes – daemon startup, reload events,
config changes – live in journald via csm.service and are not
mirrored here.
Action log
The audit log records what CSM found. The action log records what CSM did: one record per action, one file, one schema, whether the action came from the daemon, an operator command or the web UI.
/var/log/csm/actions.jsonl
csm actions # recent actions, one line each
csm actions --since 24h # RFC 3339 timestamp or a duration
csm actions --op respond.block_ip
csm actions --json # the raw records
Schema
{
"v": 1,
"ts": "2026-09-09T10:32:14.512938Z",
"hostname": "host.example.com",
"op": "respond.quarantine_file",
"actor": "daemon",
"finding_id": "8e3f1c204c1d8b95",
"target": "/home/alice/public_html/uploads/x.php",
"account": "alice",
"reason": "webshell_realtime",
"before": {"exists": true, "sha256": "5f2b...", "size": 1841, "mode": "-rw-r--r--", "uid": 1001, "gid": 1001},
"after": {"exists": false},
"result": "applied",
"recovery_path": "/opt/csm/quarantine/2026-09-09T10-32-14-x.php"
}
| Field | Meaning |
|---|---|
op | The privileged operation, using the IDs from the capability matrix. The matrix says what an operation may do; this says what it did. |
action | The specific change within an operation, such as block, unblock, apply or rollback. |
actor | daemon, cli or webui. actor_detail carries the operator’s source address or the command name. |
finding_id | The finding that caused the action, using the same ID the audit log emits, so the two streams join. |
command | The exact argv when CSM ran a program. Absent when the action used system calls only. |
before / after | The target file’s digest, size, mode and owner around the change. "exists": false after a quarantine is the record that the file is gone. |
result | applied, dry_run, failed or refused. A dry-run record says what CSM would have done. |
undo | A command that reverses the action, when an exact inverse is available. |
recovery_path | Retained quarantine content or a pre-clean backup, with a .meta sidecar. Use the quarantine recovery workflow; csm restore accepts state archives, not these files. |
Attempted operations include failures and safety refusals. An already blocked address does not produce another applied record. A successful file replacement or quarantine with a later warning stays visible with its recovery path.
A refused action is recorded too. “The safety rules stopped this” and “CSM never looked” are different operational states, and only one of them is a reason to change the configuration.
v is the schema version. It is bumped only on an incompatible change, so a
parser can pin on v: 1 and ignore unknown keys.
Long reasons and error messages are truncated with an explicit marker so one action cannot make the history unreadable. Targets and command arguments are preserved.
A file larger than 64 MiB is recorded without a digest rather than stalling the action while it is hashed. Quarantine digests come from the captured copy, after the move. Cleaning digests describe the exact bytes read and installed through the pinned file handles. Directories and symlinks have no digest.
What is covered
The capability matrix has an “Action record” column. Six operations write here today: firewall blocks and unblocks (automatic and operator), whole-ruleset changes, file quarantine, surgical file cleaning, and process termination. Everything else appears in the daemon log only. The column is the authoritative list. Tests exercise the operation paths, including failures and concurrent blocks.
The firewall keeps its own audit file because the web UI and the API read it. Its entries also appear here; dry-run decisions, failed attempts and whole ruleset apply/rollback records are recorded directly on this stream. Startup rollback records are written before the daemon requests a restart.
actor uses the caller’s attribution when available and otherwise identifies
the process performing the action. Some web UI requests therefore record as
daemon; the web UI’s own action log retains the operator’s source address.
Rotation
The file rotates to actions.jsonl.1 at 10 MB. csm actions reads both.
Readers and writers coordinate rotation through actions.jsonl.lock.
Readers pin both files before releasing the lock, so rotation cannot hide
records and a slow reader does not hold up a writer.
CSM handles this rotation itself; the installed logrotate configuration leaves
the live file alone and archives actions.jsonl.1, which the sink never writes
to again, for 90 days. Older history is compressed there and is not read by
csm actions; ship the stream to a SIEM if you need it queryable.
Recording is best effort. Sink errors or panics do not change an action’s outcome. A write waits at most 250 ms, with at most 64 writes outstanding; a stalled sink or saturation can lose records. Normal CLI writes finish before the command exits. Keep this log on a local filesystem and monitor write-error warnings. The stream is not a transactional guarantee that every host change survives a crash.
Building & Testing
Toolchain and prerequisites
Use the Go version required by go.mod (currently 1.27.1), including its
formatter. A newer local formatter can disagree with the pinned CI linter.
For an installed Go launcher that supports toolchain selection:
export GOTOOLCHAIN=go1.27.1
export PATH="$(go env GOROOT)/bin:$PATH"
go version
Linux tests need PHP CLI and python3-cryptography for the shipped PHP runtime
and release verification regressions. The release-verifier tests skip when the
module is absent and fail when CSM_REQUIRE_PYTHON_VERIFIER=1, which the CI
test jobs set so a missing module cannot pass as a silent skip. The Web UI
JavaScript tests need Node 24 or newer the same way: they skip without it and
fail when CSM_REQUIRE_NODE=1, which the CI test job sets. Builds with yara,journal,bpf also need
CGO, pkg-config, YARA-X 1.20.0 and the systemd
development library. Use the release builder or the documented test images.
CI selects the module toolchain with GOTOOLCHAIN=auto; the older Go versions
in its bootstrap images are not the module requirement.
Build
# Standard build (no YARA-X)
go build ./cmd/csm/
# Build with YARA-X support (requires libyara_x_capi)
CGO_LDFLAGS="$(pkg-config --libs --static yara_x_capi)" go build -tags yara ./cmd/csm/
Test
The default CI job runs every package with the race detector, no extra build tags, and a 30-minute per-package timeout. Run its exact test command on Linux:
go test -v -race -timeout=30m -covermode=atomic -coverprofile=coverage.out -coverpkg=./internal/... ./...
make test and the test step of make ci use -short for local iteration;
they do not reproduce the full CI suite. Neither default command tests the
shipped optional backends. The additional required production job runs:
scripts/production-tests.sh portable
That script selects all packages with yara,journal,bpf, -race, -count=1,
-p=2 and -timeout=30m, retains JSON execution evidence, and checks the test
inventory. See production and kernel tests for the
builder, tagged lint/security commands and required kernel execution.
On macOS, use scripts/go-linux.sh for Linux code. Build the repository’s
PHP-enabled test image once with an available container builder:
docker build -f build/Dockerfile.systemd-test -t csm-linux-test .
GO_LINUX_RUNTIME=docker GO_LINUX_IMAGE=csm-linux-test scripts/go-linux.sh \
bash -ec 'apt-get update -qq && apt-get install -y --no-install-recommends python3-cryptography
go test -v -race -timeout=30m -covermode=atomic -coverprofile=coverage.out -coverpkg=./internal/... ./...'
The wrapper derives the default Go version from go.mod, shares persistent
caches across worktrees, and grants fanotify/nftables capabilities. Its plain
Go image does not include PHP, python3-cryptography or the production CGO
libraries. The command above installs Python’s verifier in the disposable
test container. An image built
with another runtime can be selected through GO_LINUX_RUNTIME and
GO_LINUX_IMAGE; local test execution still goes through the wrapper.
Kernel firewall regressions use isolated network namespaces:
scripts/go-linux.sh go test -tags nftkernel ./internal/firewall -race -count=1
The separate kernel gate requires an isolated Linux runner with the documented boot capabilities. A successful macOS or default-tag suite does not demonstrate BPF LSM attachment.
Fuzz
CSM has fuzz targets for parsers that read attacker-controlled input, including Exim mainlog lines, Dovecot maillog lines, Apache Combined Log Format, /proc/net/tcp rows, wp-config.php bodies, /etc/shadow, auditd comm fields, and finding messages coming back from the WebUI.
Fuzz targets live in *fuzz*test.go files across the internal packages and scripts. Fuzz targets do two things:
- Their seed corpus runs as part of the normal test suite.
go test ./...executes every seed, so a known-bad input stays a regression test forever. - The actual fuzzer runs with
-fuzz=FuzzFoo.
Run a target for a fixed time while investigating:
go test ./internal/checks/... -run=^$ -fuzz=^FuzzExtractPHPDefine$ -fuzztime=30s
Run only the seeds:
go test -run=Fuzz ./...
If the fuzzer finds a crasher it writes the failing input to testdata/fuzz/FuzzFoo/<hash>. Commit that file alongside the fix and the input becomes a permanent seed.
Adding a fuzz target:
func FuzzMyParser(f *testing.F) {
// Seeds: real-world valid shape, empty, malformed.
f.Add("valid input")
f.Add("")
f.Add("corrupt/truncated")
f.Fuzz(func(t *testing.T, s string) {
_ = myParser(s) // must not panic on any input
})
}
Keep the target tight: call one function, assert it returns. Output verification belongs in a regular test.
Lint
make lint # must pass before push
make fmt-check # checks tracked Go files with the pinned formatter
make lint uses repo-local cache directories under .cache/ and a five-minute
timeout, matching .golangci.yml and CI. Install the pinned tools with
make tools; golangci-lint is 2.11.4. It lints the Linux build, so a macOS
host checks the code that ships rather than reporting its linux-only callers
as unused. Production-tag lint needs the Linux CGO libraries described in
production tests.
make sec, make vuln, and make check-fixtures are the local security,
vulnerability and fixture checks. For an aggregate local check use make ci,
then run the full default and production commands above. Passing local checks
does not replace the required kernel and cloud jobs.
Linter config in .golangci.yml: errcheck, govet, staticcheck, unused, ineffassign, gocritic, misspell, bodyclose, nilerr.
CI/CD
GitLab CI (.gitlab-ci.yml) is the internal build pipeline. It runs lint/test/package jobs, publishes internal packages, mirrors to GitHub, and creates the public GitHub release artifacts.
| Stage | What it does |
|---|---|
| .pre | Release preflight rejects version tags without a usable cPanel image. |
| lint | Pinned golangci-lint and formatter, vet, blocking gosec/govulncheck, fixture privacy, and Prometheus config validation. |
| test | Full default-tag race/coverage suite (30-minute package timeout), shipped-tag lint/security/race gate, required kernel/service gate, and four pinned clean-application gates. |
| build-image | Build CSM builder Docker image with YARA-X (manual trigger) |
| build | amd64 and arm64 release binaries with YARA-X CGO and the yara journal bpf build tags. arm64 builds use QEMU/buildx. |
| package | RPM + DEB via nFPM |
| integration | Spin up cloudv-1 AlmaLinux and Ubuntu hosts plus the configured clean cPanel image via phctl, install the pipeline-built amd64 packages, run the integration test binary, collect coverage, and confirm every test server was really deleted. main integration is manual and can omit cPanel; version tags require the cPanel image, baseline-to-candidate package upgrade, installer, WHM, mail watcher and forward guard checks. |
| sign | Detached signatures on release artifacts |
| publish | Internal GitLab Generic Package Registry (versioned + latest) |
| repo | Publish RPM/DEB to the public mirrors.pidginhost.com apt/dnf repos |
| pages | Docs + coverage HTML (GitLab Pages preview) |
| cleanup | Remove old package versions |
| release | GitLab release on tags matching v* |
| github | Mirror to GitHub + upload release artifacts (auto on tag push) |
Public Releases
To cut a release:
- Move the
[Unreleased]heading inCHANGELOG.mdto the new version (e.g.[2.4.2] - YYYY-MM-DD), commit asrelease: cut X.Y.Z.CHANGELOG.mdholds one block of ten minor versions. The first release of a new block (3.40.0, 3.50.0, 4.0.0) moves the previous block and its link references todocs/changelog/<first>-<last>.mdand adds that file to the archive list at the top ofCHANGELOG.md. - Tag and push:
git tag vX.Y.Z git push origin main vX.Y.Z - Wait. The tag pipeline runs integration, publishes packages to the mirror, creates the GitHub release, and uploads every artifact including the fresh
merged-coverage.out. No manual pipeline clicks needed.
Tag pipelines require INTEGRATION_CPANEL_IMAGE to name a clean cPanel CI
image, or CSM_RELEASE_WITHOUT_CPANEL to state why no image is available.
INTEGRATION_CPANEL_PACKAGE optionally selects its compute package and
defaults to cloudv-2. A release taken without cPanel coverage records
cpanel_coverage: "absent" in dist/cpanel-release.json; see
cPanel release tests.
Tag-specific publish dependencies require preflight, fixtures, corpus,
production tags, kernel tests, signed artifacts and integration. Repository
publication and the GitLab release depend on publish; the GitHub release
also directly requires integration and the test gates. A missing csm-kernel
runner leaves publication pending. A missing cPanel image fails tag preflight.
See cPanel release tests for image acceptance and
the roadmap
for remaining operational readiness work. These configured dependencies are
not evidence of a successful live release run.
The coverage badge rebuilds automatically once the GitHub release exists, because the Pages workflow fetches merged-coverage.out from the latest release that carries one (it walks back through releases if the newest is missing the asset).
Installs and upgrades on end-user servers come from the GitHub release artifacts or the apt/dnf mirror. The internal GitLab package registry is operational tooling only.
Code Conventions
- Imports: stdlib, blank line, third-party, blank line, internal. Use
goimports -local github.com/pidginhost/csm - Errors: Return up the call stack. Wrap with
fmt.Errorf("context: %w", err) - Store:
store.Global()singleton bbolt DB. Always nil-check. - State:
state.Storehandles finding dedup, alert throttling, baseline tracking, latest findings persistence. Passed to subsystems at init - Web UI: Vanilla JS, no framework, no build step, and syntax up to ES2019 (a test rejects later syntax). Tabler CSS framework. Each script keeps its helpers in a function scope and adds only to the
CSMnamespace; the shared runtime iscsm-core.js,csm-format.js,csm-page.jsandcsm-live.js. UseCSM.get()/CSM.post()/CSM.delete()for API calls. Escape string-built markup withCSM.esc(); prefer DOM APIs for attacker-controlled values. Run the script tests withnode --test ui/. - API responses: answer through
writeJSONError,writeOK,writeItems/writeAll/writeCappedandwriteJSON, so every route follows the contract in API Reference: JSON errors,okon actions, lists underitems, UTC instants, durations in seconds and severity labels. - Logging: New code should use
internal/log(wrapslog/slog). Legacyfmt.Fprintf(os.Stderr, "[%s] ...", ts())call sites remain valid until migrated.
Attack event storage
Attack events live in attacks:events; attacks:events:ip stores empty values
under <ip>/<TimeKey> keys. The writer chooses an unused primary key inside
the write transaction because batch counters can repeat for the same timestamp.
Primary rows, index entries and the event count are updated atomically, including
count-cap pruning.
Address queries walk the index newest-first and resolve primary rows until the requested limit is met. Older index entries can still contain a full event copy; the reader falls back to that copy if the primary row is missing, malformed or belongs to another address. Both forms must match the requested address. The UTC time-key migration preserves index values verbatim, so readers must keep handling both forms until older entries age out.
Structured Logging (slog)
Legacy daemon call sites emit log lines via fmt.Fprintf(os.Stderr, "[%s] ...", ts()). The internal/log package provides a drop-in slog wrapper so operators can opt into JSON output for log-shipping pipelines (Loki, ELK, Datadog) without a big bang migration.
Operator controls
Two environment variables, read once at daemon startup:
| Variable | Values | Default | Effect |
|---|---|---|---|
CSM_LOG_FORMAT | text, json | text | Output handler |
CSM_LOG_LEVEL | debug, info, warn, error | info | Minimum log level |
Set via systemd drop-in:
# /etc/systemd/system/csm.service.d/logging.conf
[Service]
Environment="CSM_LOG_FORMAT=json"
Environment="CSM_LOG_LEVEL=info"
Then systemctl daemon-reload && systemctl restart csm.
Writing new logging code
import csmlog "github.com/pidginhost/csm/internal/log"
csmlog.Info("scan complete", "findings", len(f), "duration_ms", d.Milliseconds())
csmlog.Warn("log not found, will retry", "path", path, "retry_in", "60s")
csmlog.Error("alert dispatch failed", "err", err, "channel", "email")
Keys should be snake_case. Values should be machine-parseable (numbers, strings, booleans) – avoid formatted strings when you can pass the raw value.
Migrating legacy call sites
Migration is incremental and optional. The legacy format stays valid. Start with the hottest subsystems (alert dispatch, firewall operations, WAF handlers) where structured fields provide the most value, then work outward. Do not batch-convert – each subsystem should get a dedicated commit with before/after log samples in the PR description.
Keep the [TIMESTAMP] prefix of journalctl lines readable by humans: slog’s text handler uses time=... level=... msg=... which is also human-parseable, so journalctl viewers still work.
YARA-X Worker Process
CSM runs YARA-X in a supervised child process by default (since the
2026-04-23 default-flip). The goal is blast-radius control: a cgo
crash inside yara_x_capi (the 2026-04-16 production incident) stays
contained to the child and the daemon keeps its fanotify watchers,
log watchers, and firewall engine alive. Wiring lives in
internal/daemon/yara_backend.go; process supervision lives in
internal/yaraworker/supervisor.go. Completed work is recorded in
CHANGELOG.md and git history.
The knob is a tri-state *bool: omit it (or set true) for the
default-on child process; set false to fall back to the in-process
scanner.
signatures:
# yara_worker_enabled: true # default; omit for default-on
# yara_worker_enabled: false # explicit opt-out -> in-process
When on, daemon startup:
- Does not call
yara.Init()in the daemon process. - Builds a
yaraworker.Supervisorand callsStart(ctx). - The supervisor executes the running daemon binary with
yara-worker, the worker socket path and the configured rules directory. - Supervisor waits for the worker’s first
Pingbefore returning. - Installs itself as
yara.SetActive(...)so the existingyara.Active()callers (fanotify, rule reload) route transparently through the IPC.
Operator view:
ps axfshows the daemon with onecsm yara-workerchild.- New socket:
/var/run/csm/yara-worker.sock(0600, root-only). - Crashes produce a Critical
yara_worker_crashedfinding (rate- limited to one per minute) and restart with exponential backoff (1 s, 2 s, 4 s, capped at 60 s). Restarts reset to 1 s after the worker stays up for 30 s. - The crash finding is emitted before a restart is attempted, while YARA scans cannot run. Scans can resume once a replacement worker serves requests; they do not wait for the 30 s health check. The finding does not confirm a successful restart or recovery.
csm doctorreportswatcher: yara_workeras failed from a crash until a restarted worker has stayed up for 30 s, so a worker that keeps crashing shortly after each restart keeps doctor failing. This also covers crashes between initial readiness and backend activation. Shutdown waits for any in-flight recovery callback to finish.- A
csm update-rulesrun that completes triggers the supervisor’s in-processReload(the worker recompiles). Escalate to a full worker restart from Go code viaSupervisor.RestartWorker(). An explicit restart also marks the watcher failed until the replacement stays up.
Emailav under worker mode: the IPC wire format carries string-valued
rule metadata on every match (yaraipc.Match.Meta /
yara.Match.Meta). The emailav adapter consumes
Meta["severity"] via yara.Active(), so both in-process and worker
backends produce the same verdict shape. Non-string metadata (ints,
floats, bytes) is deliberately dropped at the worker boundary; add a
typed value struct here only if a future consumer actually needs one.
Testing:
- Unit-level:
internal/yaraipc(protocol framing + round-trip) andinternal/yaraworker(handler adapter, Run, supervisor). The supervisor tests re-invoke the test binary as a mock worker via the standardTestMain+ env-var helper-process pattern, including a realSIGKILL-driven signal-death test that exercises thesyscall.WaitStatus.Signaled()branch. - Integration: staged in the GitLab pipeline’s
integrationstage against AlmaLinux, Ubuntu, and the configured clean cPanel release image.
Building the Documentation
cd docs
mdbook build # generates docs/book/
mdbook serve # local preview at http://localhost:3000
Clean application corpus
The same job measures the shipped rules against long uninterrupted base64, hex and word runs, dense variable calls, and PDF-shaped streams. It discards one warm-up and scores the fastest of two further scans against a 5-second per-file budget. Each attempt has a 10-second engine timeout; a timeout never counts as a completed scan, and a cold timeout alone cannot fail the gate. Both scored attempts timing out fails it. The job shares the heavy-test resource group to avoid competing with other heavy tests in this project.
A rule with no literal atom to match on can still pass match tests while
stalling mail delivery. Investigate slow rules with yr scan --profiling.
Run the gate and its fixture/timing checks locally with YARA-X installed:
CGO_LDFLAGS="$(pkg-config --libs --static yara_x_capi)" go test -count=1 -v -tags yara ./internal/yara -run 'TestShippedRulesScanWithinBudget|TestRuleScanBudget'
Every pipeline runs the required clean-corpus gate in the production YARA-X builder image. Package publication and GitHub releases depend on its success.
See cPanel release tests for the required image, upgrade baseline, release dependencies and retained evidence.
Production tag selection, execution artifacts, and the required isolated kernel runner are described in Production build and kernel tests.
make check-fixtures is also a blocking GitLab job and publication dependency.
It checks all tracked and unignored files under testdata and fixtures,
including files without extensions. Scanner failures stop the check; reports
identify the file and line without printing the suspected address. The
fixture sanitisation rules
describe the additional manual privacy review.
Clean application corpus
The required test:clean-corpus job downloads the public vendor ZIP archives in
scripts/clean-corpus/manifest.json. The manifest pins the version, official
HTTPS source, SHA-256, license file, and exact regular-file count. Archives are
cached, rechecked before extraction, and unpacked into a new private directory
on every run. Missing input, checksum mismatch, missing license, empty or
incomplete inventory, unsafe archive entries, and read failures stop the job.
No downloaded PHP or JavaScript is executed.
Every source names the CMS it belongs to in a cms field, and the manifest
lists each supported CMS that has no pinned source yet under pending, with
the roadmap item that blocks it. The manifest is version 2; the validator
refuses older versions, a source without a CMS, a supported CMS that is
neither sourced nor pending, and one that is both, so a database scanner for
a new CMS cannot ship without a corpus decision. A pending entry records
absent evidence only: no scanner, admission rule, hit budget or result filter
reads it, and files from that CMS are treated exactly as before on a server.
When a source for a pending CMS lands, its pending entry is removed in the
same change. The manifest, not this page, is the record of which CMS has a
false-positive gate.
Ordinary unit tests validate the manifest offline, including with
go test -trimpath. Invalid metadata is rejected before cache, extraction or
report files are created or changed. Command tests check the full archived manifest
and every inventory record; these checks do not replace the detector gate.
The initial corpus contains 10,178 files: WordPress 6.8.2 (GPL-2.0-or-later), WooCommerce 9.9.5 (GPL-3.0-or-later), and Elementor 3.29.2 (GPL-3.0-only). The archives retain their license notices, including notices for bundled components. These are fixed test versions, not installation recommendations. Source URLs and license file locations are in the manifest.
Run the same gate locally with the Go toolchain from go.mod and the matching
YARA-X C library installed:
bash scripts/clean-corpus-test.sh
The script compares the installed C library version with both the Go binding
and production builder pins. It runs YARA, YAML signatures, PHP taint, and
JavaScript taint gates. Ordinary unit tests may omit external corpus inputs;
the required job sets CSM_CORPUS_REQUIRED=1, so missing input is fatal.
Package publication and GitHub releases explicitly depend on the job.
corpus-results/ is retained for one year, including on failures:
manifest.jsonidentifies the authenticated archives and licenses.inventory.jsonlists every extracted file, its size and SHA-256.engines.txtrecords the commit, toolchain, YARA-X and parser versions.- Engine JSON reports include counts, per-rule hits including zeroes,
thresholds, and taint status counts.
tests.jsonlrecords execution.
YARA and both taint engines require zero findings. YAML retains the existing
reviewed false-positive budgets in internal/signatures/corpus_fp_gate_test.go;
new rules default to zero. The pinned corpus currently produces 10 YAML hits
across five rules, all within those existing budgets. This is a regression
gate, not a claim that every engine has zero false positives. A second YAML
check fails if any regex would be skipped on a corpus file it matches.
The signature engines apply the production default 16 MiB bound and archive admission policy; 10,157 inputs reach each engine. PHP taint reads 10,171 inputs within its source limit, including 163 that complete analysis. JavaScript admission receives all 10,178 files and completes analysis on 289; seven are oversized and twelve have parse errors. These JavaScript gaps have explicit ceilings in the required gate; canceled, resource-limited and panic outcomes have a zero ceiling. Neither oversized nor unparsed inputs count as proof of successful analysis. The required taint gates each need at least 100 completed analyses as well as their corpus inventory checks.
To update the corpus, download a versioned archive from its official source, review its license and file inventory, and commit the new manifest pin with the measurements. Investigate new matches and narrow detectors using content semantics. Do not add scan-path exclusions or raise budgets to pass a run. The gate also compiles deliberately overbroad YARA and YAML rules and verifies that their clean-file matches are rejected.
Recorded finding streams
Calibrating correlation (the coordinated-attack threshold, corroboration
grading, sequence rules) needs what a real host produced over days: every
finding with its check, severity, timestamp, owner and text, including the
false-positive floods and the long-lived rows that never clear. CSM already
writes that stream: every dispatched finding lands in
/var/log/csm/audit.jsonl (rotated copies are gzip files beside it). A
recorded stream is a copy of those files with every identity removed.
scripts/finding-stream turns the raw files into an anonymized stream:
# On the operator machine, after copying the files read-only from the host:
go run ./scripts/finding-stream anonymize \
--salt-file ~/.local/share/csm/finding-streams/salt \
--out ~/.local/share/csm/finding-streams/host-a.jsonl.gz \
raw/audit.jsonl raw/audit.jsonl-*.gz
What the tool replaces, in every structured field and in free text:
- Host names become
host-<id>, account namesacct-<id>, domainsdom-<id>.example, mailboxesuser-<id>@dom-<id>.example(a mailbox truncated after the@keeps the sameuser-<id>). The<id>is derived from an HMAC of the value under a private salt, so the same salt maps one account to one pseudonym on every host and streams can be joined without knowing who is who. Names are replaced wherever they sit: in/home/<account>/paths,Account:lines, process context, LiteSpeed vhost tokens, and inside longer tokens such asexample.com-ssl_logorcp1.log. The host’s short name (its first label, when it carries a digit) maps to the same pseudonym as the full name. Any other token that looks like a domain (two or more labels, alphabetic last label that is not a file extension) is mapped too, even if no field named it; that over-reaches on a few dotted names likeoptions.optionand is accepted. Domain-shaped names inside filenames are replaced even when no structured field named them. Only pseudonyms actually emitted by this run are exempt; a raw name beginning withhost-oracct-is still scrubbed. - IPv4 addresses map into 198.18.0.0/15 and IPv6 addresses into
2001:db8::/32, both reserved and never routed; a raw value always maps to
the same pseudonym. The IPv4 range holds 131,072 addresses, so distinct
addresses can share a pseudonym. A manifest counts distinct addresses and
distinct pseudonyms per family, and a replay cannot tell merged addresses
apart.
Loopback addresses, system users such as
rootornobody, and bare numeric uids are kept: they identify nobody and carry meaning. Equivalent IPv6 spellings share one pseudonym across all streams; IPv4 addresses written in IPv6 form use the IPv4 pseudonym. Addresses inside filenames and before numeric rotation suffixes are also replaced. This can replace address-shaped version numbers; privacy takes priority over preserving ambiguous numeric text. - The details of a credential-leak finding are dropped entirely, and generic
password=,secret:andtoken=material is blanked anywhere, including quoted keys and values containing spaces. - Finding ids become
fid-<id>, a salted id of at least 128 bits that keeps case, so an action row can name its finding without exposing the raw id. - Timestamps, check names, severities, path structure below the account, plugin and file names, and process names are kept unless they contain an identity: they are what calibration reads. Paths and mailboxes in process command lines and parent processes also teach the scrubber which identities to remove elsewhere.
Before writing, the tool scans its own output for every identity it learned from structured fields, paths and mail addresses, for domain-shaped names, for raw ids (including inside longer tokens), and for any mailbox or address outside the reserved ranges. It refuses to write if it finds one. The raw-id scan includes check names and severity labels; these fields otherwise keep their original vocabulary. Structured finding ids must be ids this run emitted. A name glued to underscores or file extensions is still found. The summary prints row counts, the number of distinct checks, replacement counts, the time span and a salt fingerprint, never identities, check names, paths or the salt itself. An error names only the stream, the file’s position on the command line, the line number and a fixed reason.
Every row must parse as exactly one JSON object of the known schema, with a supported version and a timestamp field. Rows that older CSM versions wrote with a zero time are kept, counted in the manifest and left out of time spans. Unknown or repeated fields, nulls, data after the object and values over the size limits refuse the whole run. Malformed Unicode, non-JSON whitespace and timestamps that cannot be written back as JSON also refuse the run before a salt is created. Typed rows accept only the exact pseudonym spellings emitted during that run.
Handling rules:
- Copy the audit files read-only (
taroverssh, orscp) into a local directory with mode 0700, run the tool, then delete the raw copies. Nothing runs on the monitored host. - The salt file is created exclusively on first use with mode 0600. Existing salts must be regular files without group or other access; symlinks are refused. A concurrent run that reads an unfinished salt fails without replacing it; retry after the creator finishes. Keep it private and reuse it for every host whose stream should be joinable with the others; losing it makes new recordings unjoinable with old ones.
- Outputs cannot replace an input, the salt, the input manifest or each
other, whether named by path, through a symlinked directory or as a hard
link, and an existing output must be a regular file. These checks run
before the salt is created, including for missing directories and paths
containing
... Case-only and Unicode-equivalent path variants are refused on every platform, as are paths that would need a file to also serve as a directory. Every output is written to a private temporary file beside its destination, and all are complete before any is renamed into place; if publishing fails partway, the outputs already replaced are restored. New directories have mode 0700 and outputs mode 0600, even when replacing a less restricted file. - Filesystem cleanup failures are reported without claiming success. If rollback fails, remaining recovery copies are retained. If every output was published but backup cleanup fails, the error explicitly says the outputs were published; check the manifest and remove the leftover backups before sharing the directory. Other cleanup errors report unchanged outputs with temporary files left to remove.
- Recorded streams stay outside the repository entirely, in a private
directory such as
~/.local/share/csm/finding-streams/with mode 0700. A pseudonymized stream still describes real incidents on a real host, and this repository is public. Do not keep them in an ignored directory inside the checkout:git clean -fdxremoves ignored files too, and a recording that took a host weeks to accumulate is not reproducible from anywhere else.
Joining actions and firewall entries
The action log (actions.jsonl) and the firewall audit log
(<state>/firewall/audit.jsonl) say what CSM did about the findings. The same
run can anonymize both next to the findings, so the three streams share
pseudonyms. A joined run needs a manifest, and a manifest needs a build of a
known commit without local changes, so build the tool from a clean checkout
first:
go build -o /tmp/finding-stream ./scripts/finding-stream
/tmp/finding-stream anonymize \
--salt-file ~/.local/share/csm/finding-streams/salt \
--out host-a/findings.jsonl.gz \
--actions raw/actions.jsonl --actions-out host-a/actions.jsonl.gz \
--firewall-audit raw/firewall-audit.jsonl --firewall-out host-a/firewall.jsonl.gz \
--manifest host-a/manifest.json \
raw/audit.jsonl raw/audit.jsonl-*.gz
Action and firewall rows are rebuilt from a closed list of fields rather than scrubbed:
- An action row keeps its time, operation and action, actor kind, result, a
reason category, whether an error occurred, whether the file existed before
and after, a block lease, and salted ids for the finding, incident, action
and the action it undoes. Addresses map as in findings, networks keep their
prefix length and endpoints their port and protocol. File paths, process ids
and other targets become
tid-<id>. Command lines, error and reason text, undo commands, recovery paths, digests, sizes, modes and owners are dropped. Account names are always mapped, system users included. - A firewall row keeps its time, action, target, reason category, source and lease. Firewall entries carry no ids, so nothing joins them to actions or findings; the manifest says so rather than guessing from times or addresses.
- An operation, action, actor, result or source the tool has not been reviewed against refuses the run instead of passing through.
The manifest is written last and is the bundle’s completion marker. It lists every input and output with its stream kind, position, SHA-256 of the exact bytes, record count and time span, the tool’s source revision, the salt fingerprint, join counts, counts of discarded fields, action results by value, and coverage for each stream. Check every output’s digest against it; a bundle without a manifest, or with a digest that does not match, is incomplete.
- An action is matched when it names a finding id present in the recording, and nothing more. Actions naming a finding the recording lacks, and actions naming none, are counted apart.
- Repeated finding rows are kept and counted. A durable action row repeated exactly is a retransmission and does not count as another outcome; rows that share an action id and version but differ are reported as conflicting, never resolved by taking the latest.
- A result is what the writer recorded. An applied block or a firewall entry is an observation, not a verified effect or a reviewed correct action.
--input-manifest takes the collector’s inventory of what the host has:
{"v": 1, "streams": [
{"kind": "findings", "availability": "present", "sha256": "<digest of the copied file>", "records": 1200},
{"kind": "actions", "availability": "not_recorded"},
{"kind": "firewall_audit", "availability": "absent"}
]}
Every supplied file must appear as a present entry with its digest and record
count, and every present entry must be supplied. absent means the collector
looked and found none; not_recorded means the host does not keep that
stream. Ledger and review streams can only be stated absent or not recorded.
Without an inventory, a stream that was not supplied is reported as not
supplied.
What a recording does and does not contain
The audit log is the dispatch record: it holds findings that were alerted, after deduplication. It is not the persisted latest-state set, which is larger because every scan re-emits the findings it still sees and the merge refreshes their timestamps. On one production host a single sweep dispatched 49 critical findings while the active set carried 157 refreshed critical rows.
A replay therefore understates how full the persisted correlation window gets. Compare windows against each other, and treat an absolute rate from a replay as a property of the replay, not of the live host.
Replaying a stream through correlation
scripts/correlation-calibrate replays a recording through the production
cross-account correlation so its thresholds can be re-derived from what hosts
produced rather than from an assumption:
go run ./scripts/correlation-calibrate \
~/.local/share/csm/finding-streams/host-a.jsonl.gz
It reports how often account-and-check pairs repeat, which checks dominate
the input, and what the aggregates did under both derivations:
per dispatch batch, and over the persisted latest-state set. For each it prints
a threshold sweep using the same eligible, windowed accounts as the aggregate.
--window selects the persisted correlation bound: it defaults to
one hour, and --window 0 reproduces the original unbounded correlation.
Longer windows are applied directly, without the production default limiting
them. The source set keeps the store’s retention and size limit. Batch
correlation counts all qualifying rows grouped into that dispatch batch.
Repeated pairs can include distinct findings on the same account; this metric
does not establish how many rows are re-reports of the same finding.
Batch boundaries are inferred from timestamp gaps. Persisted correlation is recomputed at every recorded arrival, including ignored checks that only advance time. Raised duration is measured between those observations, not against an invented scan schedule. Recordings do not contain empty scans, purges or dismissals, so this replay models latest-state accumulation rather than reconstructing every store transition. Rows with no timestamp are counted and skipped because they have no replay position; runtime correlation still counts unstamped stored rows.
The tool needs no host access and reads nothing but the recording. It uses cPanel account roots supplied directly to correlation, without platform discovery on the replay machine.
Replaying scan admission
scripts/response-replay replays a recording through a model of how
automatic scan blocks are admitted today: the hourly block limit, the retry
queue and its age limit, and eviction at the temporary deny limit. It reports
what the model would have done with the recorded findings, not what the host
did, and not whether any block was right.
go build -o /tmp/response-replay ./scripts/response-replay
/tmp/response-replay --findings host-a/findings.jsonl.gz --out replay/host-a.json \
--max-blocks-per-hour 200 --deny-temp-ip-limit 500 --seed 1 --hour-zone UTC \
--manifest host-a/manifest.json
- The live path tries queued candidates in Go map order. The model tries them
in a reproducible random order chosen by
--seed, so compare several seeds rather than reading one run as the order the host used. --hour-zoneis the zone the host’s clock runs in: the hourly limit resets when that zone’s hour changes.Localis refused because it would depend on the machine running the replay.- A recorded block from a challenge timeout, central intel, a credential spray or an incident is applied as recorded: it takes a firewall slot but not the hourly budget. Only a single IP address is accepted as the block target; malformed targets are counted as unclassified rows. A permanent block is counted, not modelled.
- The audit log does not record a finding’s structured source address, so
a reputation finding, which names its address only there and in its
message, has no address in a replay by default and is counted as missing
one.
--reconstruct-reputation-sourcerecovers the address from that message’s fixed form; the report then counts the recovered rows and lists the reconstruction among its assumptions. - Batches are inferred from equal timestamps. Recordings hold no empty scans, so queued work drains only when another finding arrives. Rows without a timestamp are counted and left out.
- The report lists the effective policy with every default it applied, the counts, queue delay and eviction residence distributions, blocks per elapsed hour, and the assumptions and mechanisms the model leaves out: infrastructure and allowlist protection, verdict callbacks, subnet blocks, permanent escalation, retries and failures, and manual unblocks. It carries no address, id, name or text from the recording.
- Candidates the queue dropped (aged out or overflowed) are reported apart from work still queued when the recording ended, which the end of the recording cut off rather than the queue.
- The demand is the audit log’s record of dispatched findings, not every finding the admission path was called with.
- The report counts the recorded subnet, ASN-crawl, netblock, permanent and dry-run block rows the model leaves out, and the findings that feed the subnet paths. ASN-crawl subnets spend the same hourly budget the model gives single addresses, so free scan slots in a replay are an upper bound.
- Hourly distributions include empty elapsed hours across the full recorded time span, without allocating a sample for each empty hour. The first stamped row establishes the replay clock, including dates before year one.
- With
--manifest, the report checks that the manifest describes this exact recording and carries its coverage and recorded outcomes beside the replay. Without one, the report says the recording has no statement of what was collected with it. Negative outcome and join counts are refused. - A report is written only to a new path: it never replaces an existing file, and it cannot alias its recording or manifest through symbolic links, parent-directory traversal or hard links. It is staged beside the resolved destination and published as a private file, refusing the run if another writer creates the destination in the meantime. A temporary file cleanup failure is reported with whether the report was published.
- Like a manifest, a report needs a build of a known commit without local changes.
cPanel release tests
Version tags require either a cPanel integration image or a recorded reason for
releasing without one. release-preflight runs before builds, and integration
repeats the preflight before allocating servers. The tag-specific publication
dependencies require that integration succeeds; package registry publication,
repository publication and releases cannot use an AlmaLinux/Ubuntu-only result
unless the omission was acknowledged. Main-branch integration remains manual and
may run without cPanel when no image is configured.
Releasing without cPanel coverage
CSM has no licensed cPanel image while no cPanel licence is available for
disposable CI clones. Set the protected variable CSM_RELEASE_WITHOUT_CPANEL
to a sentence stating why, for example no licensed cPanel image available.
A reason must have at least 12 characters after trimming, contain at least two
words, and fit on one line. Blank values and bare flags such as 1 are rejected.
The reason is published in the pipeline log and in
dist/cpanel-release.json as cpanel_coverage: "absent".
Do not read such a release as cPanel-tested. WHM plugin installation, mail integration, cPanel platform paths, service confinement under a real cPanel layout, and upgrade behaviour are unvalidated in that pipeline, and the findings that depend on them (F19, F22, F25) stay open. Clear the variable as soon as an image exists.
The CSM release maintainer owns the CI variables and image refresh schedule.
The PidginHost cloud image administrator owns capture and publication of the
private reusable image. The release maintainer must verify a clone before
setting INTEGRATION_CPANEL_IMAGE to its immutable ID or versioned slug.
Do not set it to a base AlmaLinux image: the package test requires a working
cPanel installation and fails if it is absent.
Image preparation
- Allocate a dedicated AlmaLinux 9 x86_64 build VM with the CI SSH key.
The default test plan is
cloudv-2(8 GB RAM, 80 GB disk); override withINTEGRATION_CPANEL_PACKAGEif needed. Follow the supported OS, hostname, storage and licensing requirements in the cPanel installation guide. - Install full cPanel/WHM using the official installer. Complete initial setup with EA4 Apache, Exim, Dovecot, rsyslog and a valid license arrangement for disposable clone IPs. DNSOnly is insufficient. Do not install CSM or create tenant accounts on the image.
- Verify
systemctl is-active exim dovecot,exim -bV,doveconf -n, and the WHM AppConfig tools./var/log/maillogmust exist. Ensure cloud-init preserves a valid cPanel hostname and injects the dedicated CI SSH key into the existingphuseraccount with passwordless sudo. - Have the cloud image administrator capture the VM using the provider’s image preparation process. Remove build credentials and tenant data; regenerate host keys and machine identity on clones. Record the OS, cPanel version, image build date and responsible maintainer in the image metadata. Never publish an image containing a customer backup.
- Publish the private image under an immutable ID or versioned slug and
verify that the CI account sees it in
phctl compute image list. The current CLI can create VM snapshots but cannot promote them to reusable OS images; that step requires cloud image administration access. - Set the protected GitLab project variable
INTEGRATION_CPANEL_IMAGE. Run the manual integration job once these checks are on main before relying on the image for a tag release.
Refresh the image for cPanel/OS security updates and review it at least monthly. License activation, SSH injection, empty CSM state and mail service readiness must work on a new clone, not only on the image-build VM.
What the release job checks
The cPanel VM starts without CSM. The job downloads the previously released
RPM pinned in build/integration-upgrade.json and verifies its SHA-256. It
installs that version, writes operator configuration and a state sentinel,
starts its daemon, then upgrades to the RPM built by the current pipeline.
It checks that configuration, drop-ins and state survive and that the installed
binary has the exact SHA-256 of the current build artifact.
The candidate’s package-mode installer is exercised on that cPanel host, and the service is restarted using its installed unit. Dedicated tests require:
- cPanel and Exim platform detection with existing platform-derived paths.
- Working Exim and Dovecot services and configuration commands.
- Executable WHM CGI, registration reported by the read-only WHM application-list API, and the expected redirect.
- A running notify service with strict filesystem protection, healthy state, and an attached mail-log watcher.
The ordinary real-system integration suite then runs on that same host and on the AlmaLinux and Ubuntu hosts. No test sends customer mail; alert delivery and automatic response are disabled in the integration drop-in. The cPanel package test only runs with an explicit disposable-server marker and refuses an image with CSM or its state already present.
Update the baseline pin deliberately after reviewing the released RPM and its
published SHA-256. Never resolve a moving latest package during this test.
Evidence and cleanup
Integration retains cPanel package output, service/journal diagnostics,
baseline and candidate status snapshots, and cpanel-release.json for one
year. The JSON identifies the candidate commit, binary/package hashes,
image, plan and upgrade baseline. Its package_checks field covers the
package phase; the integration job must also pass the later real-system tests
and verified server cleanup before publication is allowed.
Server IDs are recorded immediately after creation. Verified cleanup runs
before successful completion, with an after_script retry for errors or
cancellation. A failed package check retains diagnostics and still reaches
that cleanup. Image preparation and a successful live cPanel job are required
operational steps; compiling the integration binary locally does not replace
them.
The candidate daemon also enables and removes the forward guard through a test drop-in and service restart. This runs the actual cPanel rebuild through the packaged service sandbox and checks that removal preserves operator configuration.
Production build and kernel tests
The default suite tests the portable stubs. test:production also runs the full
repository with yara,journal,bpf, the tags used by shipped Linux binaries.
It uses the release builder’s YARA-X C library and verifies that its version
matches both the Go module and the builder recipe. PHP 8.2 is installed for
runtime tests. The same job runs the pinned linter and security analyzer with
all shipped tags; scanning APIs retain documented exceptions for intentional
reads of caller-selected local files. Rule trust checks cover every accepted
filename extension, regardless of letter case. Run the same gate in a suitable Linux image with:
scripts/production-tests.sh portable
Each gate writes the selected test inventory, Go JSON test events, engine
versions, and commit ID under production-results/. The inventory comes from
go list and the selected Go test files across all packages. A newly added
package or tagged test is included automatically. Verification rejects missing
tests, package failures, and skipped required regressions. Other environment
skips remain visible in the JSON events. These are not kernel coverage.
Kernel runner
test:kernel is a required job on a dedicated Linux shell runner tagged
csm-kernel. Provision an ephemeral VM for each job, locked to this project,
with local Docker, cgroup v2, BTF, BPF LSM enabled in the boot-time LSM list,
BPF ring buffers, fanotify, and nftables. Give it access to pull the pinned
builder image. It must have no unrelated workloads or host credentials.
The ordinary Kubernetes runner does not satisfy this contract automatically.
Provision it from the cloud catalogue’s alma9 image (AlmaLinux 9, x86_64),
the closest available match to the EL production family; alma10, ubuntu26
and debian13 also satisfy the contract. Enable BPF LSM before registering the
runner, since it is not in the default boot-time LSM list:
grubby --update-kernel=ALL --args="lsm=capability,yama,selinux,bpf"
reboot
cat /sys/kernel/security/lsm # must list bpf
stat -fc %T /sys/fs/cgroup # must print cgroup2fs
Register with gitlab-runner register --executor shell --tag-list csm-kernel,
lock it to this project, and disable it for other projects.
This runner does not reproduce the production kernel. Supported hosts run
CloudLinux 8 and EL8 on 4.18; no catalogue image offers that kernel. Treat a
passing test:kernel as evidence for a 5.14-or-newer kernel only. Capabilities
that differ across those kernels – pidfd_open is present on the runner and
absent on 4.18 – are probed at runtime instead and reported by csm doctor
and the health status.
The job builds build/Dockerfile.production-test on the release builder and
uses scripts/go-linux.sh to boot systemd inside a disposable container.
GO_LINUX_PRIVILEGED=1 is an explicit Docker-only test option; the wrapper’s
normal capability set remains the default. Use the same setup locally only on
a dedicated Linux test VM:
docker build --build-arg BUILDER_IMAGE=<release-builder-image> \
-f build/Dockerfile.production-test -t csm-production-test .
GO_LINUX_RUNTIME=docker GO_LINUX_PRIVILEGED=1 \
GO_LINUX_IMAGE=csm-production-test scripts/go-linux.sh \
bash scripts/systemd-account-roots-test.sh production
Systemd runs the packaged service sandbox regression and a second test service.
The second service executes the shipped tags plus nftkernel,kernelintegration.
Its inventory selects every test added by those tags across all packages, plus
explicitly required attachment tests. New kernel test names need no prefix.
Required checks include real nftables transactions in isolated network
namespaces, kernel-produced BPF ring events, AF_ALG BPF attachment and shutdown,
and journal delivery from two actual services followed by reader cancellation.
The journal test covers both empty and existing history, excludes unrelated
services, and rejects replayed records.
Both the CLI and service test binary use the shipped tags.
The kernel test service has an explicit Go workspace and toolchain selection;
EL8 system services may start without a home directory. Once tests stop and
their results and journals are saved, the collector exits the disposable
systemd manager directly with the test result, avoiding EL8 exit-target loops.
The first service verifies remediation and restore under the packaged unit,
custom account-root grants, doctor checks, and process-handle signaling.
It does not exercise the daemon’s complete watcher startup; the cPanel package
gate and cloud integration cover that separately.
Missing capabilities fail the kernel gate. A skipped AF_ALG attachment test is not accepted. Both production and kernel jobs retain artifacts for one year and are explicit dependencies of publication, including tag-specific dependencies. A missing kernel runner leaves the job pending and publication blocked.
The separate clean application gate provides pinned vendor corpus measurements. cPanel release tests provide package, upgrade, and primary-platform validation. None substitutes for another.
The journal reader starts after existing matching records and also follows services with no prior records. Package integrity rechecks preserve a reported mismatch for the original finding even if that file is now absent or no longer executable; current file mode is not evidence that the modification was repaired.
Required Queue inventory
Both runner modes validate scripts/queue-inventory.json before test selection
and execution. Its reviewed owners cover channel handoffs, durable work,
coalesced pending work and kernel buffers. Each owner records its health rows,
bounds and required publication and lifecycle regressions. The runner writes
the union of those regressions and its existing requirements to
production-results/<mode>/queue-required.json, then requires actual passing
events for every selected requirement.
The portable baseline in scripts/production-required.json also requires the
shared accounting, health reporting, doctor and gate regressions. Keep recovery
and concurrency regressions with their owner when they test an owner-specific
boundary; keep shared contracts in the baseline. A passing full suite alone
does not protect an omitted requirement against a later skip.
PHP relay publication requirements start the Linux wiring with temporary filesystem paths and a real state store. They check live registration, persistence, shutdown evidence and failed watcher attachment. They must also reject removal of either production registration; manually registering a provider in a test only proves that provider’s behavior.
The allocation scanner is available for ownership reviews:
go run ./scripts/queuegate -scan > queue-allocations.json
It scans repository Go source across all build constraints, excluding test files,
testdata, hidden directories and vendored dependencies. It records raw channel
allocations and imported queuehealth.NewChannel calls, including renamed and
dot imports. Named channel aliases resolve across every package variant in a
repository import directory; an alias that is a channel in only some variants fails
explicitly. Imported named types use the Go toolchain’s source importer.
Unresolved types, reflective allocation and indirect references to the accounted constructor fail explicitly.
Supporting a new constructor or generic constraint requires scanner tests and
an ownership review first.
Allocation identities use the file, enclosing function, assignment target,
constructor kind and ordinal. The descriptor includes the allocation expression,
expanded repository constants and assignments or local var initializers for
the capacity and its local inputs in the enclosing function. Expression grouping
and build-variant values are retained, including array lengths, literal indices and slice bounds.
Implicit constant declarations, iota, closures and named composite literals
in capacity expressions need explicit scanner support and are rejected, including
when nested inside field selections. Import aliases are resolved as package
names, including aliases that match predeclared identifiers. Runtime
collection lengths stay symbolic; only explicit construction and slice bounds
enter the local capacity-input graph. The scanner does not infer arbitrary
function bodies or type layouts and is not whole-program data-flow analysis.
A version 1 manifest contains allocations and owners. Each allocation copies
its scanner descriptor and adds a reviewed class, rationale, and queue
owner where applicable. Classes are work, lifecycle, maintenance, and
constructor. A semaphore or completion signal can order data stored elsewhere;
trace that data before deciding whether it belongs to a work owner. Every work
allocation must reference an owner. Changes, omissions and stale descriptors
fail validation.
Each owner records health rows, reviewed bounds, and separate publication
and lifecycle evidence lists. An evidence entry names a package, a top-level
Go test and its required portable or kernel mode. Owners without a channel
allocation need source anchors: a path, symbol and canonical shape. Supported
anchors include struct fields, types, variable or constant declarations, and
function signatures. Missing or changed anchors fail validation.
For a reviewed manifest, generate requirements for the existing test verifier:
go run ./scripts/queuegate -manifest scripts/queue-inventory.json \
-mode portable -base-required scripts/production-required.json \
-required-out queue-required.json
Use the resulting file as scripts/testgate -required with the selected test
inventory and actual Go JSON test events. It combines existing required tests
with the owner’s evidence for that mode. A missing, skipped or failed required
test is not accepted. A passing scanner or a test name in JSON alone does not
prove health publication or lifecycle coverage; the named tests need substantive
assertions and execution evidence. Runner regression tests exercise both modes
with real test events, including changed capacities, unclassified allocations,
missing tests and skipped required evidence.
The scanner cannot discover arbitrary queues held in maps, heaps, durable storage, kernel buffers or dependencies. Those need explicit owner entries and source anchors, plus tests that drive real admission, progress, loss and cleanup. Adding a new non-channel owner remains a code-review responsibility. No syntax inventory substitutes for reviewing the behavior of its required tests.
Release Signing
CSM has two separate signing paths:
- Package repository signing for the normal APT/DNF install path.
- Detached Ed25519 artifact signatures for raw binaries, tarballs, and package files downloaded outside the package manager.
Do not reuse keys between these paths. The package repositories use GPG because APT and DNF verify repository metadata that way. Detached release signatures use Ed25519 because the standalone install and deploy scripts verify raw artifact bytes with OpenSSL.
Status
| Surface | Key type | CI variable | Notes |
|---|---|---|---|
| APT repository metadata | GPG | CSM_GPG_SIGNING_KEY | Published by repo:publish; operators install with signed-by=/etc/apt/keyrings/csm.gpg. |
| RPM packages and repository metadata | GPG | CSM_GPG_SIGNING_KEY | Published by repo:publish; operators use gpgcheck=1 and repo_gpgcheck=1. |
Raw binaries, tarballs, .deb, .rpm siblings | Ed25519 | CSM_SIGNING_KEY | Detached .sig files for direct downloads and standalone scripts. |
| YARA Forge rule ZIPs | Ed25519 | CSM_SIGNING_KEY | Signed by the yara-forge-mirror job; clients verify via signatures.signing_key. |
The preferred operator path is the signed APT/DNF repository documented in Installation. Standalone scripts also verify detached signatures: scripts/install.sh embeds the Ed25519 public key in EMBEDDED_SIGNING_KEY. Override it at runtime with CSM_SIGNING_KEY_PEM. Current releases fail before execution if the key, signature, or compatible verifier is unavailable. Setting CSM_REQUIRE_SIGNATURES=0 does not bypass verification for current releases.
Public Key
The same Ed25519 key signs release artifacts and YARA Forge rule ZIPs.
Hex form, for signatures.signing_key in CSM config:
2d1472b2a1d9728c2717b75111487145a7863f7ce731c1b44181f7a68bb908f7
PEM form, for standalone script verification (EMBEDDED_SIGNING_KEY / CSM_SIGNING_KEY_PEM):
-----BEGIN PUBLIC KEY-----
MCowBQYDK2VwAyEALRRysqHZcownF7dREUhxRaeGP3znMcG0QYH3pou5CPc=
-----END PUBLIC KEY-----
Package Repository Signing
repo:publish runs on version tag pipelines and rebuilds the public package repositories from the current tag plus the retained historical releases.
Required protected CI variables:
| Variable | Type | Purpose |
|---|---|---|
CSM_GPG_SIGNING_KEY | File | GPG private key used to sign APT metadata, RPM packages, and RPM repo metadata. |
CSM_MIRROR_SSH_KEY | File | SSH key used to publish the mirror output. |
CSM_MIRROR_KNOWN_HOSTS | Variable | SSH host keys for the mirror host. |
The job exports the public key as csm-signing.gpg and publishes it at the mirror root so install docs can reference:
https://mirrors.pidginhost.com/csm/csm-signing.gpg
APT verifies signed repository metadata through the signed-by= keyring. DNF verifies both RPM package signatures and repository metadata via gpgcheck=1 and repo_gpgcheck=1.
Detached Artifact Signatures
sign:artifacts signs binaries and packages with the Ed25519 private key in CSM_SIGNING_KEY when that variable is present. The publish and GitHub release jobs create and sign csm-assets.tar.gz after assembling it. Each signed file gets a .sig sibling uploaded with the artifact.
Examples:
csm-X.Y.Z-linux-amd64
csm-X.Y.Z-linux-amd64.sha256
csm-X.Y.Z-linux-amd64.sig
csm-assets.tar.gz
csm-assets.tar.gz.sha256
csm-assets.tar.gz.sig
csm_X.Y.Z_amd64.deb
csm_X.Y.Z_amd64.deb.sig
csm-X.Y.Z-1.x86_64.rpm
csm-X.Y.Z-1.x86_64.rpm.sig
The signature covers the raw artifact bytes with no hashing wrapper. Verification uses:
openssl pkeyutl -verify -pubin -inkey csm-signing.pub -rawin \
-sigfile csm-X.Y.Z-linux-amd64.sig -in csm-X.Y.Z-linux-amd64
Detached Signature Setup
On a trusted workstation:
openssl genpkey -algorithm ed25519 -out csm-signing.key
openssl pkey -in csm-signing.key -pubout -out csm-signing.pub
Store the private key in GitLab as a protected CSM_SIGNING_KEY variable. Keep the private key in an offline password manager and a second secure backup location. Do not commit it.
For standalone script verification, either:
- Embed the public key PEM in
EMBEDDED_SIGNING_KEYinscripts/install.sh,scripts/deploy.sh, andscripts/deploy-gitlab.sh. - Or pass the public key at runtime with
CSM_SIGNING_KEY_PEM.
Signature verification is mandatory for current standalone releases. To also reject missing signatures on historical pre-signing releases:
curl -fsSLo /tmp/csm-install.sh https://raw.githubusercontent.com/pidginhost/csm/main/scripts/install.sh
sudo env CSM_REQUIRE_SIGNATURES=1 bash /tmp/csm-install.sh
Detached verification requires OpenSSL 3.0 or newer because the Ed25519 command uses pkeyutl -rawin. The scripts stop on older OpenSSL, including the default on EL8/CloudLinux 8, Ubuntu 20.04 and Debian 11. Use the signed APT/DNF repository on these supported platforms; its GPG verification does not depend on this OpenSSL command. The scripts never download and execute a verifier to bypass this requirement.
If a .sig file exists but verification fails, the installer aborts regardless of CSM_REQUIRE_SIGNATURES. A missing .sig (HTTP 404) is tolerated only with CSM_REQUIRE_SIGNATURES=0 and an explicitly selected complete release version older than v2.2.0, with no suffix or leading zeroes; for any later release the scripts abort even without CSM_REQUIRE_SIGNATURES, because every such release ships a signature and its absence means the download is incomplete or tampered. An unknown version or latest cannot qualify for the exception. This historical exception is checksum-only installation and is reported as unverified, not signature-verified.
The same policy applies to the csm-assets.tar.gz.sha256 checksum: releases published before checksums existed produce a warning and skip the checksum step, while CSM_REQUIRE_SIGNATURES=1 makes the missing checksum fatal. A checksum that exists but does not match always aborts.
The check action only compares checksums and reports update availability; it
does not download or execute a release binary. Installation and upgrade perform
the signature checks above.
Key Rotation
Package repository GPG key rotation:
- Generate a new GPG signing key.
- Replace
CSM_GPG_SIGNING_KEYin protected CI variables. - Publish a tag pipeline so
repo:publishexports the new public key to the mirror. - Update install docs or automation if the key URL changes.
Detached Ed25519 key rotation:
- Generate a new Ed25519 key pair.
- Replace
CSM_SIGNING_KEYin protected CI variables. - Update the embedded public key in standalone scripts, or rotate the
CSM_SIGNING_KEY_PEMvalue used by automation. - Tag a new release.
Old detached signatures remain verifiable only with the old public key. Archive old public keys alongside release metadata so historical releases can still be checked.
Manual Detached Verification
VERSION=X.Y.Z
curl -LO "https://github.com/pidginhost/csm/releases/download/v${VERSION}/csm-${VERSION}-linux-amd64"
curl -LO "https://github.com/pidginhost/csm/releases/download/v${VERSION}/csm-${VERSION}-linux-amd64.sig"
openssl pkeyutl -verify -pubin -inkey csm-signing.pub -rawin \
-sigfile "csm-${VERSION}-linux-amd64.sig" -in "csm-${VERSION}-linux-amd64"
openssl pkeyutl -rawin needs OpenSSL 3.0 or newer. EL8 and CloudLinux 8 ship
OpenSSL 1.1.1, whose command line cannot verify Ed25519 at all. Use an already
installed CSM build there, which verifies with Go’s implementation and needs no
OpenSSL:
csm verify-release csm-signing.pub \
"csm-${VERSION}-linux-amd64.sig" "csm-${VERSION}-linux-amd64"
It exits zero only for an artifact signed by the given key. The installer and
deploy scripts try OpenSSL 3.0+, then an installed csm verify-release, then
python3-cryptography, and refuse the artifact when none is available.
The Go and Python verifiers reject special files and limit input sizes: keys to 64 KiB, signatures to exactly 64 bytes, and nonempty artifacts to 512 MiB. Python verification ignores the working directory, user site packages, and Python environment settings when importing its dependencies. Install python3-cryptography as a system package.
If verification fails, treat the artifact as untrusted. Do not install it.