Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

API Reference

Machine-readable HTTPS API. All endpoints require token authentication. State-changing POST, PUT, PATCH, and DELETE requests require CSRF protection for browser cookie sessions.

Authentication

# Bearer token (header)
curl -H "Authorization: Bearer YOUR_TOKEN" https://server:9443/api/v1/status

# Cookie-based (after login)
curl -b "csm_auth=SESSION_COOKIE_FROM_LOGIN" https://server:9443/api/v1/status

Cookie-authenticated state-changing requests require the X-CSRF-Token header (obtained from the authenticated page meta tag). The token belongs to the browser session that loaded the page; another session’s token is refused, and a form field is read from the request body only. Admin-scope Bearer requests are CSRF-exempt because the Authorization header is the write credential.

Browser session management

A successful browser login exchanges an admin token for an opaque session cookie. API tokens are never valid cookies. Sessions expire on idle or absolute deadlines and all end on daemon restart. The header’s Sessions page provides the same management operations as these admin-only endpoints:

MethodPathResult
GET/api/v1/sessionsitems list of sessions with id, name, created, last_seen, expires, remote_ip, user_agent, current
DELETE/api/v1/sessions/<id>Revoke that session; an unknown ID answers 404
DELETE/api/v1/sessionsRevoke all browser sessions, including the caller’s

Revocation returns {"ok":true} only after committing to the store. Failure returns 503; invalid session IDs return 404. Lists contain no credential hashes or session verifiers. Cookie-authenticated DELETE requests require CSRF; admin bearer callers are exempt. Read-scope bearers cannot list or revoke sessions. POST /logout requires authentication and CSRF for browser callers; it revokes the current session and clears the cookie. GET /logout cannot log users out.

Token scopes

Configure tokens under webui.tokens: with a scope of admin or read:

webui:
  tokens:
    - name: "operator"
      token: "..."
      scope: admin       # full read+write
    - name: "panel-readonly"
      token: "..."
      scope: read        # status, findings, history, stats, challenge stats, blocked IPs, scan jobs, health, components, capabilities, SSE

The legacy single-token webui.auth_token: is migrated automatically to a legacy-auth-token admin entry on first start. Read-scope tokens are intended for orchestrators and dashboards that consume status, findings, history, stats, challenge stats, blocked-IP summaries, scan jobs, health, components, capabilities, and SSE events. Admin scope is still required for write routes and for sensitive reads such as quarantine, settings, firewall internals, threat-intel detail, rules, account detail, exports, incident timelines, and audit history. (ModSecurity stats/blocks/events are read scope; only the ModSecurity rules and escalation routes need admin.) metrics_token: is a separate, read-only credential for /metrics only.

Errors

Every failure answers a non-2xx status with a JSON body that has an error message. That includes CSRF, origin, rate-limit and wrong-method refusals. Some failures add detail next to it; a settings save that fails validation adds errors, a list of fields and messages. An unknown path under /api/ answers 404. No 2xx response reports a failure.

{"error": "invalid IP address"}

Actions

A request that changes state answers "ok": true with the action’s own fields, such as undo_token or warning. Work that continues after the answer, such as a daemon restart, a firewall rollback or a scan job, answers 202. A batch where some items failed answers 200 and lists the failures; a batch where nothing changed answers an error status. A read-only route answers 405 to any method but GET.

An empty fix batch is rejected. Threat actions report firewall failures, and bulk actions list invalid addresses as well as failed changes. A failed action can have applied some steps before the error; inspect the current state before retrying it. A response that cannot be encoded answers 500 with an error body.

/api/v1/firewall/check and /api/v1/firewall/unban also send "success": true for callers written against the older API. It will be removed; use the status code and ok.

Lists

Every GET route that returns a list answers a JSON object, never a bare array. The list is under items. An empty list anywhere in a response is [] and an empty map {}, never null; null means a value is not known.

  • total is the number of matches the server counted. It is larger than the length of items when the route pages or cuts the list.
  • offset and limit come with routes that page or cap the list. Capped routes that do not support paging report offset: 0.
  • truncated is true when matches were left out: past the page, past the limit, or past a scan cap. When a scan cap stopped the count, total counts only what was scanned. Incident groups also send scan_truncated to distinguish a scan cap from a page limit; their total is exact when scan_truncated is false.
  • Other keys next to items describe the whole list, such as check_types on /findings/enriched or summary on /email/forwarders.

/audit, /threat/top-attackers and /threat/events do not count every match. They send limit and truncated without total.

{"items": [{"ip": "203.0.113.9", "reason": "wp_login_bruteforce"}], "total": 1}

Times and durations

Every time is an RFC 3339 instant in UTC with sub-second precision, such as 2026-09-22T10:04:05.123456Z. The same applies to the event stream. A time that is not set is left out, never sent as 0001-01-01T00:00:00Z.

The API does not send times it formatted for reading, such as “5m ago”, “1h2m” or a clock time without a date. Clients format instants in their own time zone and count down to an expires_at themselves.

Durations are numbers of seconds in keys that end in _seconds, such as uptime_seconds, elapsed_seconds, duration_seconds, oldest_age_seconds and update_interval_seconds. Temporary whitelist responses use duration_seconds, rollback status uses remaining_seconds, and incident timelines use window_seconds. Request parameters such as hours and editable configuration values retain their documented units.

Two values keep a text form. started_at_token is an opaque token for restart polling. The temporary reason on /api/v1/firewall/check keeps its “(expires in …)” text for existing callers; expires_at carries the instant.

Severity

A severity is its label: CRITICAL, HIGH or WARNING. That covers severity, severity_max, demoted_from and the attack event sev. Clients derive colours and sort order from the label. A severity filter takes the label in any case, or the older 0, 1 or 2 level. The findings inside an email quarantine entry keep the scanning engine’s own rating.

Status & Data

GET  /api/v1/status              Full health snapshot: version, uptime, watchers, severity counts,
                                 store health, blocklist size, capabilities[], config_hash, binary_hash,
                                 automation rollout state, challenge pending count, rollback state.
                                 `started_at_token` changes after a daemon restart and is suitable
                                 for restart polling.
                                 `security_posture` is `healthy`, `warning`, or `critical` after
                                 combining daemon faults with active high/critical incidents.
                                 `incidents_open_by_severity` breaks open and contained incidents
                                 down by `critical`, `high`, and `warning`.
                                 `correlation_attribution` (present after the first active-set
                                 merge) lists per check the findings correlation could not
                                 attribute to an account: `current` for the active set now,
                                 `cumulative` since daemon start.
                                 `queues` reports named protection queues, including depth,
                                 capacity, running work, recent and cumulative drops, and lag.
                                 A degraded queue changes `status` and `security_posture`.
                                 `latest_scan` is the canonical last-scan timestamp; `last_scan_time`
                                 is a legacy alias kept for older clients and will be removed.
                                 `uptime_seconds` is the time since the daemon started.
GET  /api/v1/challenge/stats     Challenge-routing activity for the UI: `pending`, `escalated`
                                 (timeouts that became hard blocks), `routed_by_check` (per source
                                 check, since restart), and `recent` routes. Read scope.
GET  /api/v1/capabilities        Static feature list (e.g. `confd.dropins.v1`, `events.sse.v1`,
                                 `webhook.phpanel.v1`, `webui.prefs.v1`, `webui.undo.v1`,
                                 `mail.queue.composition.v1`,
                                 `detect.http_scanner_profile.v1`, `challenge.stats.v1`,
                                 `firewall.rollback.v1`, `firewall.dos_exempt.v1` on Linux builds).
                                 Use for orchestrator feature-detect.
GET  /api/v1/components          Watcher/component matrix with attachment, event, and upstream freshness state.
GET  /api/v1/events              Server-Sent Events stream of findings as they dispatch.
                                 Read-scope token sufficient. One JSON event per `data:` line.
                                 Writes and flushes have a three-second deadline. A failed
                                 write or flush closes the stream and frees its subscriber slot.
GET  /api/v1/health              Daemon health (fanotify, watchers, engines)
GET  /api/v1/findings            Current active findings
GET  /api/v1/findings/enriched   Enriched findings with GeoIP, accounts, fix info, and a list version.
                                 ?limit=N orders by severity, then newest, even when all rows fit;
                                 returns at most N rows with limit and truncated, while total and
                                 the severity counts cover all.
                                 ?fields=version returns only {version, total}, for change polling.
                                 block_ip is set only for checks that report an attacker address,
                                 the same evidence auto-block acts on
GET  /api/v1/finding-detail      Finding detail with action history (?check=&message=)
GET  /api/v1/history             Paginated history (?limit=&offset=&from=&to=&severity=&search=&checks=).
                                 checks is a comma-separated list of check names to include.
                                 total counts every match; truncated is true when matches exist past the page
GET  /api/v1/history/csv         CSV export of the newest 5,000 entries matching the /history filters
GET  /api/v1/stats               24h severity counts, accounts at risk, auto-response summary
GET  /api/v1/stats/trend         30-day daily severity counts; each `date` is a calendar day in the server's time zone
GET  /api/v1/stats/timeline      Hourly severity counts for the last 24 hours; each bucket names its `start` instant
GET  /api/v1/quarantine          Quarantined files with metadata (incl. htaccess pre_clean backups)
GET  /api/v1/quarantine-preview  Preview quarantined file content (?id=)
GET  /api/v1/db-object-backups   db_object_backups bucket (MySQL trigger/event/procedure/function drops)
GET  /api/v1/db-object-backup-preview Preview captured CREATE SQL (?key=)
GET  /api/v1/blocked-ips         Blocked IPs with reason and expiry
GET  /api/v1/accounts            Accounts a server-wide scan covers
GET  /api/v1/account             Per-account findings, quarantine, history (?name=)
GET  /api/v1/audit               Newest 200 UI audit log entries; truncated marks older ones. Each entry names the credential that acted (actor) and whether it came as an API token or a browser login (via)
GET  /api/v1/export              Export state (suppressions, whitelist)
GET  /api/v1/incident            Incident timeline (?ip=&account=&hours=)
GET  /api/v1/performance         Performance metrics snapshot (admin scope)
POST /api/v1/perf/fix-error-log  Truncate a fixed-row error_log finding (admin scope, CSRF)
POST /api/v1/perf/fix-display-errors
                                  Disable display_errors for a fixed-row config finding (admin scope, CSRF)
POST /api/v1/perf/fix-wp-cron    Disable WP-Cron and install a system cron for a perf_wp_cron finding (admin scope, CSRF)
GET  /api/v1/hardening           Last stored hardening audit report (admin scope)

Date ranges

/api/v1/history, /api/v1/email/groups and /api/v1/email/relay-abuse take from and to as a calendar date (YYYY-MM-DD) or an RFC 3339 time. A date is a day in the server’s time zone, and to includes the whole day. An RFC 3339 to is exclusive. The web UI sends RFC 3339 times so a day follows the operator’s time zone preference. A value in neither form is rejected with 400. An empty or whitespace-only bound is treated as absent. Calendar days start at their first valid local time, including when midnight is skipped or repeated; a skipped calendar date covers an empty range. Fractional seconds are preserved when filtering RFC 3339 bounds.

WordPress verification coverage

wordpress_verification contains core and plugins coverage when installation history is available. The same fields appear in csm status --json; human status shows the counts. Each entry reports verified, modified, unverified, unknown, not_wordpress and last_attempt. unknown means discovery found a directory but no completed attempt is recorded. modified means core verification completed with integrity differences, including missing core files. A verified plugin inventory can still contain vulnerable versions. The timestamp is the latest recorded attempt in that group. Counts retain the last observed results across restarts and cached scan cycles. Overlapping scans retain the latest attempt and its consecutive failure history in scan order. Discovery without an attempt does not replace verification evidence. An absent group means there is no retained installation history, not proof of coverage. An error field reports unreadable history instead of clean counts.

The first failed attempt increases unverified; repeated failures also produce an installation-specific Warning. In host scans, when one cause stops more installations than the per-cause limit, the warnings collapse into a single finding that names the cause, the total and a sample, so a host-wide fault cannot spend the whole alerts.max_per_hour budget. The limit applies across accounts for each kind of verification and cause. Account scans keep installation-specific warnings. A summary has no installation path and names an account only when all affected installations share it. These findings are separate from queue health: a command that completed with a refusal did not lose queue work. Existing queue loss accounting remains responsible for interrupted or missing command work.

Protection queue health

status.queue_health.v1 advertises the queues map on status responses. The same measurements appear under snapshot.queues in csm status --json and as named checks in csm doctor.

findings.ingest, fanotify.analyzer, fanotify.staged_packages, fanotify.dropper, fanotify.dropper_findings and spool.scanner report waiting work, capacity, running work, cumulative losses, losses during the last minute, the oldest waiting item’s age and the oldest running item’s processing time. Waiting work includes producers blocked on admission. Ingest work remains running while the dispatcher holds or processes its batch, including startup. Ages pause while the startup hold is in place and resume from its release, so a long baseline scan is not reported as a stall; work lost during the hold still counts. Releasing a hold preserves any future eligibility time for deliberately deferred work. Loss totals include the undelivered tail of a batch canceled during shutdown; scan warnings intentionally excluded from alerts do not count as lost work. Recovered file and spool scanner panics count as lost scan work. The workers continue processing later events, but repeated failures still degrade health. fanotify.reconcile reports the recovery queue in directory tasks, with 1,024 waiting slots. Repeated drops in a directory retain its original queue age. A detached scan batch remains running while new drops queue separately, including new work for a directory that is already being scanned. Eviction, work older than the recovery scan window, unreadable directories or candidate files, interrupted batches and unfinished shutdown work count as failed recovery tasks. An incomplete task may still have scanned some files; each directory task counts at most once. Directories and files that no longer exist when recovery runs hold nothing left to scan and complete the task, which keeps bulk extraction and package restores out of the loss counters. The recovery window includes its cutoff; work completed exactly at that boundary does not count as expired. Staged package verification reserves capacity for its whole running batch. Its full-queue timer starts when admission fills the queue and continues while the verifier retains those slots, including between retries. Files awaiting another attempt retain their original waiting age; retrying does not reset lag. Shutdown counts files still awaiting verification after the analyzer workers have stopped. Package metadata I/O cannot block health polling for this queue. Dropper candidates and held findings each have 16,384 waiting slots. Their lag starts after the configured unlink TTL or the finding’s 45-second grace period, respectively; intentional waiting does not count as overdue work. Detached probes and emission batches remain visible as running work. Processing time starts when each batch leaves its waiting queue, independently of the timestamp used to decide which work is eligible. Retries and refreshed observations retain earlier eligibility, and exhausted probes, capacity refusals and unfinished shutdown work count as losses. mail.delivery reports the selected mail-log reader’s 64-slot delivery queue. Its running time includes parsing and delivery of resulting findings. Losses include oversized complete file records, unreadable journal entries, canceled admission, consumer failure and buffered records abandoned on shutdown. Counters survive reader replacement and changes between file and journal sources. They count records already read, including complete records abandoned in file read-ahead, not unread source history or partial file records that have not reached a newline.

mail.file_source reports sampled untransferred bytes after the first file attachment, with depth_unit: bytes and capacity_unavailable: true. This includes disk backlog, buffered read-ahead and raw partial-line bytes, even when the reader retains only a bounded prefix of an oversized line. Transfer to mail.delivery removes the source bytes before output admission can block. These byte and record measurements are separate stages and must not be added.

A reader-owned sampler checks the current descriptor every two seconds, so appends remain visible while delivery is blocked. Health reads memory only. lag_basis: consumer_progress measures time without consumption while unread work remains. Sampling or new arrivals do not reset this clock. A partial line at EOF waits for more input and does not report a stalled consumer. Reader and sampler operations also report processing_lag after one minute without progress, including blocked reads, metadata calls and cleanup.

File I/O failures report source_io. A successful read cannot clear a failed cursor check; each operation must recover. Failed metadata sampling makes depth unavailable until a successful sample. Rotation and observed truncation retire the old generation, and late samples cannot overwrite its replacement. The initial attachment still starts at EOF; replacements start at the beginning. Copytruncate detection retains its existing limit: truncation followed by growth past the old descriptor offset between checks may be indistinguishable from an append.

Complete records already held in read-ahead count once in delivery loss when discarded. Unread disk data and unterminated fragments have no measured record count; generation changes, source failures and shutdown preserve dropped_lower_bound in the source row. A working file or journal replacement clears the retired source’s current error while retaining that history. Source path disappearance still follows the watcher’s existing grace period.

mail.journal_source reports the journal reader after its first successful attachment. The cursor exposes no exact unread-record count or queue capacity, so depth_unavailable and capacity_unavailable are always set, and lag_basis: unavailable marks the missing backlog age rather than reporting a waiting time of zero. The row measures the current operation instead. A selected entry remains in flight until delivery acquires it; known unreadable or abandoned selected entries count once in mail.delivery. The two stages must not be added together.

Journal progress follows cursor advancement, entry reads, bounded idle waits, output admission and close. One minute without operation progress reports processing_lag, including a stuck cursor with no known selected entry. Cursor, entry, wait and close failures report source_io; known abnormal exits remain visible through actual cleanup. An unknown cursor outcome or unread shutdown sets dropped_lower_bound without inventing a record count. Successful reads restore current health; historical uncertainty survives reader replacement. Successful attachment of a file source clears the retired journal’s current error. Failed attachment leaves it visible. Health reads memory only and does not call the journal or change its tail positioning, retry or delivery policy.

When their BPF backends are active, bpf.af_alg.output, bpf.connection.output, bpf.execution.output and bpf.sensitive_files.output report the 256-slot userspace delivery queues. Their losses include decoding failures, admission refusals, consumer panics and buffered output left after the reader stops. A received event remains running until evaluation and delivery to the finding queue finish, including events intentionally filtered during evaluation. Kernel ring occupancy and reservation failures are separate from these userspace measurements. bpf.connection.verdict reports the advisory annotation pool’s 256 waiting slots and up to four running callbacks. Repeated findings for a pending destination, reason and severity share one request and retain its original age. Capacity refusal, callback failure and abandoned shutdown work count as lost annotations, separately from lost findings. Callback errors release the pending key and allow a later finding to retry; a successful retry preserves earlier loss totals. Shutdown closes admission, waits for running callbacks and counts the queued requests left behind. Cached answers remain available without new work. Findings continue immediately when an annotation is unavailable. processctx.enrichment reports 1,024 waiting requests and up to two running process-context workers. It appears only after a BPF consumer initializes the shared pool; reading health does not start workers. Overflow replaces the oldest queued request and preserves the waiting age of the remaining requests. Running time includes the process read, latency observer, account resolution and cache write. Read errors and interrupted processing count as lost enrichment; vanished processes and stale or unverified identities are completed filtering. The queue dropped metric counts refused, evicted and abandoned requests; queue health also includes failures after a worker receives the request. The daemon stops the pool after BPF producers have joined, lets running reads finish and counts buffered requests discarded at shutdown. Final loss evidence remains available. This row measures enrichment requests, not each individual deadline-based file read inside a request. processctx.proc_reads reports the 64 shared slots for deadline-limited process file and symlink reads, including process-start captures on the BPF path. Waiting and running reads share this capacity. A slot remains occupied until both the underlying read and its caller have finished; an undelivered result still occupies its slot. A timeout counts the lost result immediately and keeps a blocked syscall visible as running work until it returns. Timeout and a later read failure count as one loss. Refusals at capacity and read failures count as losses; missing process files are expected churn. Synchronous reads without a deadline consume no slots. The initial process-directory check and the cached boot-time read are synchronous and are not measured by this row. smtp_rdns.resolves reports the 64 reverse-DNS lookup slots used by direct SMTP egress detection. A slot remains occupied until its resolver and caller finish. The one-second caller deadline also bounds healthy processing age; timing out counts a lost result once while keeping a resolver still running visible. Capacity refusals and resolver failures count as losses, while NXDOMAIN and successful empty responses are normal negative results. Cache hits perform no queued work. The bounded result cache is retained data, not waiting lookups. Status can initialize the empty cache but never performs DNS or waits for its cache lock. Synchronous lookups without a deadline use no slots or queue row. email_password.mailboxes reports waiting mailboxes and the five concurrent audits per scan, retaining each audit through finding collection and cache writes. Concurrent scans have no fixed combined waiting capacity. A busy pool uses each audit’s five-minute budget or shorter scan deadline; an idle slot with no dispatch progress for one minute reports backlog lag. Admission and post-audit work have their own one-minute budgets. Cancellation removes waiting demand without adding losses, while actual audits remain visible until they return. Expired work, failed verification, cache write failures and abnormal exits count once per mailbox. Unsupported, malformed and over-budget hashes remain normal incomplete-audit outcomes. Health retains loss counts after the batch drains and reads only memory. Deadline losses are recorded where an audit actually stops unfinished. A deadline arriving after a successful cache write does not turn that mailbox into a loss, and context evaluation during drain never holds the health lock. email_password.hashes reports the three shared password-verification slots. Each slot remains occupied until its KDF and caller have both finished, including after scan cancellation. email_password.waiting reports callers awaiting a slot; it sets capacity_unavailable because concurrent scans have no fixed global waiting limit. Both rows use the password audit’s five-minute check budget for lag. The occupied pool also reports sustained full capacity. A deadline while waiting counts one admission loss; a deadline after admission counts one lost verification result. A later KDF error or abnormal exit cannot count it twice. Explicit cancellation withdraws demand without loss; actual verification errors still count. Successful matches and nonmatches are normal results. Rejected input and already canceled callers start no queued work. These rows contain no password, hash, mailbox or account data, and status never runs a verification. The outer mailbox scan’s discovery, network enrichment and persistence are separate work from these hash slots. checks.executions reports dispatched check functions across host and account scans. An execution remains present until both its function and caller finish, including while its result awaits consumption or its function outlives a timeout. Lag is evaluated against each call’s original deadline, including a shorter parent deadline; a delayed function start does not reset it. Timeouts, panics and abandoned results count once per execution. Explicit cancellation withdraws demand without counting a loss, but a later function failure still counts. This aggregate sets capacity_unavailable: separate scans have their own wrapper limits, and timed-out functions can outlive those slots. It does not measure checks awaiting dispatch, scan-job persistence or later automatic actions. Status reads memory without waiting for a check or accessing the state database. checks.plugin_inventory reports sites waiting for the five inventory workers per refresh, retaining each site through command execution, result collection and cache storage or failed-result cleanup. Concurrent refreshes have no fixed combined capacity. Inventory uses a four-minute budget for its two bounded commands, or the shorter check deadline; admission and result handling each have one minute. A full pool stays healthy within those budgets. A free slot with no dispatch progress for one minute reports backlog lag. Command, decoding and storage failures count once per site, including when cleanup also fails. A command that ran and returned output (on stdout or stderr) with an error, such as a tree wp-cli will not inventory, answered the check and counts no loss. Missing command output still counts as lost work, including when the command exited with an error. Cancellation adds no losses; deadline withdrawal counts unfinished sites. Buffered sites stay visible until the original workers stop consuming and the refresh joins them. Actual commands remain in flight until they return. Health reads memory only and retains loss evidence after recovery. Shared-refresh waiters remain owned by checks.executions; they do not create another set of site jobs. Optional domain lookup fallback and plugin metadata enrichment keep their existing behavior. checks.wordpress_core reports installations waiting for the five checksum workers per scan, retaining each through its command, integrity findings and verified-file caching. The unit is installations; concurrent scans have no fixed combined capacity. Command lag uses the two-minute command budget or a shorter parent deadline. Result and cache work use one minute without progress. A full finite batch stays healthy within those budgets; a free worker with no dispatch progress for one minute reports backlog lag. Returned operational failures and abandoned work count once per installation. A command killed by a signal counts as failed work while retaining any partial integrity findings. Recognized integrity results from a completed command, including deliberately filtered output, complete without a queue loss. So does a command that ran and refused the tree, such as a directory that is not a WordPress installation or one whose configuration fails to load. Missing command output still counts as lost work. Cancellation withdraws unfinished demand without loss, while deadline withdrawal counts unfinished installations. Commands ignoring cancellation remain in flight until they return. A deadline during caching cannot undo completed verification, and an ordinary cache read failure does not turn a verified site into lost work. Health reads only metadata, including while result or cache locks are occupied, and retains confirmed losses after recovery. Discovery is separate check work. auto_block.waiting reports scan, direct-block, firewall-flush and startup observation calls waiting for shared state, with no fixed waiting capacity. auto_block.active reports the single state owner through firewall operations, state writes and cleanup. lag_basis: operation_progress times the current operation within a batch; advancing batches do not degrade solely because their total duration exceeds a minute. One minute without progress degrades the active row and any waiting callers. A free state slot with no admission for one minute also reports lag. Returned direct-block or flush errors, failed state reads and writes, and abnormal exits count once per call; protected-address refusals do not. Known write failures are recorded before readback and diagnostic output. Known errors remain visible during later cleanup, and the common loss threshold and recovery policy apply. These rows count state-lock callers, independently of persisted per-IP retry records. Health snapshots use memory only and cannot wait for the state lock or I/O.

auto_block.pending reports durable retry records, including records loaded at startup when automatic blocking is disabled. Its capacity is the retry admission limit of 1,000 records. Records selected for a cycle remain in flight until their state-file outcome is known; newly requeued records also remain in flight until persistence finishes. auto_block.candidates reports distinct IPs admitted to the current cycle, with no fixed capacity. Candidate and record counts describe different stages and must not be added together. A candidate remains visible through its firewall callback, subsequent bookkeeping and durable requeue.

Pending age uses its original queue timestamp, or first observation when that timestamp is absent. An eligible record receives a timestamp on its first requeue; records without check identity are withdrawn under the current block policy. Refreshing the reason, check or severity does not reset the original queue age. Normal quota waits stay healthy within the existing two-hour retry lifetime; older waiting records report backlog_lag. Active record and candidate work use one minute without operation progress, so advancing batches can run longer without a false warning. These measurements add no retry scheduler and do not change the hourly quota, expiry or overflow policy.

An unsuccessful attempt whose retry survives reports retry_failed without a loss. Successful blocks, dry runs and expected refusals remain completed even if later bookkeeping fails. Withdrawal under the current check policy or login-blocking setting is an expected refusal and adds no loss. Persisted record identities include check and severity, so a refused record cannot acknowledge a different eligible retry after a failed write. Confirmed removal of expired, invalid or overflowed records counts as pending loss; a fresh candidate that neither completes nor survives on disk counts as candidate loss. Existing duplicate coalescence adds no loss while a retry survives. The common loss threshold and recovery policy apply.

State read or write failures report state_io. Failed writes are read back because an error after rename can leave the new state in place. If that read also fails, depth_unavailable and dropped_lower_bound expose the uncertainty. A later successful read restores measured depth; lifetime loss totals remain lower bounds when earlier outcomes could not be established. New work with known outcomes still contributes exact losses. Completed candidate history is released after settlement, and health snapshots never read files or wait for firewall callbacks.

auto_block.cleanup reports distinct IPs awaiting bookkeeping cleanup after a successful firewall flush. It observes saved cleanup retries at startup and admits the flushed engine entries before reading the tracker. Cleanup also includes tracked blocks, with duplicate IPs counted once. Work remains in flight through store removal, threat-record cleanup and the final state-file outcome. This row has no fixed capacity and is independent of the block-candidate stages.

Cleanup waiting age starts at first observation with lag_basis: deferred_checkpoint. Age alone does not degrade health: cleanup retries run on the next explicit flush, without a background retry deadline. One minute without active operation progress reports processing_lag. Failed cleanup whose retry survives reports retry_failed; an old tracker block can retain a retry even when saving its cleanup marker failed. Completed cleanup stays acknowledged across failed saves, while a newly blocked IP creates fresh cleanup demand. Recovery first establishes the known block baseline. Cleanup of a later block has an exact outcome even when older cleanup history remains uncertain.

Cleanup state and snapshot failures report state_io. An unreadable tracker sets depth_unavailable and dropped_lower_bound; speculative waiting records are not reported as measured depth. The last bounded batch retains its identity for later read recovery. A subsequent unreadable flush replaces that history with its own batch. A readable state restores measured depth and counts confirmed unfinished work that no longer has a retry source. An incomplete pre-flush engine snapshot marks lifetime losses as a lower bound without inventing a missing count. Health reads memory only. These measurements preserve the existing flush policy, returned errors and cleanup retry behavior.

attackdb.events reports events awaiting persistence, including historical imports. The buffer has no fixed capacity. Waiting age starts at admission, independently of an event’s timestamp. A detached batch remains in flight through writes, close and error logging; newly arriving events stay in the waiting count. Processing age uses lag_basis: operation_progress and resets when a write returns, so a progressing batch does not appear stalled solely because of its total duration. Waiting age or absent write progress of one minute degrades health. Completed JSONL records accepted by the file writer and successful database writes count as persisted. Buffered data alone does not. Confirmed unwritten events count as losses, including interrupted work left buffered or unencoded, and returned encoding failures. These losses are visible before cleanup; close errors or abnormal I/O exits preserve uncertainty with dropped_lower_bound and report persistence_uncertain for one minute, ahead of a backlog or a stalled write, as the record queue already does. The common loss threshold and recovery policy apply to confirmed losses. These measurements preserve the existing write, retention and shutdown policy; they do not add retries or promise storage durability beyond the writer’s result. Memory-only databases have no persistence backlog. Health reads queue memory, independently of database state locks and I/O.

attackdb.records reports distinct IP records awaiting a saved update or deletion. Repeated changes to one IP coalesce before the snapshot; a mutation during an active write remains separate pending work. A successful deletion can also satisfy a repeated deletion. Loaded records needing normalization and expired records needing removal enter the same queue. Retained scoring records are not counted as pending persistence. There is no fixed capacity.

Waiting age starts at the first pending change. The active snapshot stays in flight through writes, error logging and retry bookkeeping. Processing age uses lag_basis: operation_progress and resets at each returned store write or delete, or at the complete flat-file write result. One minute without progress or with waiting work degrades health. Returned failures report retry_failed while the retry remains queued, preserving its original age without counting it as lost. Successful retries clear that condition. A failed shutdown flush retains the in-memory retry; the stopped background saver does not schedule another attempt.

An interrupted operation counts confirmed abandoned demand only when no pending update or deletion survives. Unreturned write outcomes report persistence_uncertain for one minute and retain dropped_lower_bound afterward. Known losses remain counted through recovery. Health reads queue metadata only, without waiting for database state or storage locks. These measurements preserve the existing persistence, retry and shutdown behavior; they do not make retained in-memory retry state durable across process exit.

block_digest.records reports the actual buffer of watched-country block records, bounded to 5,000 with drop-oldest overflow. Age starts when a record enters the buffer, independently of the block timestamp. Normal batching and a full buffer alone do not report a stall: waiting becomes overdue one minute after the configured interval. Detached digest preparation remains in flight; one minute without progress reports processing_lag. The row uses lag_basis: operation_progress. Records arriving during preparation belong to the next window. Expected per-IP coalescing and delivery filtering are successful completion, while abandoned preparation counts its discarded records as losses.

block_digest.email and block_digest.webhook report configured destinations in notifications, with capacity_unavailable. Each eligible destination owns one notification before delivery starts, including live alerts and configured empty heartbeats. Default delivery follows current alert settings at admission, attempt and interrupted cleanup. Disabled destinations add no new waiting work or loss; destinations enabled during an earlier send are still handled. Explicit destination selection keeps its error when that alert channel is disabled. One minute waiting or active reports lag. A returned sink error counts one lost notification and stays in flight through error logging. If an operation exits without returning, its attempted delivery reports delivery_uncertain for one minute and retains dropped_lower_bound; a remaining destination that was never attempted counts one confirmed loss. Known losses use the common warning threshold and survive recovery in lifetime totals.

These observations preserve the existing digest schedule, live deduplication, best-effort delivery and shutdown policy. Records retained after the ticker stops report consumer_stopped and can still be explicitly flushed; no retry or durability is added. Health reads metadata independently of collector state locks, country lookups and delivery callbacks. Disabled collectors have no rows.

state.pending reports findings parked for the next startup, bounded to the newest 10,000 findings. The store observes existing parked work when opened. Depth is the last confirmed file contents; in-flight findings include incoming appends and cleared batches still in replay. The bound applies to the stored batch, not concurrent callers or an already detached replay. Overflow and new findings that fail to reach storage count as losses. Duplicate occurrences remain separate findings.

Waiting age starts at admission or startup observation, independently of the finding timestamp. Surviving findings keep their age through later appends; evicted findings no longer determine it. Waiting uses lag_basis: deferred_checkpoint: a parked or full batch awaiting restart does not itself report a stall. state.pending_operations measures callers waiting for the state lock and active persistence or replay, in operations with no fixed bound. One minute waiting or without operation progress reports lag. Actual read, write and clear results advance progress; replay stays active through dispatch.

Read, write and clear errors report state_io. Failed mutations are read back: a returned error after replacement does not imply the new findings were lost. Readback compares the recovered JSON payload, including its normal repair of invalid UTF-8, so repaired log text does not invent uncertainty or hide overflow. Unreadable outcomes set depth_unavailable and dropped_lower_bound. A later read restores measured depth, while lifetime losses remain lower bounds. Known encoding failures and overflow are counted even when other outcomes are unknown. Indistinguishable retained payloads keep the earlier age after a failed write. An interrupted write or clear marks disk state unknown before releasing the state lock. Interrupted replay retains measured disk depth but cannot claim an exact loss for a partially dispatched batch; uncertainty reports persistence_uncertain for one minute. Confirmed losses use the common warning threshold and remain in lifetime totals after recovery.

The file format, append order, returned errors and clear-before-replay policy are unchanged. These observations add no retry or extra replay. Health snapshots read metadata only and do not wait for state locks, files or dispatch callbacks.

incident.persist.waiting reports immutable incident snapshots waiting for the ordered writer, with no fixed waiting capacity. incident.persist.active reports the single occupied writer. A writer or free-slot admission stalled for one minute reports lag. Failed writes and abnormal callback exits count once; abandoned bulk snapshots count as waiting losses and release their ordering slots for later writes. Callbacks remain owned through cleanup and error logging. Store failures retain in-memory transitions and the existing warning log. incident.persist.deferred reports coalesced bookkeeping waiting for a later mutation or explicit flush, including shutdown. Its oldest age uses lag_basis: deferred_checkpoint; age alone does not degrade health because there is no periodic flush deadline. Full snapshots supersede earlier bookkeeping; restoration and retention discard the affected markers. Memory-only correlators have no persistence work. All three rows read memory independently of state locks and database I/O. The common loss threshold and recovery policy apply to writes.

checks.file_index.waiting reports live scans waiting for the shared baseline slot, with no fixed waiting capacity. Each wait keeps its shorter parent deadline or the file-index check budget; a free slot with no admission for one minute also reports backlog lag. checks.file_index.active reports the single occupied slot. Walking and file analysis use the check’s original execution deadline (normally 15 minutes) or the shorter parent deadline. Setup, persistence and cleanup each have a separate one-minute budget. A healthy long scan alone does not report a full queue. Actual filesystem work remains visible after cancellation until it returns. Deadline withdrawal, incomplete walks, unreadable PHP content, failed executable metadata or state reads and writes, and abnormal exits count once per live scan. Executable entries disappearing during enumeration and explicit cancellation alone add no loss. The common three-loss warning threshold applies, and recovery retains total losses. A successful late baseline commit adds no loss. Findings and shrink protection keep their existing behavior. Force-file-index audits bypass this stateful slot and remain owned by the check execution row. Snapshots read memory only. checks.reputation_queries publishes accepted lookups immediately after quota reservation, including while refused lookups receive fallback scoring. It follows the accepted work through HTTP, response cleanup, buffered results, supplemental scoring and cache or quota-backoff storage. Each scan has at most five requests; concurrent scans have no fixed combined capacity. HTTP uses the client’s timeout, while cleanup, local result handling and persistence have separate one-minute budgets. Supplemental scoring uses the combined timeout of its enabled sources or the shorter parent deadline. Buffered results follow their own active consumer’s deadline; another scan cannot hide their delay. Query failures, failed storage and abnormal exits count once per result. Quota responses and refusal before dispatch remain expected outcomes with the existing quota health warning. Parent cancellation does not cancel the existing HTTP requests; they remain visible until response handling and storage finish. Successful late results add no loss. Health reads memory only and retains losses after recovery. Local feed matches, cache hits and serial pre-query discovery do not create query jobs. A failed cache write retains the findings already produced. checks.dispatch reports pending checks and occupied runner wrappers across host and account scan batches. Concurrent batches have no fixed global waiting limit, so capacity is unavailable. Lag uses consumer_progress within each batch: pending work degrades after one minute without progress when that batch has a free worker slot. A busy pool may keep working within each check’s own deadline. Setup and result handling have a separate one-minute budget; they cannot borrow a heavy check’s longer deadline. A wrapper panic or abnormal exit counts a loss, as does a parent deadline before a check can start. Explicit cancellation withdraws queued demand without loss. Execution failures belong to checks.executions; dispatch tracks the surrounding scheduling operation. Reading status never takes a scan, context or database lock. scans.jobs reports eight waiting full-scan jobs and one worker. Work remains visible through enumeration, scanning, remediation, persistence and cleanup. Its consumer_progress lag follows that worker and its own check batches: healthy check deadlines allow long scans to proceed, while an overdue child cannot be hidden by another child’s progress. Enumeration, result handling, individual file actions and database operations each have a one-minute progress budget. Thirty seconds at full waiting capacity also degrades health. scans.admission separately reports callers waiting for admission and writing their initial job record, with a one-minute lag budget and unavailable capacity. Status reads only memory; operator polling does not advance the worker’s clock. Queue refusal, failed job operations and abandoned work count once per job. Configured finding-history truncation, explicit cancellation and successful shutdown draining do not count as loss. A terminated worker refuses new jobs. At startup, both queued and running records from the previous process become errors with reason daemon_restarted; their lost requests count in job health. email_av.scans reports antivirus engine work through execution and both result handoffs. Timed-out engines remain in flight until they return; buffered results remain queued until the message scan consumes them. Execution lag uses the original per-part deadline, while result waiting and processing each have a one-minute budget. Moving a result between buffers preserves its waiting age. Concurrent messages and engines outliving timeouts have no fixed global limit, so capacity is unavailable. Timeouts, errors and abandoned work count once per engine scan; later failures cannot count the same scan twice. Detections and unavailable engines are normal outcomes, and mail verdict behavior is unchanged. Status reads only memory, and watcher restarts reuse the same engine health. php_taint.requests follows callers waiting for the worker lock, active requests and pipe operations that outlive their callers. Waiting, setup, reply handling and cleanup each have a one-minute lag budget. Active worker communication uses its configured timeout or the shorter caller deadline; a buffered reply cannot borrow a long worker timeout. The request stays in flight through reply decoding and cleanup, and until any outstanding pipe operation finishes. Capacity is unavailable because callers have no fixed global waiting limit. Worker failures, timeouts, breaker refusals and abnormal exits count once per request; known failures are recorded before cleanup. Oversize input, caller cancellation and refusal after an intentional stop do not add losses. Status reads memory without the worker lock, and shutdown retains the supervisor’s health evidence. Recovered analyzer panics also count as failed work when delivered in a valid worker reply. The report and worker reuse policy remain unchanged. central.actions reports 1,024 waiting central-intelligence actions and one running action. Backlog remains visible while the signed feed refreshes; processing time includes the action handler and its evidence delivery. Overflow, action failure and abandoned shutdown work count as losses. A protected-address refusal or absent firewall engine completes the queue task without claiming that a block happened. Challenge tasks measure delivery to the challenge list; they do not measure its later file or firewall writes. Shutdown cancels feed refreshes, waits for an already-running action and discards the remaining queue. Captured dispatch hooks refuse later work and preserve its loss count after shutdown. Logging does not reset loss totals. bot_verification.requests reports 256 waiting bot-identity requests and one running verification. Duplicate requests share their original waiting age. Running time includes DNS lookups and the cache write. Queue overflow, DNS timeouts or transient failures, failed cache or missing-PTR record writes and abandoned shutdown requests count as losses. Missing PTR records, unknown bot identities and definitive positive or negative answers remain expected outcomes. A source with a recorded missing PTR is not queued again until its one-hour suppression lapses; suppressed requests do not count as losses. For a day after the lookup, lapsed records prevent renewed pending grace without blocking DNS retries. Queue health does not turn a resolver failure into a spoof finding. Shutdown cancels DNS, waits for any active cache write and discards waiting work. Late submissions are refused, pending keys are released and loss totals remain available. abuse_reporting.ingress measures reports awaiting durable storage, with a capacity of 10,000 or the configured spool limit when smaller. Running time includes persistence to every configured target. Memory overflow, incomplete persistence and work abandoned by a worker failure count as losses; one source report counts once even when several targets fail. Outbound delivery can delay memory work, and that backlog remains visible. The queue stays open through the daemon’s final finding flush. Reporter shutdown then closes admission and persists every accepted report before closing the spool; a failed write does not discard unrelated waiting reports. Captured hooks refuse later reports. This row measures the memory queue, not reports already retained on disk for retry; normal shutdown does not count those durable reports as lost. abuse_reporting.spool measures durable reports per destination, with the configured spool capacity. Reports already being sent remain in flight through the database acknowledgment. Failed admission and discarded records count as losses; a delivery retry, failed acknowledgment or normal shutdown retains the report and does not count it as lost. An evicted record counts as lost only if no send was acknowledged during this process. An active send settles that count when it finishes; a failed retry cannot undo a prior acknowledgment. lag_basis: observed_age means waiting age starts when this process first sees the record. Existing records start at spool open; retries keep that age. Doctor labels this as observed_lag, since time spent waiting before restart is unknown. Waiting or running work degrades after two minutes, allowing for the normal one-minute delivery interval. The common full-queue and recent-loss thresholds also apply. spool_io names failed admission, read or acknowledgment operations; delivery_failed names a failed or panicking sender. These states clear when the affected operation recovers. Health reads remain independent of database writes and outbound requests. Reopening with a smaller cap preserves existing records until the next admission applies the configured overflow policy. phpanel.spool measures the durable panel webhook queue, capped at 100,000 records per state directory. In-flight work includes writes waiting to commit and deliveries awaiting database acknowledgment. Retried findings keep their original waiting age; a failed attempt is not a lost finding while its record remains durable. An overflowed record counts as lost only when no send was acknowledged during this process; an active send settles that count when it finishes. Malformed findings count once when removed from delivery; their bounded diagnostic archive is retained history, not pending work. Waiting age uses the persisted enqueue timestamp after restart. Boundary records absent from the accounting rebuilt at open are adopted with the same timestamp, so retrying them does not reset their age. Missing, damaged or future timestamps are timed from queue open or adoption. A minute of waiting or processing degrades the row; delivery and database errors remain visible until the affected operation succeeds. Health reads use memory only, so a stalled database or collector cannot block status. Disabling delivery preserves the durable backlog and cumulative loss; enabling it resumes the stored work. Stopped queue instances refuse late admissions. actionlog.writes measures the 64 process-wide action-log write slots. A sink write that outlives the caller’s 250ms wait budget stays in flight until the sink and any panic reporting finish; a batch of actions serialising behind one file lock is normal, and only a minute without the write returning degrades the row. A caller deadline after admission does not count as loss, since the sink may still record the action. Refused admission, sink errors and panics count as lost records, once per record. Changing or disabling the sink preserves outstanding work and cumulative loss. Health reads do not wait for the sink. Normal writes still complete before the caller returns; saturated or stalled recording keeps the existing caller budget. events.deliveries aggregates live event-stream subscribers in one row without client identities. Capacity sums their buffers (64 findings per daemon stream); depth counts waiting deliveries and in-flight work includes encoding, writing and flushing. One full subscriber can degrade this row even while others drain. The common one-minute lag, 30-second fullness and recent-loss thresholds apply. Overflow, encoding errors and failed streams count as losses; a failed stream also counts its abandoned buffered events. Normal request cancellation and server shutdown withdraw pending demand without adding losses. Cumulative loss survives subscriber removal, and an outstanding write remains visible until it returns. Closing the bus preserves buffered work for consumers still draining it. HTTP writes retain their three-second deadline; successful delivery here means the write and flush returned, not that the remote application acknowledged it. Each active BPF backend also exposes a .kernel row. depth_unit: bytes labels ring occupancy and capacity. lag_basis: consumer_progress means lag_seconds measures time without observed consumption while data remains, starting at the first sample that sees pending data. It is not an event age. One minute without progress degrades the row; recently observed reservation failures use the same loss threshold as userspace queues. Kernel loss counters are separate from decode failures and userspace admission loss. After a reader unmaps its ring, depth_unavailable: true and lag_basis: unavailable prevent zero fields from claiming an empty live ring. The last counter sample adds submitted records the reader never consumed. Submission accounting precedes publication, so a fast reader cannot consume an event before it has been counted. dropped_lower_bound: true marks this final shutdown total: kernel detachment can leave callbacks finishing after the sample. Doctor prints dropped>=... for that bound. A counter lookup failure or an incoherent occupancy reading marks the measurement unavailable and degrades the row as measurement_unavailable once it lasts half a minute, so one artefact of reading a live ring raises nothing; a failed final sample degrades at once and remains degraded. A stopped reader is reported instead of the measurement artefacts it causes. Invalid occupancy cannot inherit fullness or consumer stall alarms from an earlier sample; independently counted losses remain visible. The required kernel suite fills the shipped connection program’s ring with real non-root connect calls and checks reservation loss and retained output. fanotify.kernel and spool.kernel report pending notification records and use the same consumer-progress lag measurement. Their group capacity is not exposed by the kernel: capacity_unavailable: true marks it as unknown, and doctor prints depth=N/unknown records. The current system queue limit may differ from the limit captured when the watcher was created. Their loss totals are lower bounds: each overflow record proves at least one loss, and shutdown adds the records known to be unread before closing the descriptor. Events can still arrive between that sample and close. An unavailable pending-record measurement degrades the row once it lasts half a minute; the reading taken at close degrades at once and remains degraded. Records the kernel has already dropped are reported ahead of an unreadable depth. Closing a descriptor leaves known zero occupancy. fanotify.reader and spool.reader track batches after a kernel read, until all records have been dispatched or filtered. A stalled batch remains visible even when the kernel queue is empty. Reader losses count batches interrupted by consumer failure, separately from kernel-record losses. Spool replacements retain loss and running-batch evidence, while the new descriptor starts its own occupancy and progress measurements. Health reads, event reads and permission responses cannot use a descriptor after close. The required kernel suite verifies pending records, stalls, drain recovery and unread shutdown loss using real fanotify events.

forwarder.kernel and phprelay.kernel report inotify backlog in bytes, because records include variable-length filenames. Capacity is unknown and lag follows observed consumer progress. Overflow markers each prove at least one lost event. Closing with unread bytes adds one further known loss; the byte count does not reveal the number of discarded events. These totals use dropped_lower_bound: true. forwarder.reader and phprelay.reader track the running read batch, including synchronous callbacks. Failed callbacks count as interrupted batches. PHP relay replacements retain earlier losses with fresh occupancy measurements. Descriptor reads, watch changes, polling and close share one lifetime guard.

phprelay.index.persistence reports the message attribution writer’s 4,096 waiting slots. Writes remain running from channel receipt through the pending batch and its database transaction. Batches contain at most 256 writes, also during an explicit flush. Each flush covers the backlog present when the writer accepts it; concurrent arrivals cannot extend that flush indefinitely. Refused submissions, encoding failures and writes in a failed transaction count as losses. Failed writes are settled before emitting their error finding, so a blocked reporter cannot hide the loss. A failed transaction counts each of its writes once and does not prevent subsequent batches from committing. Shutdown closes admission and drains accepted writes; later submissions count as refused work. The existing persistence dropped metric counts refused submissions, while queue health also includes writes that fail after admission.

Some rows carry advisory. Their work is best effort: a client that stops reading its event stream, an unreachable panel asked for an optional annotation, and expired process context reads all lose detail around findings that are still detected, stored and delivered. An advisory row reports its own degradation with the same evidence, but leaves the host status and security posture unchanged, warns instead of failing csm doctor, and raises no notification.

Each degraded row names its condition: backlog_lag for the oldest waiting item past its budget, processing_lag for the oldest running item, queue_full for capacity held continuously, dropped_work for recent losses, consumer_stalled for a sampled queue whose consumer stopped advancing, reader_stopped for a released kernel descriptor, measurement_unavailable for a reading that stayed unavailable, and persistence_uncertain for a write whose outcome is unknown. The owner-specific conditions retry_failed, source_io, spool_io, state_io, delivery_failed, delivery_uncertain and consumer_stopped are described with their rows above.

Doctor names each age by what it measures: observed_lag for an age taken at observation, consumer_stall for consumer progress, operation_stall for the current operation, deferred_age for work parked until a restart, and lag=unavailable where no age exists.

A queue becomes degraded after three losses in a minute, thirty seconds continuously full, or a minute waiting or processing. These are operational alert budgets, not measured throughput guarantees. Health is computed directly from the counters, independently of finding delivery. The daemon records and dispatches protection_queue_degraded at most once per queue every five minutes while pressure remains, then one protection_queue_recovered event. The five-minute bound spans recoveries, so a queue that clears and degrades again inside the window stays degraded in status without a second event, and no recovery event follows a degradation the bound suppressed. These are CSM health events and do not feed account-compromise correlation or automatic response. Recovery preserves cumulative loss evidence; restarting a spool watcher also preserves it. Restarting the daemon resets the counters.

Inspect worker errors and CPU, memory and I/O pressure when a queue degrades. Reduce competing bulk work and confirm the queue drains and recent losses stop. This surface currently covers finding delivery, file and spool kernel readers and scanners, recovery scans, staged package verification, dropper processing, BPF queues, process context enrichment and deadline reads, mail-log delivery, forwarder and PHP relay notification queues, and PHP relay index persistence; other bounded queues remain in the roadmap.

GeoIP

GET  /api/v1/geoip               IP geolocation (?ip=&detail=1)
POST /api/v1/geoip/batch         Batch GeoIP lookup (body: {"ips":["192.0.2.1"]}, maximum 500)

Threat Intelligence

GET  /api/v1/threat/stats        Attack stats, type breakdown, hourly trend
GET  /api/v1/threat/top-attackers Top attacking IPs with GeoIP (?limit=)
GET  /api/v1/threat/ip           IP threat lookup (?ip=)
GET  /api/v1/threat/events       IP event history (?ip=&limit=)
GET  /api/v1/threat/whitelist    Whitelisted IPs
GET  /api/v1/threat/db-stats     Attack database statistics
POST /api/v1/threat/block-ip     Block IP for 24 hours
POST /api/v1/threat/block-ip-permanent Block IP with no expiry
POST /api/v1/threat/whitelist-ip       Permanent whitelist
POST /api/v1/threat/temp-whitelist-ip  Temporary whitelist (with expiry)
POST /api/v1/threat/clear-ip           Clear IP from attack database
POST /api/v1/threat/unwhitelist-ip     Remove from whitelist
POST /api/v1/threat/bulk-action        Bulk block (24h or permanent) / whitelist across many IPs

The 24 hour endpoint returns 409 Conflict when the address already has a permanent or longer firewall block. Unblock it explicitly before shortening its lifetime. Bulk actions use action: "block", "block_permanent", or "whitelist"; refused blocks are excluded from count and explained in warnings. Repeated canonical IPs count once. Permanence cannot be selected through an extra request field.

Bulk undo restores the original per-IP firewall deadlines. A later change to any target invalidates that undo action; it cannot restore dismissed evidence or replace the later decision.

Firewall

GET  /api/v1/firewall/status         Config, blocked/allowed counts
GET  /api/v1/firewall/allowed        Whitelisted IPs
GET  /api/v1/firewall/subnets        Blocked subnets
GET  /api/v1/firewall/audit          Firewall audit log
GET  /api/v1/firewall/check          Check if IP is blocked (?ip=)
POST /api/v1/block-ip                Block an IP
POST /api/v1/unblock-ip              Unblock an IP
POST /api/v1/unblock-bulk            Bulk unblock IPs
POST /api/v1/firewall/allow-ip       Allow an IP
POST /api/v1/firewall/remove-allow   Remove IP from allow list
POST /api/v1/firewall/deny-subnet    Block subnet
POST /api/v1/firewall/remove-subnet  Remove subnet block
POST /api/v1/firewall/flush          Clear all blocks
POST /api/v1/firewall/unban          Unblock IP + flush cphulk
POST /api/v1/firewall/cphulk-clear   Flush cphulk bans only

The audit log reports each timestamp as an RFC 3339 instant in UTC and the block or allow lifetime as duration_seconds. Blocked addresses, subnets and allow rules carry expires_at, left out for a permanent entry.

ModSecurity

GET  /api/v1/modsec/stats              WAF statistics (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET  /api/v1/modsec/blocks             Blocked requests log, aggregated per IP, with resolved source country (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET  /api/v1/modsec/events             WAF event details with resolved source country (read scope). Accepts ?window=1h|6h|24h, ?severity=warning|high|critical.
GET  /api/v1/modsec/rules              Loaded rules list
POST /api/v1/modsec/rules/apply        Apply the set of disabled rules and reload
GET  /api/v1/modsec/rules/escalation   Rule IDs excluded from firewall escalation, sorted
POST /api/v1/modsec/rules/escalation   Exclude one rule from escalation or turn it back on

POST /api/v1/modsec/rules/escalation takes {"rule_id": 900112, "escalate": false}. The rule ID must be a CSM rule (900000-900999). escalate: false adds the exclusion and escalate: true removes it; other excluded rules are left as they are.

A failed rules reload reports rolled_back: true only when restoring the previous overrides succeeds. A rollback failure is reported in the response and audit log; inspect the overrides before trying another reload.

Rules & Suppressions

GET  /api/v1/rules/status        YAML/YARA rule counts, version
GET  /api/v1/rules/list          Rule files
GET    /api/v1/suppressions      Suppression rules
POST   /api/v1/rules/reload      Reload signature rules from disk
POST   /api/v1/suppressions      Add a suppression rule
DELETE /api/v1/suppressions      Delete a suppression rule by id

POST /api/v1/suppressions takes {"check", "path_pattern", "reason"}. The path pattern is a glob and must be valid. A rule that covers every path of a check, hiding all its findings and stopping their remediation, needs "all_paths": true and no path_pattern; an empty pattern without it returns 400. check must be a check name, not a pattern: letters, digits, _, ., : and -. A name that is neither a known check nor the check of a current finding is saved, and the response carries a warning saying the rule matches nothing yet.

Email

GET  /api/v1/email/stats         Email scanning statistics
GET  /api/v1/email/forwarders    Mail forwarder inventory with destination providers and local-copy flags (read scope)
GET  /api/v1/email/deferrals     Outbound deferral rollup by provider and sending IP with reason codes, parsed from exim_mainlog (read scope)
GET  /api/v1/email/queue-composition  Mail queue makeup: real vs null-sender bounce backscatter, frozen count, oldest age, top stuck recipients (read scope)
POST /api/v1/email/queue/flush-backscatter  Request removal of frozen null-sender messages from the exim queue on cPanel hosts; returns the count of targeted messages no longer queued, reports incomplete verification as 500, or returns 503 when unavailable (admin scope, CSRF)
GET  /api/v1/email/held          Forward copies held by the forward guard (admin scope)
POST /api/v1/email/held/{id}/release   Re-inject a held forward copy to its external recipient (admin scope, CSRF)
DELETE /api/v1/email/held/{id}   Discard a held forward copy (admin scope, CSRF)
GET  /api/v1/email/groups        Server-grouped action rows (kind=compromised_account|spam_outbreak|auth_failure|queue_alert|malware) with from/to/limit (read scope)
GET  /api/v1/email/relay-abuse   Outbound PHP-mail abuse detections (spam outbreaks, high-volume scripts/accounts) with per-site script breakdown; from/to/limit (read scope)
GET  /api/v1/email/quarantine    Quarantined email list
GET  /api/v1/email/av/status     Email AV watcher status
GET  /api/v1/email/quarantine/{id}          One quarantined message
POST /api/v1/email/quarantine/{id}/release  Release a quarantined message (admin scope, CSRF)
DELETE /api/v1/email/quarantine/{id}        Delete a quarantined message (admin scope, CSRF)

Hardening

GET  /api/v1/hardening           Load last hardening audit report (admin scope)
POST /api/v1/hardening/run       Run hardening audit and save report (admin scope, CSRF)

Scan Jobs

csm scan --full enqueues full-scan jobs that run inside the daemon and persist to the store. Jobs are report-only unless an account-scope request sets quarantine: true; server-wide jobs reject quarantine.

GET  /api/v1/scan-jobs              List full-scan jobs (read scope)
GET  /api/v1/scan-jobs/{id}         Job status and stored report (read scope)
GET  /api/v1/scan-jobs/{id}/findings
                                      Paginated findings for one job (?offset=&limit=, limit 500 by
                                      default and at most 5000; truncated marks more pages) (read scope)
POST /api/v1/scan-jobs              Enqueue a full-scan job (admin scope, CSRF)
POST /api/v1/scan-jobs/{id}/cancel  Cancel a queued or running job (admin scope, CSRF)

Verified Bots

GET  /api/v1/verified-bots        Configured verified-crawler allowlist plus live verification state (admin scope)
POST /api/v1/verified-bots/apply  Validate, apply, and reload an edited verified-bots list (admin scope, CSRF)

Actions

POST /api/v1/fix                      Apply fix for a finding
POST /api/v1/fix-bulk                 Bulk fix multiple findings
POST /api/v1/dismiss                  Dismiss one finding {key} or up to 500 {keys}; returns undo_token
POST /api/v1/scan-account             On-demand account scan
POST /api/v1/verify-finding           Re-check a single finding on demand (admin scope, CSRF)
POST /api/v1/quarantine-restore       Restore quarantined file
POST /api/v1/quarantine/bulk-delete   Bulk-delete quarantined files; returns count and the ids it could not delete
POST /api/v1/db-object-backup-restore Restore a dropped MySQL object from its db_object_backups record
POST /api/v1/test-alert               Send test alert through all channels
POST /api/v1/import                   Import state bundle (suppressions, whitelist)

fix and fix-bulk act on the file the stored finding names. A request may repeat that path in file_path, but a different path is refused, and a target is never a remediation root itself (/home, /tmp, /var/tmp, /dev/shm) or an account’s home directory.

verify-finding returns the verifier verdict in checked, resolved, demote, and detail. When that verdict also changes the stored finding, severity_change is demoted or restored. A demote verdict without severity_change means no stored severity changed; for example, the finding was already demoted or a scan replaced its snapshot while verification was running. Callers must not report a state change from the verdict alone.

Settings

GET  /api/v1/settings             List editable config sections
GET  /api/v1/settings/<section>   Read a config section (secrets redacted)
POST /api/v1/settings/<section>   Update a config section (safe fields reload, restart fields queue)
POST /api/v1/settings/restart     Request a daemon restart. Returns 202 with `started_at_token`
                                  for polling until the restarted daemon reports a new marker.
POST /api/v1/settings/firewall/tentative-apply  Save firewall config, restart, and arm rollback timer
GET  /api/v1/settings/firewall/rollback         Read pending rollback state
POST /api/v1/settings/firewall/confirm          Confirm tentative firewall changes
POST /api/v1/settings/firewall/revert           Revert tentative firewall changes now

Sections map to top-level config keys: alerts, auto_response, challenge, reputation, performance, infra_ips, sentry, etc. Writes persist to csm.yaml, re-sign the integrity hash, and hot-reload where possible; restart-required changes are queued for /api/v1/settings/restart. Invalid field values return 422 and do not touch disk.

Fields marked file_only in the schema are shown but refused on write with a 422: anything that names a command, an executable, a file path, a socket or an environment variable the daemon acts on as root can only be changed in csm.yaml. Changing reputation.upstream.url or reputation.rspamd.url also returns 422 unless the same request enters the effective credential for that address again, and is refused outright when the credential comes from an environment variable. A credential entered in the request must match the resulting configuration after conf.d merging; an overridden replacement cannot authorize sending the stored credential to another address. Firewall tentative apply is restart-class by design: it snapshots the previous config, writes the new one, restarts the daemon, and auto-reverts unless the operator confirms before the timer expires.

Operator preferences

Per-operator state (UI density, timestamp display, default auto-refresh, saved filter views) is keyed server-side by SHA-256 of the auth token, so preferences follow the operator across browsers and devices without the daemon ever storing the raw credential. Capability flag: webui.prefs.v1. These endpoints require admin scope because they read or mutate operator-private UI state.

GET    /api/v1/prefs/user        Read this operator's UI preferences
PUT    /api/v1/prefs/user        Replace the prefs blob (CSRF on cookie sessions)
GET    /api/v1/prefs/views       List saved views; `?page=findings` filters by page
PUT    /api/v1/prefs/views       Upsert one view {page, name, params} (CSRF on cookie sessions)
DELETE /api/v1/prefs/views       Delete one view {page, name} (CSRF on cookie sessions)

Response shape for GET /api/v1/prefs/user:

{
  "density":       "comfortable",
  "timezone":      "local",
  "auto_refresh":  "on",
  "table_columns": { "findings-table": ["check","severity","when"] }
}

density is comfortable or compact. timezone is server, local, or an IANA-shaped zone string (e.g. Europe/Bucharest). auto_refresh is on or off. Server-side sanitisation drops any other value. Unset prefs encode as empty strings; the UI applies comfortable, local, and on defaults.

Response shape for GET /api/v1/prefs/views:

{
  "items": [
    {
      "name": "Critical SSH",
      "page": "findings",
      "params": { "severity": "critical", "check": "smtp_bruteforce" },
      "updated": "2026-05-25T21:47:35Z"
    }
  ],
  "total": 1
}

Saved views are operator-scoped and capped at 200 per operator. The saved view collection is stored as one 64 KiB preference blob. page and params keys must be simple identifiers: ASCII letters, digits, underscore, hyphen, or dot, up to 64 bytes. Each view has at most 32 params, and param string values are capped at 256 bytes. name must be 1-80 bytes with no control characters. PUT and DELETE return "ok": true on success.

Bulk-action undo

Finding dismissals, bulk threat block / whitelist and bulk firewall unblock responses return an undo_token when the daemon queues an inverse operation server-side for 30 seconds. The UI surfaces a banner with the same TTL; CLI callers can act on the token through the endpoints below. Each successful undo writes an undo_<original_action> audit entry. Capability flag: webui.undo.v1. These endpoints require admin scope because they read or mutate operator-private action state.

GET  /api/v1/undo/pending    Latest pending undo entry for this operator (empty object if none)
POST /api/v1/undo/run        Consume an entry and dispatch its inverse {id}; empty id uses latest

Non-empty response shape for GET /api/v1/undo/pending:

{
  "id": "188d1f2a6c8b0000",
  "action": "threat_bulk_block",
  "inverse": "threat_bulk_unblock",
  "summary": "Blocked 2 IPs",
  "recorded_at": "2026-05-26T00:07:09Z",
  "expires_at": "2026-05-26T00:07:39Z"
}

POST /api/v1/undo/run returns {ok, action, inverse, count} on success, or 410 Gone when the entry is missing, already consumed, or past its 30-second TTL. Recognised inverse action keys are threat_bulk_unblock, threat_bulk_block, threat_bulk_unwhitelist, threat_bulk_whitelist, firewall_bulk_reblock, and finding_undismiss. Other bulk actions (quarantine delete, generic fix) do not surface an undo token because they have no clean inverse.

Dismissals accept at most 500 keys per request and deduplicate repeated keys. Undo skips findings changed by a later dismissal, successful re-check or baseline reset, and count reports only the dismissals it actually reversed. Newer scan copies remain in the current findings list.

Finding fields

Every finding in /api/v1/findings, /api/v1/events, and the JSONL audit log carries optional correlation fields when CSM can attribute them:

FieldMeaning
tenant_idTenant attribution from the verdict callback or panel-side webhook reply
domainDomain associated with the event (e.g. PHP-relay scriptKey host, mailbox domain)
mailboxMailbox attribution (e.g. mail brute-force target, PHP-relay envelope-from)
relay_totalPHP-relay trigger count for the path that fired
relay_breakdownPHP-relay script samples that contributed to the alert, with script key, hit count, last seen time, and a bounded sample subject when available

Fields are omitted when the daemon could not attribute them. Orchestrators should treat absence as “unknown,” not “global.”

Cleanup fields

GET /api/v1/quarantine also powers the Cleanup page’s file-backup list. Entries include:

FieldMeaning
kindquarantine or pre_clean
live_stateoriginal_missing, live_differs, original_not_file, archive_missing, archive_not_file, or unknown. Byte-identical restored entries are hidden.

GET /api/v1/db-object-backups returns restored and restored_at when a captured MySQL trigger/event/procedure/function backup has already been replayed.

Incidents

GET /api/v1/incidents/groups is a read-scope rollup of active incidents by kind and source. It accepts status=active|all|open|contained|resolved|dismissed, kind, limit and offset, allowing a credential spray to render as one row per attacker rather than one row per target. total counts the groups and scanned_incidents the incidents they came from; truncated is true when the scan cap left incidents out.

GET /api/v1/incidents

Returns one page of incidents sorted by updated_at descending, with items, total, offset, limit and status. limit defaults to 50 and is at most 500. status filters by open, contained, resolved, dismissed, or active for open and contained; empty means all.

GET /api/v1/incidents/<id>

Returns one incident by id. 404 if not found.

POST /api/v1/incidents/<id>/status

Body:

{"status": "resolved", "details": "operator-marked"}

Status values: open, contained, resolved, dismissed. Closing an incident (resolved/dismissed) means future findings for the same correlation key start a fresh incident. Reopening an incident binds the same key again. Incident JSON includes correlation_key when CSM has a stored account, mailbox, domain, process, or remote-IP key.