Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Upgrading

Use the signed APT/DNF repository for maintained package upgrades. Standalone deploy scripts require successful detached signature verification for current releases. They verify with OpenSSL 3.0+ where available, otherwise with the installed CSM binary’s csm verify-release, otherwise with python3-cryptography, so upgrades keep working on EL8/CloudLinux 8 (OpenSSL 1.1.1) including the first upgrade to a build that provides csm verify-release. When no verifier is present the upgrade stops; disabling CSM_REQUIRE_SIGNATURES does not enable unsigned current upgrades. Set CSM_VERIFIER_BINARY to select a specific CSM binary. See Release signing for the historical-release exception.

Use the same signed repository that installed CSM:

sudo apt update && sudo apt install --only-upgrade csm   # Debian/Ubuntu
sudo dnf upgrade csm                                     # RHEL family

The package preserves the operator config and state, updates runtime assets, re-signs integrity metadata, and restarts the daemon when it was already active.

Standalone installations

sudo /opt/csm/deploy.sh upgrade

The upgrade refuses to install a release older than the one running. To roll back deliberately, run it with CSM_ALLOW_DOWNGRADE=1.

The helper:

  1. Downloads and verifies the binary and supporting assets
  2. Stages the UI, rules, PAM files, and deploy helper before downtime
  3. Stops the daemon and keeps the previous binary and assets as rollback material
  4. Activates the staged release and rehashes the config once
  5. Restarts the daemon, checks sustained liveness, and runs csm doctor

The health gate waits 20 seconds by default. Set CSM_UPGRADE_HEALTH_SETTLE to an integer from 1 to 3600 seconds for a longer or shorter observation window. Invalid values fail the gate. Doctor warnings do not trigger rollback; failed checks do.

If activation, rehash, startup, or the health gate fails, the helper stops the new daemon, restores the previous binary and assets, re-signs the restored binary hash, and starts the previous version. It exits nonzero and reports where it retained recovery material. If recovery itself fails, it reports an incomplete rollback that needs operator attention.

Optional nightly upgrades

Packages ship a disabled sample at /opt/csm/configs/cron/csm-auto-upgrade. It runs only after an operator installs it into /etc/cron.d/:

csm version && /opt/csm/deploy.sh check
sudo install -m 0644 /opt/csm/configs/cron/csm-auto-upgrade /etc/cron.d/csm-auto-upgrade

Confirm the host runs a published release first. The checksum check reports any different build as an available update. The downgrade guard rejects a lower version, but a development build with the same version can still be replaced by the published release.

The job invokes the standalone deploy helper every night between 03:30 and 04:30, with a random delay and a nonblocking lock. This updates runtime files directly; it does not update APT/DNF’s installed package version. Hosts that need package-manager ownership and version tracking should use their package upgrade automation instead.

Output goes to /var/log/csm/auto-upgrade.log. Failures also produce cron mail to root; configure and test root’s mail alias before enabling the job. This notification does not depend on CSM’s alert channels or on successful daemon recovery. Successful runs and an already-held lock produce no mail. Staggering spreads upgrades over an hour but does not stop a bad release from reaching the fleet.

Troubleshooting

“store: opening bbolt: timeout” – Most operator commands that need live state now route through the control socket at /var/run/csm/control.sock. This error should only appear from commands that intentionally open the bbolt file directly, such as csm store compact, csm store import, csm store reset-bot-verify, csm db-clean --drop-object, or a second daemon start while one daemon already owns the database.

Fix: stop the daemon before direct-store maintenance commands, then retry:

systemctl stop csm
csm store compact
systemctl start csm

If systemctl says CSM is stopped but bbolt still times out, find the process holding /var/lib/csm/state/csm.db and stop that process after review. Do not delete csm.lock; it is only the daemon instance guard and does not release bbolt’s file lock.

“csm: daemon not running” – CLI commands that talk to the daemon exit 2 with this message when the control socket is missing. This includes csm run*, csm check*, csm baseline, csm status, csm firewall ..., csm store export, csm export --since, and csm phprelay .... Start the daemon with systemctl start csm. Bootstrap commands that run before the daemon exists (csm install, csm validate, csm config schema, csm verify, csm rehash) do not require it.

Never delete csm.db – it contains all historical findings, firewall state, email forwarder baselines, and per-account data. If you delete it, the web UI will show empty data until the next full scan cycle (up to 60 minutes for deep scan findings). Restore from backup when possible; for an intentional reset, run csm baseline --confirm rather than removing the database by hand.

Config changes require rehash – After editing a restart-required field in csm.yaml or any conf.d drop-in, run csm rehash once, validate, then restart. Hot-reload-safe changes can use systemctl reload csm; the daemon validates and re-signs the accepted config itself. A restart that fails with conf.d hash mismatch means a drop-in changed since the last signing: rehash if the change was intentional, and list fragments an integration rewrites on its own under confd.integrity_exempt so they stop needing one. csm doctor reports the mismatch while the daemon is still up.

FHS migration (state, config, drop-ins, and profiles)

Current packages use FHS paths for state, config, drop-ins, and shipped profiles. Legacy main configs continue to work during the transition.

ConcernLegacy pathCurrent path
Drop-in fragmentsn/a/etc/csm/conf.d/*.yaml
State directory/opt/csm/state/var/lib/csm/state
Shipped profilesn/a/usr/lib/csm/profiles
Binary/opt/csm/csm/opt/csm/csm (unchanged)
Main config/opt/csm/csm.yaml/etc/csm/csm.yaml
Legacy config pathn/a/opt/csm/csm.yaml symlink

The package postinstall creates the FHS directories with the right ownership. If /opt/csm/csm.yaml is a real file and /etc/csm/csm.yaml is absent or still the shipped placeholder, the package copies the legacy config into /etc/csm/csm.yaml and then replaces the old path with a symlink. If both paths are real files with different operator content, CSM refuses the implicit default path until you move one aside or pass --config <path>. Copies that differ only in the integrity hashes CSM writes itself (binary_hash, config_hash, confd_hash) are treated as the same configuration: the daemon starts from /etc/csm/csm.yaml, and the next csm rehash replaces the legacy copy with the symlink.

The daemon copies a non-empty legacy /opt/csm/state/ into the new state directory on first start, but only when the new directory is empty (so a partial migration cannot corrupt it). The legacy directory is left in place; remove it after you have verified the new install.

Operators upgrading by manual binary swap (without re-running the package postinstall) keep the legacy state path if state_path: /opt/csm/state is pinned in the existing csm.yaml. To move state to the FHS layout, either reinstall the package or create the directories by hand and remove the state_path: override.

systemd Type=notify drop-in

The packaged unit file is Type=notify with WatchdogSec=300. The daemon signals READY=1 after watchers attach and pings WATCHDOG=1 on schedule, so systemctl is-active reflects truth and the watchdog kills a hung daemon.

Older units shipped Type=simple. The watchdog still functions because the daemon pings regardless of unit type, but systemctl status only sees the process, not “watchers attached.” If you need the new behavior on an older unit, drop in:

# /etc/systemd/system/csm.service.d/notify.conf
[Service]
Type=notify
NotifyAccess=main

Then systemctl daemon-reload && systemctl restart csm. Verify with systemctl show csm -p Type -p StatusText.

Auto-response dry-run safety default

auto_response.dry_run defaults to true when the key is absent. The daemon records every IP it would have blocked but does not touch nftables. If your auto_response: block sets enabled: true and block_ips: true but does not set dry_run, add dry_run: false explicitly before relying on auto-block. Verify with:

csm status --json | jq '.auto_response_dry_run, .dry_run_blocks'
csm firewall status            # check that "Recently Blocked" picks up new entries after the restart

Manual csm firewall ... operations bypass dry-run and always apply.