Upgrading
Use the signed APT/DNF repository for maintained package upgrades. Standalone
deploy scripts require successful detached signature verification for current
releases. They verify with OpenSSL 3.0+ where available, otherwise with the installed
CSM binary’s csm verify-release, otherwise with python3-cryptography, so
upgrades keep working on EL8/CloudLinux 8 (OpenSSL 1.1.1) including the first
upgrade to a build that provides csm verify-release. When no verifier is
present the upgrade stops; disabling CSM_REQUIRE_SIGNATURES does not enable
unsigned current upgrades. Set CSM_VERIFIER_BINARY to select a specific CSM
binary. See
Release signing for the historical-release exception.
Package installations (recommended)
Use the same signed repository that installed CSM:
sudo apt update && sudo apt install --only-upgrade csm # Debian/Ubuntu
sudo dnf upgrade csm # RHEL family
The package preserves the operator config and state, updates runtime assets, re-signs integrity metadata, and restarts the daemon when it was already active.
Standalone installations
sudo /opt/csm/deploy.sh upgrade
The upgrade refuses to install a release older than the one running. To roll back deliberately, run it with CSM_ALLOW_DOWNGRADE=1.
The helper:
- Downloads and verifies the binary and supporting assets
- Stages the UI, rules, PAM files, and deploy helper before downtime
- Stops the daemon and keeps the previous binary and assets as rollback material
- Activates the staged release and rehashes the config once
- Restarts the daemon, checks sustained liveness, and runs
csm doctor
The health gate waits 20 seconds by default. Set CSM_UPGRADE_HEALTH_SETTLE to an integer from 1 to 3600 seconds for a longer or shorter observation window. Invalid values fail the gate. Doctor warnings do not trigger rollback; failed checks do.
If activation, rehash, startup, or the health gate fails, the helper stops the new daemon, restores the previous binary and assets, re-signs the restored binary hash, and starts the previous version. It exits nonzero and reports where it retained recovery material. If recovery itself fails, it reports an incomplete rollback that needs operator attention.
Optional nightly upgrades
Packages ship a disabled sample at /opt/csm/configs/cron/csm-auto-upgrade. It runs only after an operator installs it into /etc/cron.d/:
csm version && /opt/csm/deploy.sh check
sudo install -m 0644 /opt/csm/configs/cron/csm-auto-upgrade /etc/cron.d/csm-auto-upgrade
Confirm the host runs a published release first. The checksum check reports any different build as an available update. The downgrade guard rejects a lower version, but a development build with the same version can still be replaced by the published release.
The job invokes the standalone deploy helper every night between 03:30 and 04:30, with a random delay and a nonblocking lock. This updates runtime files directly; it does not update APT/DNF’s installed package version. Hosts that need package-manager ownership and version tracking should use their package upgrade automation instead.
Output goes to /var/log/csm/auto-upgrade.log. Failures also produce cron mail to root; configure and test root’s mail alias before enabling the job. This notification does not depend on CSM’s alert channels or on successful daemon recovery. Successful runs and an already-held lock produce no mail. Staggering spreads upgrades over an hour but does not stop a bad release from reaching the fleet.
Troubleshooting
“store: opening bbolt: timeout” – Most operator commands that need live state now route through the control socket at /var/run/csm/control.sock. This error should only appear from commands that intentionally open the bbolt file directly, such as csm store compact, csm store import, csm store reset-bot-verify, csm db-clean --drop-object, or a second daemon start while one daemon already owns the database.
Fix: stop the daemon before direct-store maintenance commands, then retry:
systemctl stop csm
csm store compact
systemctl start csm
If systemctl says CSM is stopped but bbolt still times out, find the process holding /var/lib/csm/state/csm.db and stop that process after review. Do not delete csm.lock; it is only the daemon instance guard and does not release bbolt’s file lock.
“csm: daemon not running” – CLI commands that talk to the daemon exit 2 with this message when the control socket is missing. This includes csm run*, csm check*, csm baseline, csm status, csm firewall ..., csm store export, csm export --since, and csm phprelay .... Start the daemon with systemctl start csm. Bootstrap commands that run before the daemon exists (csm install, csm validate, csm config schema, csm verify, csm rehash) do not require it.
Never delete csm.db – it contains all historical findings, firewall state, email forwarder baselines, and per-account data. If you delete it, the web UI will show empty data until the next full scan cycle (up to 60 minutes for deep scan findings). Restore from backup when possible; for an intentional reset, run csm baseline --confirm rather than removing the database by hand.
Config changes require rehash – After editing a restart-required field in csm.yaml or any conf.d drop-in, run csm rehash once, validate, then restart. Hot-reload-safe changes can use systemctl reload csm; the daemon validates and re-signs the accepted config itself. A restart that fails with conf.d hash mismatch means a drop-in changed since the last signing: rehash if the change was intentional, and list fragments an integration rewrites on its own under confd.integrity_exempt so they stop needing one. csm doctor reports the mismatch while the daemon is still up.
FHS migration (state, config, drop-ins, and profiles)
Current packages use FHS paths for state, config, drop-ins, and shipped profiles. Legacy main configs continue to work during the transition.
| Concern | Legacy path | Current path |
|---|---|---|
| Drop-in fragments | n/a | /etc/csm/conf.d/*.yaml |
| State directory | /opt/csm/state | /var/lib/csm/state |
| Shipped profiles | n/a | /usr/lib/csm/profiles |
| Binary | /opt/csm/csm | /opt/csm/csm (unchanged) |
| Main config | /opt/csm/csm.yaml | /etc/csm/csm.yaml |
| Legacy config path | n/a | /opt/csm/csm.yaml symlink |
The package postinstall creates the FHS directories with the right ownership. If /opt/csm/csm.yaml is a real file and /etc/csm/csm.yaml is absent or still the shipped placeholder, the package copies the legacy config into /etc/csm/csm.yaml and then replaces the old path with a symlink. If both paths are real files with different operator content, CSM refuses the implicit default path until you move one aside or pass --config <path>. Copies that differ only in the integrity hashes CSM writes itself (binary_hash, config_hash, confd_hash) are treated as the same configuration: the daemon starts from /etc/csm/csm.yaml, and the next csm rehash replaces the legacy copy with the symlink.
The daemon copies a non-empty legacy /opt/csm/state/ into the new state directory on first start, but only when the new directory is empty (so a partial migration cannot corrupt it). The legacy directory is left in place; remove it after you have verified the new install.
Operators upgrading by manual binary swap (without re-running the package postinstall) keep the legacy state path if state_path: /opt/csm/state is pinned in the existing csm.yaml. To move state to the FHS layout, either reinstall the package or create the directories by hand and remove the state_path: override.
systemd Type=notify drop-in
The packaged unit file is Type=notify with WatchdogSec=300. The daemon signals READY=1 after watchers attach and pings WATCHDOG=1 on schedule, so systemctl is-active reflects truth and the watchdog kills a hung daemon.
Older units shipped Type=simple. The watchdog still functions because the daemon pings regardless of unit type, but systemctl status only sees the process, not “watchers attached.” If you need the new behavior on an older unit, drop in:
# /etc/systemd/system/csm.service.d/notify.conf
[Service]
Type=notify
NotifyAccess=main
Then systemctl daemon-reload && systemctl restart csm. Verify with systemctl show csm -p Type -p StatusText.
Auto-response dry-run safety default
auto_response.dry_run defaults to true when the key is absent. The daemon records every IP it would have blocked but does not touch nftables. If your auto_response: block sets enabled: true and block_ips: true but does not set dry_run, add dry_run: false explicitly before relying on auto-block. Verify with:
csm status --json | jq '.auto_response_dry_run, .dry_run_blocks'
csm firewall status # check that "Recently Blocked" picks up new entries after the restart
Manual csm firewall ... operations bypass dry-run and always apply.