Enhanced server status
Use server-status-enhanced.sh. It is a standalone script; no companion Python file or pip packages are needed. It preserves the previous merged script's checks and host notes. Both original attachments and server-status-merged.sh remain unchanged.
bash server-status-enhanced.sh
bash server-status-enhanced.sh staging
bash server-status-enhanced.sh production
bash server-status-enhanced.sh both
The default is automatic host-profile selection, falling back to both profiles for an unrecognized machine. Profiles determine website targets and labels; local checks always inspect the machine running the script.
Added checks
- HTTPS GET responses for every site named in the original host notes, including the four staging sites missing from the original DNS-check array. Reports final status, elapsed time, redirect count, and destination. TLS verification stays enabled; redirects to HTTP are refused. Unexpected cross-host redirects, authentication responses, and responses taking 3 seconds or longer receive warnings.
- Served certificate validation and expiry: warnings at 30 and 14 days; urgent failures at 7 days, expiry, or invalid hostname/trust. Checks the certificate presented on port 443, possibly a CDN/proxy certificate rather than the origin server's local certificate.
- Backup destination readability, optional expected mount presence, last successful backup age, and latest recorded job result.
- Disk-space and inode thresholds: warning at 85%, failure at 95%. Immutable squashfs/ISO images are excluded on Linux because they normally appear full.
- Ubuntu reboot-required marker and requesting packages; RPM needs-restarting fallback when available.
- Systemd timer schedule, triggered service's last result, and failed units. A default “success” with no execution timestamp is treated as unverified. Cron listings are preserved, but generic cron success cannot be inferred without job-specific evidence.
- All Docker containers, including stopped containers, health-check status, restart counts, exit codes, and out-of-memory exits. Running containers without a health check are explicitly unverified.
- MySQL/MariaDB and PostgreSQL
SELECT 1probes with existing client credentials. No password prompt or credential discovery. Access/authentication problems are distinguished from confirmed probe failures. - Last 24 hours of accessible kernel journal entries for out-of-memory kills, disk/filesystem errors, USB resets/disconnects, and thermal throttling.
- Time synchronization via timedatectl or chronyc.
- Tailscale local status and health warnings, plus an optional configured peer ping.
The final Needs attention summary groups FAIL, WARN, and COULD NOT CHECK, followed by totals including OK. It summarizes the added health checks; the original inventory output remains above it and is not automatically included in those totals. Missing tools, permissions, credentials, or configuration do not count as a pass.
Settings
Edit the settings block near the top, or set these environment variables:
| Setting | Default / purpose |
|---|---|
STATUS_BACKUP_DESTINATION | Empty; local or mounted directory containing backups |
STATUS_BACKUP_STATUS_FILE | Empty; JSON status written by the backup job |
STATUS_BACKUP_MOUNT | Empty; expected mount point, if backups use a separate mounted destination |
STATUS_BACKUP_MAX_HOURS | 36; maximum age of the last reported successful backup |
STATUS_TAILSCALE_PEER | Empty; expected peer hostname or Tailscale IP |
STATUS_DISK_WARN_PERCENT | 85 |
STATUS_DISK_FAIL_PERCENT | 95 |
STATUS_SLOW_SECONDS | 3 |
Backup settings should describe the machine on which the script runs. Leave them empty until you know the correct locations. Remote backup APIs and object-storage destinations are not inferred or queried automatically; a local job-result file can still provide evidence for those jobs.
The backup job should maintain a status file with this structure:
{
"last_success_epoch": 1700000000,
"last_run_status": "success"
}
The timestamp above is an example, not a current backup result. The job must write its actual last-success Unix timestamp and update last_run_status to success, failed, or running for each attempt. Preserve the prior successful timestamp when a later attempt fails. Write the file atomically at a stable location readable by the status script. This script never fabricates or updates backup evidence.
A readable backup directory and recent success report do not prove archive integrity or restoreability. Continue testing restores separately. The script does not write a test file to the destination or verify its write permissions.
PostgreSQL respects existing PGHOST, PGPORT, PGUSER, PGDATABASE, and standard credential configuration; default database is postgres. MySQL/MariaDB uses its usual client option files. Containerized or nondefault database targets may need client configuration before these probes can work.
Dependencies and behavior
Added checks require Python 3 (3.7 or later); HTTPS response checks also need curl. Other checks use the corresponding installed tools: df, systemctl, Docker, database clients, journalctl, timedatectl/chronyc, and Tailscale. Missing tools produce COULD NOT CHECK and the report continues. Nothing is installed automatically.
New external commands have time limits; HTTPS and TLS probes run sequentially. Unreachable sites can make a report take several minutes. The added checks are observational and do not repair services, renew certificates, run backups, restart containers, or change the clock.
The previous script's behavior remains: package metadata refresh attempts, the Certbot renewal-timer enable attempt, and interactive cleanup of old report logs. Therefore the complete script is not strictly read-only, even though the newly added checks are. A completed report exits zero as before; findings are in the summary, not encoded in the process exit status.
Validation
- Bash syntax and embedded Python syntax passed.
- 22 focused tests covered website errors/redirects/timeouts, certificate thresholds, backup freshness/mount/result failures, Linux/BSD disk output, container states, denied database access, timer failures, restricted logs, clock status, Tailscale, reboot requests, and summary completeness.
- Six full-script simulations passed across the original Linux/macOS shell branches and auto/production/both profiles. HTTPS/TLS and machine commands were mocked; no real sites, databases, or services were contacted.
- All original named report sections, original host fields, and historical update-reference text remained present.
- As in the previous merge validation, the sandbox did not permit
/dev/fdprocess substitution. The test copy omitted tee transport and supplied an empty file to the legacy enabled-service loop. The delivered script keeps these operations; they passed syntax checking but were not exercised end-to-end here. - No live-server execution has been performed.
Implementation references: Python TLS verification, Docker container listing, and Tailscale CLI.
#!/usr/bin/env bash
# server-status — one-shot health snapshot + logfile
# Copy to: /usr/local/bin/server-status && sudo chmod +x /usr/local/bin/server-status
# Philosophy: chatty, human-friendly banners + thorough checks, always-on
# Safe on boxes missing tools (skips gracefully)
set -o errexit
set -o pipefail
set -o nounset
VERSION="2.2.0-health-checks"
FQDN="$(hostname -f 2>/dev/null || hostname)"
echo "===== LOGFILE ===="
# ---------- logfile ----------
echo "===== SAVE LOG FILE ====="
STAMP="$(date +"%Y-%m-%d-%H%M%S")"
mkdir -p "${HOME}/server-status"
: > "${HOME}/server-status/.write-check-$$"
rm -f "${HOME}/server-status/.write-check-$$"
LOGFILE="${HOME}/server-status/server-status-${FQDN}-${STAMP}.log"
printf "Log file: %s\n" "$LOGFILE"
# tee everything
exec > >(tee -a "${LOGFILE}") 2>&1
echo " +++++++++++++++++++ if you copy a new version of html builder on this server +++++++++++++++++++ "
echo "sudo chown -R spiffy-root:www-data /var/www/htmlbuilder.shawns-machine.com/"
echo "sudo chmod -R g+w /var/www/htmlbuilder.shawns-machine.com/embeds/"
echo "sudo chmod g+w /var/www/htmlbuilder.shawns-machine.com/"
echo ""
echo "That gives www-data group write access to the embeds directory where PHP needs to create salt/hash/workspace files. "
echo "The rest of the site stays readable by www-data but only writable by spiffy-root (you)."
echo ""
echo "The .htaccess rewrite rules wont work — youre on nginx, not Apache. Nginx doesnt read .htaccess files."
echo "You need to add rewrite rules to your nginx config for this site. On the Mac Mini:"
echo ""
echo "sudo nano /etc/nginx/sites-available/htmlbuilder.shawns-machine.com"
echo "Correct stuff is added into the server block"
# ====== ORIGINAL HOST PROFILES (edit these notes per host) ======
# Usage: bash server-status-enhanced.sh [auto|staging|production|both]
# auto matches a recorded IPv4 address or known host label; otherwise shows both.
STAGING_BANNER_IP="192.168.1.9/24 metric 100 fd5f:7f58:cc54:f270:1a7e:b9ff:fe03:3d95/64 fe80::1a7e:b9ff:fe03:3d95/64"
STAGING_NICKNAME="staging-mini - 2018 mac mini staging"
STAGING_BANNER_DESC="WHAT THIS BOX DOES — hosts staging sites: htmlbuilder.shawns-machine.com recipes.shawns-machine.com tagger.shawns-machine.com vault.shawns-machine.com matomo.shawns-machine.com, flap.shawns-machine.com, wp-flap.shawns-machine.com"
STAGING_EXTRA_DOMAINS=("matomo.shawns-machine.com" "flap.shawns-machine.com" "wp-flap.shawns-machine.com" )
PRODUCTION_BANNER_IP="192.168.1.80/24 metric 100 fd5f:7f58:cc54:f270:6afe:f7ff:fe0f:fbfc/64 fe80::6afe:f7ff:fe0f:fbfc/64"
PRODUCTION_MAC="68:FE:F7:0F:FB:FC"
PRODUCTION_NICKNAME="mimi-prod"
PRODUCTION_BANNER_DESC="PRODUCTION LIVE SITE SERVER - tracker.ferociousbutterfly.com, wp-flap-flap.ferociousbutterfly.com, flap.ferociousbutterfly.com, wp-flap.ferociousbutterfly.com"
PRODUCTION_EXTRA_DOMAINS=("tracker.ferociousbutterfly.com" "wp-flap-flap.ferociousbutterfly.com" "flap.ferociousbutterfly.com" "wp-flap.ferociousbutterfly.com" )
BANNER_NAME="$(hostname -s 2>/dev/null || hostname)"
PROFILE="${1:-auto}"
if [[ "$PROFILE" == auto ]]; then
host_addresses="$(hostname -I 2>/dev/null || true)"
case " $host_addresses " in
*" 192.168.1.9 "*) PROFILE=staging ;;
*" 192.168.1.80 "*) PROFILE=production ;;
*) case "$BANNER_NAME" in
staging-mini|mini-stage) PROFILE=staging ;;
mimi-prod|mini-prod) PROFILE=production ;;
*) PROFILE=both ;;
esac ;;
esac
fi
MAC=""
case "$PROFILE" in
staging)
BANNER_IP="$STAGING_BANNER_IP"; NICKNAME="$STAGING_NICKNAME"
BANNER_DESC="$STAGING_BANNER_DESC"
EXTRA_DOMAINS=("${STAGING_EXTRA_DOMAINS[@]}") ;;
production)
BANNER_IP="$PRODUCTION_BANNER_IP"; NICKNAME="$PRODUCTION_NICKNAME"
BANNER_DESC="$PRODUCTION_BANNER_DESC"; MAC="$PRODUCTION_MAC"
EXTRA_DOMAINS=("${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
both)
BANNER_IP="(see recorded host profiles below)"; NICKNAME="$BANNER_NAME"
BANNER_DESC="Combined staging and production reference; live checks inspect this machine"
EXTRA_DOMAINS=("${STAGING_EXTRA_DOMAINS[@]}" "${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
*) echo "Usage: $0 [auto|staging|production|both]" >&2; exit 2 ;;
esac
# ================================================================
# ====== ADDED HEALTH-CHECK SETTINGS ======
# Edit these values or provide environment variables when running the script.
# Backup status is evidence supplied by YOUR backup job (see health-check-notes.md).
export STATUS_BACKUP_DESTINATION="${STATUS_BACKUP_DESTINATION:-}"
export STATUS_BACKUP_STATUS_FILE="${STATUS_BACKUP_STATUS_FILE:-}"
export STATUS_BACKUP_MOUNT="${STATUS_BACKUP_MOUNT:-}"
export STATUS_BACKUP_MAX_HOURS="${STATUS_BACKUP_MAX_HOURS:-36}"
export STATUS_TAILSCALE_PEER="${STATUS_TAILSCALE_PEER:-}"
export STATUS_DISK_WARN_PERCENT="${STATUS_DISK_WARN_PERCENT:-85}"
export STATUS_DISK_FAIL_PERCENT="${STATUS_DISK_FAIL_PERCENT:-95}"
export STATUS_SLOW_SECONDS="${STATUS_SLOW_SECONDS:-3}"
# Check every site named in the original notes, including those absent from DNS arrays.
STAGING_HEALTH_DOMAINS=("htmlbuilder.shawns-machine.com" "recipes.shawns-machine.com"
"tagger.shawns-machine.com" "vault.shawns-machine.com" "${STAGING_EXTRA_DOMAINS[@]}")
case "$PROFILE" in
staging) HEALTH_DOMAINS=("${STAGING_HEALTH_DOMAINS[@]}") ;;
production) HEALTH_DOMAINS=("${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
both) HEALTH_DOMAINS=("${STAGING_HEALTH_DOMAINS[@]}" "${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
esac
# ================================================================
echo "===== Helpers ===="
# ---------- helpers ----------
have() { command -v "$1" >/dev/null 2>&1 ; }
ts() { date +"%Y-%m-%d %H:%M:%S %Z"; }
section() { printf "\n===== %s =====\n" "$*"; }
kv() { printf "%-26s %s\n" "$1" "$2"; }
# Always show footer/log path even if we exit early
trap '
ec=$?
echo
if (( ec == 0 )); then
section "DONE (exit $ec)"
else
section "⚠️ EARLY EXIT (code $ec)"
echo "Some part of the script terminated prematurely — review the last few lines above for the failure point."
fi
kv "Logfile" "${LOGFILE:-}"
kv "Version" "${VERSION}"
echo "===== ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
' EXIT
echo "===== ${NICKNAME} ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
echo "===== install: sudo cp server-status /usr/local/bin/ && sudo chmod +x /usr/local/bin/server-status ====="
echo "===== NOTES (edit in script header) ====="
section "RECORDED HOST NOTES — labels from the original files"
kv "Selected profile" "$PROFILE"
kv "Staging nickname" "$STAGING_NICKNAME"
kv "Staging addresses" "$STAGING_BANNER_IP"
kv "Staging purpose" "$STAGING_BANNER_DESC"
kv "Production nickname" "$PRODUCTION_NICKNAME"
kv "Production addresses" "$PRODUCTION_BANNER_IP"
kv "Production MAC" "$PRODUCTION_MAC"
kv "Production purpose" "$PRODUCTION_BANNER_DESC"
echo "HTML builder/nginx notes above describe the staging host."
section "HELPERS LOADED"
kv "Now" "$(ts)"
echo "===== OS Identity ===="
# ---------- OS identity ----------
section "OS, HOST, FQDN, KERNEL"
OS="$(uname -s)"
HOST="$(hostname -s 2>/dev/null || hostname)"
FQDN="$(hostname -f 2>/dev/null || hostname)"
KERNEL="$(uname -srmo 2>/dev/null || uname -a)"
kv "Host" "${HOST}"
kv "FQDN" "${FQDN}"
kv "Kernel" "${KERNEL}"
if [[ "${OS}" == "Darwin" ]]; then
kv "OS" "$(sw_vers 2>/dev/null | paste -sd ' ' -)"
else
if have lsb_release; then
kv "OS" "$(lsb_release -d | cut -f2-)"
elif [[ -f /etc/os-release ]]; then
. /etc/os-release; kv "OS" "${PRETTY_NAME:-unknown}"
else
kv "OS" "unknown"
fi
fi
echo "===== PACKAGE MANAGER ===="
# ---------- package manager detect ----------
section "DETECT PACKAGE MANAGER"
echo "Checks for:"
echo " • apt → Debian / Ubuntu / Mint"
echo " • dnf → Fedora / RHEL / AlmaLinux / Rocky (modern)"
echo " • yum → CentOS / RHEL (legacy)"
echo " • pacman → Arch / Manjaro"
echo " • zypper → openSUSE / SLES"
echo " • apk → Alpine Linux"
echo " • brew → macOS (Homebrew)"
PM="unknown"; PKG_LIST_CMD=""; PKG_CHECK_CMD=""; PKG_INSTALL_CMD=""; NEEDS_RESTARTING=""
if have apt-get; then
PM="apt"; PKG_LIST_CMD="apt-mark showmanual"
PKG_CHECK_CMD="dpkg -s"; PKG_INSTALL_CMD="apt install -y"; NEEDS_RESTARTING=""
elif have dnf; then
PM="dnf"; PKG_LIST_CMD="dnf history userinstalled"
PKG_CHECK_CMD="rpm -q"; PKG_INSTALL_CMD="dnf install -y"; NEEDS_RESTARTING="needs-restarting"
elif have yum; then
PM="yum"; PKG_LIST_CMD="yum history userinstalled"
PKG_CHECK_CMD="rpm -q"; PKG_INSTALL_CMD="yum install -y"; NEEDS_RESTARTING="needs-restarting"
elif have pacman; then
PM="pacman"; PKG_LIST_CMD="comm -23 <(pacman -Qqe | sort) <(pacman -Qqg base base-devel | sort)"
PKG_CHECK_CMD="pacman -Qi"; PKG_INSTALL_CMD="pacman -S --noconfirm"
elif have zypper; then
PM="zypper"; PKG_LIST_CMD="zypper se -i -s"
PKG_CHECK_CMD="rpm -q"; PKG_INSTALL_CMD="zypper in -y"
elif have apk; then
PM="apk"; PKG_LIST_CMD="apk info -v"
PKG_CHECK_CMD="apk info -e"; PKG_INSTALL_CMD="apk add"
elif [[ "${OS}" == "Darwin" ]] && have brew; then
PM="brew"; PKG_LIST_CMD="brew list"
PKG_CHECK_CMD="brew ls --versions"; PKG_INSTALL_CMD="brew install"
fi
kv "Package manager" "${PM}"
echo "===== DEPENDENCIES ===="
# ---------- dependencies (soft) ----------
section "MISSING DEPENDENCIES (script helpers only)"
missing=0
need_pkgs=(dmidecode smartmontools lm_sensors bind-utils)
[[ "${PM}" == "apt" ]] && need_pkgs=(dmidecode smartmontools lm-sensors dnsutils)
[[ "${PM}" == "brew" ]] && need_pkgs=(smartmontools coreutils gnu-sed)
# ensure needs-restarting helper on rpm families (often from yum-utils)
[[ "${PM}" == "dnf" || "${PM}" == "yum" ]] && need_pkgs+=("yum-utils")
for pkg in "${need_pkgs[@]}"; do
if ! ${PKG_CHECK_CMD:-echo} "$pkg" >/dev/null 2>&1; then
echo "⚠️ Not installed: ${pkg} (install: sudo ${PKG_INSTALL_CMD:-} ${pkg})"
missing=1
fi
done
[[ $missing -eq 0 ]] && echo "All helper packages present."
if have sensors && have systemctl; then
systemctl is-enabled --quiet lm_sensors 2>/dev/null || echo "ℹ️ sensors present; enable: sudo systemctl enable --now lm_sensors"
fi
if [[ -n "${NEEDS_RESTARTING}" ]] && have "${NEEDS_RESTARTING}"; then
${NEEDS_RESTARTING} -r >/dev/null 2>&1 || true
fi
echo "===== SYSTEM INFO ===="
# ---------- system identity ----------
echo "===== Hostname ===="
section "HOSTNAME INFO"
if have hostnamectl; then hostnamectl || echo "Hostname details unavailable."; else
kv "Timestamp" "$(ts)"; kv "Host" "${HOST}"; kv "FQDN" "${FQDN}"; kv "Kernel" "${KERNEL}"
fi
echo "===== CPU ===="
# ---------- CPU ----------
section "CPU INFO"
if have lscpu; then
lscpu | grep -E 'Model name|Socket|Core|Thread|CPU\(s\)|MHz|Vendor|Flags' || true
else
kv "CPU" "$(sysctl -n machdep.cpu.brand_string 2>/dev/null || true)"
fi
echo "===== MEMORY ===="
# ---------- memory ----------
section "PHYSICAL MEMORY"
if [[ "${OS}" == "Darwin" ]]; then
kv "Memory (GB)" "$(sysctl -n hw.memsize | awk '{printf "%.1f",$1/1024/1024/1024}')"
else
if have dmidecode; then
sudo dmidecode --type 17 | grep -E 'Size:|Speed:|Locator:' | grep -v "No Module Installed" || true
fi
fi
echo "===== DISKS ===="
# ---------- disks ----------
section "PHYSICAL DISKS"
have lsblk && lsblk -o NAME,SIZE,TYPE,MOUNTPOINT,MODEL,FSTYPE,TRAN || echo "Disk inventory unavailable."
echo "===== SMART STATUS ===="
section "SMART STATUS"
smart_one() {
local dev="$1"
if have smartctl; then
echo "Disk: ${dev}"
sudo smartctl -H "${dev}" 2>/dev/null | grep -E 'SMART overall|SMART Health|result' || echo "No SMART data available"
echo
fi
}
for d in /dev/sd? /dev/nvme?n1; do [[ -e "$d" ]] && smart_one "$d"; done
# ---------- temps ----------
section "CPU TEMPERATURE"
if have sensors; then
sensors | grep -E 'Core|temp1' || true
else
echo "sensors not found (install lm_sensors and run: sudo sensors-detect --auto)"
fi
# ---------- network ----------
section "NETWORK INTERFACES"
if have ip; then
ip addr show | grep -E '^[0-9]+:|inet ' || true
else
ifconfig 2>/dev/null | grep -E 'flags|inet ' || true
fi
# ===== USERS & ACCESS AUDIT =====
section "USERS & ACCESS AUDIT"
OS="$(uname -s)"
# Small allow-list for macOS; edit per box
# (Linux path below doesn't use this flagging)
if [[ "$OS" == "Darwin" ]]; then
ALLOW_USERS=("spiffy" "encrypto" "Guest")
fi
# --- Human users (login-able)
section "HUMAN USERS (login shells)"
unknown_users=0
if [[ "$OS" == "Darwin" ]]; then
dscl . list /Users | while read -r u; do
uid=$(dscl . -read "/Users/$u" UniqueID 2>/dev/null | awk '{print $2}' || true)
shell=$(dscl . -read "/Users/$u" UserShell 2>/dev/null | awk '{print $2}' || true)
[[ -z "$uid" || "$uid" -lt 500 ]] && continue
[[ "$shell" == "/usr/bin/false" || "$shell" == "/usr/sbin/nologin" ]] && continue
flag=" **UNKNOWN**"
for allowed_user in "${ALLOW_USERS[@]}"; do
if [[ "$u" == "$allowed_user" ]]; then flag=""; break; fi
done
if [[ -n "$flag" ]]; then unknown_users=1; fi
printf "%-18s uid=%-6s shell=%s%s\n" "$u" "$uid" "$shell" "$flag"
done
else
if have getent; then
getent passwd | awk -F: '($7~/bash|zsh|sh|fish/) && ($3>=1000){printf "%-18s uid=%-6s shell=%s\n",$1,$3,$7}'
else
grep -E '/bin/(bash|zsh|sh|fish)' /etc/passwd | awk -F: '($3>=1000){printf "%-18s uid=%-6s shell=%s\n",$1,$3,$7}'
fi
fi
# --- Who is admin/sudo
section "ADMIN/SUDO GROUP MEMBERS"
if [[ "$OS" == "Darwin" ]]; then
echo "macOS admin group:"
dscl . read /Groups/admin GroupMembership 2>/dev/null || echo "(no info)"
else
if getent group sudo >/dev/null 2>&1; then
echo "sudo group: $(getent group sudo | awk -F: '{print $4}')"
fi
if getent group wheel >/dev/null 2>&1; then
echo "wheel group: $(getent group wheel | awk -F: '{print $4}')"
fi
fi
# --- Sudoers presence (filenames only)
section "SUDOERS FILES (filenames only)"
[[ -r /etc/sudoers ]] && echo "/etc/sudoers present"
if [[ -d /etc/sudoers.d ]]; then
echo "/etc/sudoers.d entries:"
ls -1 /etc/sudoers.d 2>/dev/null || true
else
echo "no /etc/sudoers.d directory"
fi
# --- Current & recent logins
section "CURRENT SESSIONS"
who || true
section "RECENT LOGINS (last 10)"
last -a | head -10 2>/dev/null || echo "no 'last' data"
# --- Cron overview (safe listing)
echo "===== CRON OVERVIEW ===="
section "CRON OVERVIEW"
echo "/etc/cron.* directories:"
ls -1 /etc/cron.* 2>/dev/null || echo "(no cron.* dirs)"
echo "/var/spool/cron (filenames):"
ls -1 /var/spool/cron 2>/dev/null || true
echo "/var/spool/cron/crontabs (filenames):"
ls -1 /var/spool/cron/crontabs 2>/dev/null || true
echo "===== DEFAULT ROUTE ===="
section "DEFAULT ROUTE"
if have ip; then ip route show default || true
elif [[ "${OS}" == "Darwin" ]]; then route -n get default 2>/dev/null | awk '/gateway|interface/{print}' || true; fi
echo "===== DNS RESOLVERS ===="
section "DNS RESOLVERS"
[[ -f /etc/resolv.conf ]] && awk '/^nameserver/{print $2}' /etc/resolv.conf | paste -sd ',' - | sed 's/^/nameservers: /'
echo "===== LISTENING SERVICES ===="
section "LISTENING SERVICES"
if have ss; then
ss -tulpn 2>/dev/null | sed -n '1,120p' || true
elif have lsof; then
lsof -nP -iTCP -sTCP:LISTEN 2>/dev/null | sed -n '1,120p' || true
else
netstat -an 2>/dev/null | grep -i LISTEN | sed -n '1,120p' || true
fi
set +e
echo "--- External listeners (0.0.0.0 / :::) ---"
if command -v ss >/dev/null 2>&1; then
ss -tulpn 2>/dev/null \
| awk '$0 ~ /LISTEN/ && $0 ~ /(0\.0\.0\.0:|:::)/ { print }'
else
netstat -tulpn 2>/dev/null \
| awk '$0 ~ /LISTEN/ && $0 ~ /(0\.0\.0\.0:|:::)/ { print }'
fi
set -e
echo "===== SYSTEM UPDATES ===="
# ---------- updates ----------
section "CHECKING FOR SYSTEM UPDATES"
case "$PM" in
apt) sudo -n apt-get update >/dev/null 2>&1 || true; apt list --upgradable 2>/dev/null | sed -n '1,120p' || true ;;
dnf|yum) ${PM} check-update || true ;;
pacman) pacman -Qu || true ;;
zypper) zypper list-updates || true ;;
apk) apk update >/dev/null 2>&1 || true; apk upgrade -s || true ;;
brew) brew update >/dev/null 2>&1 || true; brew outdated || true ;;
*) echo "Unknown package manager." ;;
esac
echo "===== FIREWALL ===="
# ---------- firewall ----------
section "FIREWALL STATUS"
if command -v ufw &> /dev/null; then
echo "UFW detected"
sudo ufw status verbose || echo "Could not get UFW status"
elif command -v firewall-cmd &> /dev/null; then
echo "firewalld detected"
echo "--- Active zones ---"; sudo firewall-cmd --get-active-zones || true
echo "--- Ports ---"; sudo firewall-cmd --list-ports || true
echo "--- Services ---"; sudo firewall-cmd --list-services || true
echo "--- Rich rules ---"; sudo firewall-cmd --list-rich-rules || true
elif command -v iptables &> /dev/null; then
echo "iptables detected"
sudo iptables -L -n -v | sed -n '1,20p' || true
elif command -v nft &> /dev/null; then
echo "nftables detected"
sudo nft list ruleset | sed -n '1,40p' || true
else
echo "No firewall detected (ufw/firewalld/iptables/nftables)"
fi
echo "===== SYSTEM STATUS ===="
# ---------- system health ----------
section "SYSTEM STATUS"
uptime || true
if have free; then section "MEMORY USAGE"; free -h; fi
if have df; then section "DISK USAGE"; df -hT 2>/dev/null || df -h; fi
section "SERVICES (RUNNING)"
have systemctl && systemctl list-units --type=service --state=running | sed -n '1,200p' || true
section "FAILED SERVICES"
have systemctl && systemctl --failed || true
section "ENABLED AT BOOT"
have systemctl && systemctl list-unit-files --state=enabled --type=service | sed -n '1,200p' || true
section "RECENT CRITICAL LOGS"
have journalctl && journalctl -p 3 -xb | sed -n '1,200p' || true
echo "===== WEB SERVERS ===="
# ---------- web servers ----------
section "WEB SERVER CHECK"
if have systemctl; then
if systemctl is-active --quiet nginx; then
echo "NGINX is active ✅"
(ls -la /etc/nginx/sites-enabled/ 2>/dev/null || echo "No sites-enabled")
(ls -la /etc/nginx/sites-available/ 2>/dev/null || echo "No sites-available")
nginx -v 2>&1 || true
else
echo "NGINX is NOT running ❌"
fi
if systemctl is-active --quiet httpd; then
echo "Apache (httpd) is active ✅"
ls -la /etc/httpd/conf.d/ 2>/dev/null || true
httpd -v 2>/dev/null | head -1 || true
elif systemctl is-active --quiet apache2; then
echo "Apache (apache2) is active ✅"
ls -la /etc/apache2/sites-enabled/ 2>/dev/null || true
apache2 -v 2>/dev/null | head -1 || true
else
echo "Apache is NOT running ❌"
fi
fi
echo "===== AUTOSTART SERVICES (DETAILED) ====="
section "AUTOSTART SERVICES - What runs on boot?"
if have systemctl; then
echo ""
echo "📋 ALL ENABLED SERVICES (auto-start on boot)"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
systemctl list-unit-files --state=enabled --type=service --no-pager --no-legend | \
awk '{printf "%-50s %s\n", $1, $2}' || echo "Service inventory unavailable."
echo ""
echo "⚠️ ENABLED BUT NOT RUNNING (might be an issue)"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
while read -r svc; do
if ! systemctl is-active --quiet "$svc" 2>/dev/null; then
status=$(systemctl is-active "$svc" 2>/dev/null || true)
printf "%-50s [%s]\n" "$svc" "$status"
fi
done < <(systemctl list-unit-files --state=enabled --type=service --no-pager --no-legend | awk '{print $1}')
echo ""
echo "🔧 CUSTOM/USER SERVICES (likely added by you)"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
echo "Services in /etc/systemd/system/ (custom):"
if ls /etc/systemd/system/*.service 2>/dev/null | grep -v '@' >/dev/null; then
for svc_file in /etc/systemd/system/*.service; do
[[ -f "$svc_file" ]] || continue
svc_name=$(basename "$svc_file")
enabled=$(systemctl is-enabled "$svc_name" 2>/dev/null || true)
active=$(systemctl is-active "$svc_name" 2>/dev/null || true)
printf "%-50s [enabled: %-8s active: %-8s]\n" "$svc_name" "$enabled" "$active"
# Show ExecStart line (shows what command runs)
if [[ -f "$svc_file" ]]; then
exec_line=$(grep -E '^ExecStart=' "$svc_file" 2>/dev/null | sed -n '1p' || true)
[[ -n "$exec_line" ]] && echo " └─ $exec_line"
fi
done
else
echo " (no custom services found)"
fi
echo ""
echo "📖 QUICK REFERENCE - Managing services"
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
cat <<'SVCHELP'
sudo systemctl status # check status
sudo systemctl enable # auto-start on boot
sudo systemctl disable # disable auto-start
sudo systemctl restart # restart service
sudo journalctl -u -f # view logs (live)
sudo systemctl daemon-reload # after editing .service files
SVCHELP
else
echo "systemctl not available"
fi
echo "===== PHP ===="
# ---------- php ----------
section "PHP CHECK"
if have php; then php -v | sed -n '1,2p' || true; else echo "PHP not installed."; fi
for svc in php-fpm php8.3-fpm php8.2-fpm php8.1-fpm; do
have systemctl && systemctl is-active --quiet "$svc" 2>/dev/null && echo "Service running: $svc"
done
echo "===== WEB ROOTS ===="
# ---------- web roots ----------
section "WEB CONTENT"
if [[ -d /var/www ]]; then
echo "/var/www:"; ls -la /var/www/ || true; echo
[[ -d /var/www/html ]] && { echo "/var/www/html:"; ls -la /var/www/html || true; }
elif [[ -d /usr/share/nginx/html ]]; then
echo "/usr/share/nginx/html:"; ls -la /usr/share/nginx/html || true
else
echo "No standard web roots found."
fi
echo "===== CERTIFICATES ===="
# ---------- certs ----------
section "SSL CERTIFICATES"
if have certbot; then
certbot --version 2>/dev/null || true
sudo certbot certificates 2>/dev/null || echo "No certs or insufficient permissions."
systemctl enable --now certbot-renew.timer 2>/dev/null || true
else
echo "Certbot not installed."
fi
echo "===== ACCID SUBDOMAINS ===="
# ---------- accid subdomain inventory ----------
section "ACCID SUBDOMAINS (*.shawns-machine.com)"
echo "Per-host nginx configs (excluding wildcard):"
nginx_found=0
for d in /etc/nginx/conf.d /etc/nginx/sites-enabled /etc/nginx/sites-available; do
[[ -d "$d" ]] || continue
for f in "$d"/*shawns-machine.com*; do
[[ -e "$f" ]] || continue
bn="$(basename "$f")"
[[ "$bn" == *wildcard* ]] && continue
echo " $f"
nginx_found=1
done
done
[[ $nginx_found -eq 0 ]] && echo " (none)"
echo
echo "Let's Encrypt certs for *.shawns-machine.com:"
cert_found=0
if [[ -d /etc/letsencrypt/live ]]; then
for dir in /etc/letsencrypt/live/*shawns-machine.com*; do
[[ -d "$dir" ]] || continue
name="$(basename "$dir")"
expiry="$(sudo openssl x509 -enddate -noout -in "$dir/cert.pem" 2>/dev/null | cut -d= -f2- || true)"
printf " %-50s expires %s\n" "$name" "${expiry:-unknown}"
cert_found=1
done
fi
[[ $cert_found -eq 0 ]] && echo " (none)"
echo
echo "Per-host Apache vhosts (filename-matched):"
apache_found=0
for d in /etc/apache2/sites-enabled /etc/apache2/sites-available /etc/httpd/conf.d; do
[[ -d "$d" ]] || continue
for f in "$d"/*shawns-machine.com*; do
[[ -e "$f" ]] || continue
bn="$(basename "$f")"
[[ "$bn" == *wildcard* ]] && continue
echo " $f"
apache_found=1
done
done
[[ $apache_found -eq 0 ]] && echo " (none)"
echo
echo "/var/www/ dirs (full-hostname style):"
www_found=0
for d in /var/www/*shawns-machine.com*; do
[[ -d "$d" ]] || continue
echo " $d"
www_found=1
done
[[ $www_found -eq 0 ]] && echo " (none)"
echo "===== DATABASES ===="
# ---------- databases ----------
section "DATABASE CHECK"
if have systemctl; then
if systemctl is-active --quiet mariadb || systemctl is-active --quiet mysql; then
echo "MySQL/MariaDB active ✅"; mysql --version 2>/dev/null || true
elif systemctl is-active --quiet postgresql; then
echo "PostgreSQL active ✅"; psql --version 2>/dev/null || true
else
echo "No active DB service detected ❌"
fi
fi
echo "===== DOCKER ===="
# ---------- docker ----------
section "DOCKER"
if have docker; then
docker --version
docker ps --format "table {{.ID}}\t{{.Image}}\t{{.Status}}\t{{.Ports}}" || true
else
echo "Docker not installed."
fi
echo "===== USERS / LAST / CRON ===="
# ---------- users / last / cron ----------
section "USER ACCOUNTS"
echo "Users with login shells:"
grep -E '/bin/(bash|zsh|sh|fish)' /etc/passwd | cut -d: -f1 || true
section "LAST LOGINS"
last | head -10 || true
section "CRON JOBS"
echo "System crontabs:"
ls -la /etc/cron.d/ /etc/cron.daily/ /etc/cron.hourly/ /etc/cron.monthly/ /etc/cron.weekly/ 2>/dev/null || true
echo
echo "User crontabs:"
for user in $(cut -f1 -d: /etc/passwd); do
[[ -f "/var/spool/cron/${user}" ]] && echo "crontab for ${user}"
[[ -f "/var/spool/cron/crontabs/${user}" ]] && echo "crontab for ${user}"
done
echo "===== MANUALLY INSTALLED PACKAGES ===="
# ---------- manually installed ----------
section "MANUALLY INSTALLED PACKAGES (first page)"
if [[ -n "${PKG_LIST_CMD}" ]]; then
# shellcheck disable=SC2086
eval "$PKG_LIST_CMD" 2>/dev/null | head -50 || true
else
echo "(no list command for ${PM})"
fi
echo "(The first 50 lines are shown here and saved in the log.)"
echo "===== DNSMASQ ===="
# ---------- dnsmasq ----------
section "DNSMASQ"
if have systemctl && systemctl is-active --quiet dnsmasq; then
echo "dnsmasq is active ✅"
else
echo "dnsmasq is NOT running ❌"
fi
echo
echo "/etc/dnsmasq.d/*.conf:"
if ls /etc/dnsmasq.d/*.conf 1>/dev/null 2>&1; then
grep -H 'address=' /etc/dnsmasq.d/*.conf || echo "No 'address=' entries found."
else
echo "No dnsmasq config files found."
fi
echo "===== DNS RESOLUTION TEST ===="
section "DNS RESOLUTION TEST"
# Only use the domains that are set at the top
EXTRA_DOMAINS_USE=("${EXTRA_DOMAINS[@]}")
resolve_one_domain() {
local d="$1" iplist=""
if have dig; then
iplist="$(dig +short A "$d" 2>/dev/null | grep -E '^[0-9.]+$' | sort -u | paste -sd ',' - || true)"
[[ -z "$iplist" ]] && iplist="$(dig +short AAAA "$d" 2>/dev/null | grep -E '^[0-9a-fA-F:]+$' | sort -u | paste -sd ',' - || true)"
elif have getent; then
iplist="$(getent ahostsv4 "$d" 2>/dev/null | awk '{print $1}' | sort -u | paste -sd ',' - || true)"
[[ -z "$iplist" ]] && iplist="$(getent ahostsv6 "$d" 2>/dev/null | awk '{print $1}' | sort -u | paste -sd ',' - || true)"
elif have host; then
iplist="$(host -t A "$d" 2>/dev/null | awk '/has address/{print $NF}' | sort -u | paste -sd ',' - || true)"
[[ -z "$iplist" ]] && iplist="$(host -t AAAA "$d" 2>/dev/null | awk '/has IPv6 address/{print $NF}' | sort -u | paste -sd ',' - || true)"
else
echo "No dig/getent/host available to test DNS."
return 0
fi
if [[ -n "$iplist" ]]; then
echo "$d -> $iplist"
else
echo "$d -> (no answer)"
fi
}
for domain in "${EXTRA_DOMAINS_USE[@]}"; do
resolve_one_domain "$domain"
done
echo "===== TOP PROCESSES ===="
# ---------- top processes ----------
section "TOP PROCESSES (by RAM)"
if [[ "$OS" == Darwin ]]; then
ps -amr | sed -n '1,11p' || true
else
have ps && ps aux --sort=-%mem | awk 'NR==1||NR<=11{print}' || true
fi
echo ""
echo "===== CLEAR LOGS ====="
if [[ -t 0 ]]; then
ans=""
read -rp "Clear old logs in ${HOME}/server-status/ (keep this run)? [y/N]: " ans || ans=""
if [[ "$ans" =~ ^[Yy]$ ]]; then
for old_log in "${HOME}"/server-status/server-status-*.log; do
[[ -f "$old_log" ]] || continue
[[ "$old_log" == "$LOGFILE" ]] && continue
rm -f -- "$old_log"
done
echo "Logs cleared ✅"
else
echo "Logs kept ✅"
fi
else
echo "(non-interactive run; skipping clear prompt)"
fi
echo ""
# --- AlmaLinux update reference (pure info, not executed) ---
echo ""
echo "===== ALMALINUX SYSTEM UPDATE REFERENCE ====="
echo "Historical notes retained verbatim; version/current-state claims are not live checks."
cat <<'EOF'
📋 Current state & support
--------------------------
- AlmaLinux OS 9 is under active support until **31 May 2027**, and security support until **31 May 2032**.
- AlmaLinux 9.6 (current minor release) includes updated OpenSSL, kernel, and other security packages.
- There are ongoing kernel and library security updates for AlmaLinux 9 (e.g. buffer overflow, use-after-free fixes).
- Stay on a stable minor version (9.5/9.6) and apply regular security updates.
🔍 What to check on your machine before updating
------------------------------------------------
cat /etc/almalinux-release
uname -r
# (Shows AlmaLinux version and kernel version — you want something like “AlmaLinux release 9.x”.)
sudo dnf updateinfo list security # lists pending security updates
sudo dnf check-update # lists all available updates
🛡 What to aim for / avoid
--------------------------
- Stick to official AlmaLinux repos (no beta/testing).
- Always reboot after kernel updates.
- Backup configs before updating (nginx, firewall, etc).
- Schedule updates during a maintenance window.
🧩 Step-by-Step: Safe Update Routine
-----------------------------------
sudo dnf clean all
sudo dnf makecache
sudo dnf check-update
sudo dnf update -y
# After updating:
sudo systemctl daemon-reexec
sudo systemctl daemon-reload
sudo systemctl restart firewalld
sudo systemctl restart nginx
# (or Apache/Redis/etc as appropriate)
sudo dnf autoremove -y
sudo dnf history list
🧭 Notes
--------
- If you saw a “dnf-makecache failed” warning earlier, it usually fixes itself after cleaning cache.
- Kernel updates may appear — always reboot after:
sudo reboot
- 2012 Intel CPUs are fully supported; AlmaLinux 9 packages use x86_64-v2 baseline.
⚙️ Optional: Security-only updates
----------------------------------
sudo dnf update --security
EOF
echo "===== UBUNTU SYSTEM UPDATE REFERENCE ====="
echo "Historical notes retained verbatim; version/current-state claims are not live checks."
cat <<'EOF'
🐧 Ubuntu System Update Reference
=================================
📋 Current state & support
--------------------------
- Ubuntu **24.04 LTS (Noble Numbat)** is under *standard support until April 2029* and *ESM security support until April 2034*.
- Your 6.8 kernel is current and supported; kernel 6.11 is expected in 2026.
- Firmware (Aug 2025) is recent — no urgent action required unless fwupd prompts an upgrade.
🔍 What to check on your machine before updating
------------------------------------------------
lsb_release -a
uname -r
# Show available updates:
sudo apt list --upgradable
# (Optional) Security-focused dry run:
sudo unattended-upgrade --dry-run | grep Inst
🛡 What to aim for / avoid
--------------------------
- Stick with **official Ubuntu repos only** (no “proposed” or unstable PPAs).
- Always **reboot after kernel or libc updates**.
- Backup critical configs (nginx, Apache, etc.) before large upgrades.
- Run updates via terminal for better visibility.
🧩 Step-by-Step: Safe Update Routine
-----------------------------------
sudo apt update
sudo apt upgrade -y
sudo apt full-upgrade -y # optional: ensures kernel & meta packages are current
sudo apt autoremove -y
sudo apt autoclean
# After updates:
sudo systemctl daemon-reexec
sudo systemctl daemon-reload
sudo systemctl restart networking
sudo systemctl restart nginx
# (or apache2, redis, etc. as appropriate)
# If a new kernel is installed:
sudo reboot
🧭 Notes
--------
- Ubuntu backports all critical CVEs to the current LTS kernel; staying on 6.8.x is safe.
- T2-based Mac minis benefit from kernel updates for USB/Fan control.
- If prompted to upgrade firmware:
sudo fwupdmgr refresh
sudo fwupdmgr get-updates
sudo fwupdmgr update
⚙️ Optional: Security-only updates
----------------------------------
sudo apt update
sudo apt install unattended-upgrades
sudo unattended-upgrade
EOF
section "SCRAPER REFERENCE — printed examples, not executed"
cat <<'SCRAPER_HELP'
cd /var/www/htmlbuilder.shawns-machine.com/ultimate_site_parser
source ../venv/bin/activate
# Detailed parsing check:
python - <<'PYTHON'
from core import UltimateSiteParser
import requests
url = 'https://scribbled.space/new-normal-post/'
parser = UltimateSiteParser('https://scribbled.space')
print('Testing parsing after structural element fix...')
print('='*60)
response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10)
modules = parser.parse_page(response.text)
print(f'Framework detected: {parser.framework}')
print(f'Modules extracted: {len(modules)}')
if modules:
print(f'\nFirst 10 modules:')
for idx, module in enumerate(modules[:10], 1):
mod_type = module['type']
layout = module.get('layoutClass', '')
print(f'{idx}. {mod_type:8s} [{layout:15s}]', end='')
if mod_type == 'text':
content = module['content'][:50].replace('\n', ' ')
print(f' - "{content}..."')
elif mod_type == 'image':
src = module['src'].split('/')[-1]
print(f' - {src}')
elif mod_type == 'hero':
print(f' - Hero section')
print(f'\n✅ Successfully extracted {len(modules)} modules!')
else:
print('❌ No modules extracted')
PYTHON
# Quick module-count check (same directory and virtual environment):
python - <<'PYTHON'
from core import UltimateSiteParser
import requests
url = 'https://scribbled.space/new-normal-post/'
parser = UltimateSiteParser('https://scribbled.space')
response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10)
modules = parser.parse_page(response.text)
print(f'Modules: {len(modules)}')
PYTHON
SCRAPER_HELP
# ---------- added health checks (Python 3 standard library, no pip packages) ----------
section "ADDED HEALTH CHECKS"
echo "Checking selected sites plus this machine. Added checks do not perform repairs."
if have python3; then
if python3 - "${HEALTH_DOMAINS[@]}" <<'SERVER_STATUS_HEALTH_PY'
# Embedded standard-library checks. Commands are read-only and have time limits.
import datetime
import json
import math
import os
from pathlib import Path
import re
import shutil
import subprocess
import sys
import time
from urllib.parse import urlsplit
RESULTS = []
ENV = dict(os.environ, LC_ALL='C', SYSTEMD_PAGER='', SYSTEMD_COLORS='0')
def report(level, name, detail):
# Keep command output and host-supplied strings from adding terminal controls.
detail = re.sub(r'[\x00-\x1f\x7f]', ' ', str(detail)).strip()
RESULTS.append((level, name, detail))
print('[{}] {}: {}'.format(level, name, detail), flush=True)
def run(args, timeout=15, env=None):
try:
p = subprocess.run(args, stdin=subprocess.DEVNULL, stdout=subprocess.PIPE,
stderr=subprocess.PIPE, universal_newlines=True,
timeout=timeout, env=env or ENV)
return p.returncode, p.stdout.strip(), p.stderr.strip()
except subprocess.TimeoutExpired:
return 124, '', 'Timed out after {} seconds'.format(timeout)
except OSError as e:
return 127, '', str(e)
def problem(code, out, err):
return (err or out or 'command returned {}'.format(code))[:350]
def available(name):
return shutil.which(name) is not None
def number(name, default, minimum=0, maximum=None):
raw = os.environ.get(name, str(default))
value = float(raw)
if not math.isfinite(value) or value < minimum or (maximum is not None and value > maximum):
raise ValueError('{} is outside the supported range'.format(name))
return value
def websites(domains):
if not available('curl'):
report('COULD NOT CHECK', 'Websites', 'curl is not installed')
return
slow = number('STATUS_SLOW_SECONDS', 3, 0.1)
for domain in domains:
code, out, err = run(['curl', '--disable', '--silent', '--show-error', '--location',
'--max-redirs', '5', '--connect-timeout', '5', '--max-time', '15',
'--proto', '=https', '--proto-redir', '=https',
'--output', os.devnull, '--write-out',
'%{http_code}\t%{time_total}\t%{url_effective}\t%{num_redirects}',
'https://' + domain + '/'], timeout=18)
if code:
report('FAIL', 'HTTPS ' + domain, problem(code, out, err))
continue
try:
status, elapsed, final, redirects = out.rsplit('\t', 3)
status, elapsed, redirects = int(status), float(elapsed), int(redirects)
target = urlsplit(final).hostname
except (ValueError, TypeError):
report('COULD NOT CHECK', 'HTTPS ' + domain, 'Unrecognized curl response')
continue
detail = 'HTTP {}; {:.2f}s; {} redirect(s); final {}'.format(status, elapsed, redirects, final)
level = 'OK'
if status >= 500 or status == 0 or status == 404:
level = 'FAIL'
elif not 200 <= status < 300:
level = 'WARN'
detail += '; may need authentication or a host-specific expected status'
if target != domain:
if level == 'OK':
level = 'WARN'
detail += '; redirected to a different hostname — confirm this is intended'
if elapsed >= slow:
if level == 'OK':
level = 'WARN'
detail += '; slower than {:.1f}s threshold'.format(slow)
report(level, 'HTTPS ' + domain, detail)
print('These are HTTPS GET checks from this machine; they do not test login flows or outside-network reachability.')
# Runs in a child so DNS resolution as well as the TLS handshake is time-bounded.
TLS_PROBE = '''import json, socket, ssl, sys
try:
context = ssl.create_default_context()
with socket.create_connection((sys.argv[1], 443), timeout=5) as sock:
with context.wrap_socket(sock, server_hostname=sys.argv[1]) as tls:
cert = tls.getpeercert()
print(json.dumps({"expires": ssl.cert_time_to_seconds(cert["notAfter"])}))
except ssl.SSLCertVerificationError as exc:
print(json.dumps({"verification_error": str(exc)}))
except Exception as exc:
print(json.dumps({"connection_error": str(exc)}))
'''
def certificates(domains):
for domain in domains:
code, out, err = run([sys.executable, '-c', TLS_PROBE, domain], timeout=12)
if code:
report('COULD NOT CHECK', 'Certificate ' + domain, problem(code, out, err))
continue
data = json.loads(out)
if 'verification_error' in data:
report('FAIL', 'Certificate ' + domain, data['verification_error'])
elif 'connection_error' in data:
report('COULD NOT CHECK', 'Certificate ' + domain, data['connection_error'])
else:
days = (float(data['expires']) - time.time()) / 86400
expiry = datetime.datetime.fromtimestamp(data['expires'], datetime.timezone.utc).isoformat()
if days <= 0:
level, note = 'FAIL', 'expired'
elif days <= 7:
level, note = 'FAIL', 'urgent: 7-day expiry threshold'
elif days <= 14:
level, note = 'WARN', '14-day expiry threshold'
elif days <= 30:
level, note = 'WARN', '30-day expiry threshold'
else:
level, note = 'OK', 'chain and hostname validated'
report(level, 'Certificate ' + domain,
'{:.1f} days remaining; expires {}; {}'.format(days, expiry, note))
print('Certificate checks inspect the certificate served on port 443, which may belong to a proxy/CDN.')
def backups():
destination = os.environ.get('STATUS_BACKUP_DESTINATION', '')
status_file = os.environ.get('STATUS_BACKUP_STATUS_FILE', '')
if not destination:
report('COULD NOT CHECK', 'Backup destination', 'Not configured: set STATUS_BACKUP_DESTINATION')
else:
path = Path(destination).expanduser()
try:
# Do not write test files to a backup destination.
with os.scandir(str(path)) as entries:
next(entries, None)
mount = os.environ.get('STATUS_BACKUP_MOUNT', '')
if mount:
mount_path = Path(mount).expanduser().resolve()
if not os.path.ismount(str(mount_path)):
report('FAIL', 'Backup destination', '{} is not mounted'.format(mount_path))
elif os.path.commonpath([str(path.resolve()), str(mount_path)]) != str(mount_path):
report('FAIL', 'Backup destination', 'Destination is outside the configured backup mount')
else:
report('OK', 'Backup destination', '{} is readable; expected mount is present'.format(path))
else:
report('OK', 'Backup destination', '{} is readable; mount presence not configured'.format(path))
except FileNotFoundError:
report('FAIL', 'Backup destination', '{} is missing or disconnected'.format(path))
except OSError as e:
report('COULD NOT CHECK', 'Backup destination', str(e))
if not status_file:
report('COULD NOT CHECK', 'Backup success/freshness', 'Not configured: set STATUS_BACKUP_STATUS_FILE')
return
try:
data = json.loads(Path(status_file).expanduser().read_text())
# A report from the backup job is stronger evidence than archive-file mtime.
last_success = float(data['last_success_epoch'])
age_hours = (time.time() - last_success) / 3600
max_hours = number('STATUS_BACKUP_MAX_HOURS', 36, 0.1)
if not math.isfinite(last_success) or last_success <= 0:
raise ValueError('last_success_epoch must be a positive Unix timestamp')
state = data.get('last_run_status')
if state not in ('success', 'failed', 'running'):
raise ValueError('last_run_status must be success, failed, or running')
if age_hours < -0.1:
report('WARN', 'Backup freshness', 'Success timestamp is in the future; check clock/status file')
else:
report('WARN' if age_hours > max_hours else 'OK', 'Backup freshness',
'Last reported success {:.1f} hours ago; limit {:.1f} hours'.format(age_hours, max_hours))
report({'success': 'OK', 'failed': 'FAIL', 'running': 'WARN'}[state],
'Backup job result', 'Latest recorded run: ' + state)
print('Backup-job evidence only; archive integrity and a successful restore still need separate verification.')
except (OSError, ValueError, KeyError, TypeError) as e:
report('COULD NOT CHECK', 'Backup success/freshness', 'Cannot read valid backup-job status: ' + str(e))
def storage():
warn = number('STATUS_DISK_WARN_PERCENT', 85, 1, 100)
critical = number('STATUS_DISK_FAIL_PERCENT', 95, warn, 100)
for inode, label in [(False, 'Disk space'), (True, 'Inodes')]:
args = ['df', '-Pi' if inode else '-Pk']
if sys.platform == 'linux':
args += ['-x', 'squashfs', '-x', 'iso9660'] # Immutable images are normally 100% full.
code, out, err = run(args, timeout=15)
if code:
report('COULD NOT CHECK', label, problem(code, out, err))
continue
count = 0
for line in out.splitlines()[1:]:
# Locate the percentage column: BSD df -Pi includes block and inode columns.
columns = line.split()
matches = [i for i, v in enumerate(columns) if re.fullmatch(r'\d+%', v)]
if not matches:
continue
index = matches[-1] if inode else matches[0]
used = int(columns[index][:-1])
mount = ' '.join(columns[index + 1:])
count += 1
level = 'FAIL' if used >= critical else 'WARN' if used >= warn else 'OK'
report(level, label + ' ' + mount, '{}% used (warn {:.0f}%, fail {:.0f}%)'.format(used, warn, critical))
if not count:
report('COULD NOT CHECK', label, 'No usable filesystem percentages returned')
def reboot():
if sys.platform != 'linux':
report('COULD NOT CHECK', 'Reboot required', 'Linux reboot markers do not apply on this OS')
return
marker = Path('/run/reboot-required')
if marker.exists():
report('WARN', 'Reboot required', 'The OS requests a reboot')
try:
print('Packages requesting it: ' + Path('/run/reboot-required.pkgs').read_text().strip())
except OSError:
print('Package list unavailable.')
elif available('apt-get'):
report('OK', 'Reboot required', 'No /run/reboot-required marker; this does not prove all processes use current libraries')
elif available('needs-restarting'):
code, out, err = run(['needs-restarting', '-r'])
report('OK' if code == 0 else 'WARN' if code == 1 else 'COULD NOT CHECK',
'Reboot required', out or err or ('Not requested' if code == 0 else 'Reboot requested'))
else:
report('COULD NOT CHECK', 'Reboot required', 'No supported reboot-status provider')
def properties(unit, fields):
args = ['systemctl', 'show', unit, '--no-pager']
for field in fields:
args += ['--property=' + field]
code, out, err = run(args)
if code:
raise RuntimeError(problem(code, out, err))
return dict(line.split('=', 1) for line in out.splitlines() if '=' in line)
def timers():
if not available('systemctl'):
report('COULD NOT CHECK', 'Scheduled jobs', 'systemctl is not available; cron inventory is above')
return
code, out, err = run(['systemctl', 'list-timers', '--all', '--no-pager'])
if code:
report('COULD NOT CHECK', 'Scheduled jobs', problem(code, out, err))
return
print(out)
code, out, err = run(['systemctl', 'list-units', '--all', '--type=timer', '--no-legend', '--plain', '--no-pager'])
if code:
report('COULD NOT CHECK', 'Timer results', problem(code, out, err))
return
units = [line.split()[0] for line in out.splitlines() if line.strip()]
if not units:
report('COULD NOT CHECK', 'Timer results', 'No loaded timers found')
for unit in units:
try:
data = properties(unit, ['Triggers', 'ActiveState'])
if data.get('ActiveState') != 'active':
report('WARN', unit, 'Timer is {}; may be intentionally disabled'.format(data.get('ActiveState', 'unknown')))
services = [s for s in data.get('Triggers', '').split() if s.endswith('.service')]
if not services:
report('COULD NOT CHECK', unit, 'No triggered service found')
for service in services:
result = properties(service, ['Result', 'ExecMainStatus', 'ExecMainStartTimestamp', 'ActiveState'])
if result.get('Result') and result['Result'] != 'success':
report('FAIL', service, 'Last recorded result: {}; exit {}'.format(result['Result'], result.get('ExecMainStatus', '?')))
elif not result.get('ExecMainStartTimestamp'):
report('COULD NOT CHECK', service, 'No recorded execution; default success is not proof the job ran')
elif result.get('ActiveState') in ('activating', 'deactivating', 'active'):
report('WARN', service, 'Still active/in progress; completion is not established')
elif result.get('Result') == 'success':
report('OK', service, 'Last recorded run succeeded; started ' + result['ExecMainStartTimestamp'])
else:
report('COULD NOT CHECK', service, 'No recorded result')
except RuntimeError as e:
report('COULD NOT CHECK', unit, str(e))
code, out, err = run(['systemctl', '--failed', '--no-legend', '--plain', '--no-pager'])
if code:
report('COULD NOT CHECK', 'Failed units', problem(code, out, err))
elif out:
report('FAIL', 'Failed units', out)
else:
report('OK', 'Failed units', 'No failed systemd units reported')
print('Timer results are the last records retained by systemd, not a complete job history. Cron success needs job-specific evidence.')
def containers():
if not available('docker'):
report('COULD NOT CHECK', 'Docker health', 'Docker is not installed')
return
code, out, err = run(['docker', 'ps', '-aq'])
if code:
report('COULD NOT CHECK', 'Docker health', problem(code, out, err))
return
ids = out.split()
if not ids:
report('OK', 'Docker inventory', 'No containers found; expected application inventory is not configured')
for cid in ids:
code, out, err = run(['docker', 'inspect', '--format', '{{json .}}', cid])
if code:
report('COULD NOT CHECK', 'Container ' + cid, problem(code, out, err))
continue
item = json.loads(out)
name = item.get('Name', cid).lstrip('/')
state = item.get('State', {})
health = state.get('Health', {}).get('Status', 'not configured')
status = state.get('Status', 'unknown')
restarts = item.get('RestartCount', 0)
detail = 'state={}; health={}; restarts={}; exit={}; OOM={}'.format(
status, health, restarts, state.get('ExitCode', '?'), state.get('OOMKilled', False))
if state.get('OOMKilled') or health == 'unhealthy' or status in ('dead', 'restarting'):
level = 'FAIL'
elif status != 'running' or restarts or health == 'starting':
level = 'WARN'
elif health == 'healthy':
level = 'OK'
else:
level = 'COULD NOT CHECK'
detail += '; process is running but application health is not established'
if status == 'exited' and state.get('ExitCode', 0):
level = 'FAIL'
report(level, 'Container ' + name, detail)
print('Stopped containers and historical restarts may be intentional; review before taking action.')
def databases():
checked = False
for label, clients, units in [('MySQL/MariaDB', ['mariadb', 'mysql'], ['mariadb', 'mysql']),
('PostgreSQL', ['psql'], ['postgresql'])]:
client = next((c for c in clients if available(c)), None)
active = available('systemctl') and any(run(['systemctl', 'is-active', '--quiet', u])[0] == 0 for u in units)
if not client and not active:
continue
checked = True
if not client:
report('COULD NOT CHECK', label + ' query', 'Service detected but database client is missing')
continue
env = dict(ENV)
if client == 'psql':
args = [client, '-X', '-w', '-At', '-d', os.environ.get('PGDATABASE', 'postgres'), '-c', 'SELECT 1;']
env['PGCONNECT_TIMEOUT'] = '5'
env['PGOPTIONS'] = env.get('PGOPTIONS', '') + ' -c statement_timeout=5000'
else:
args = [client, '--connect-timeout=5', '--batch', '--skip-column-names', '-e', 'SELECT 1;']
code, out, err = run(args, timeout=10, env=env)
if code == 0 and out.strip() == '1':
report('OK', label + ' query', 'SELECT 1 succeeded using existing client credentials/defaults')
elif re.search(r'access denied|authentication|password|role .* does not exist|database .* does not exist|permission denied', err, re.I):
report('COULD NOT CHECK', label + ' query', 'Existing credentials/database selection do not permit the probe; no password requested')
elif active or code == 124:
report('FAIL', label + ' query', problem(code, out, err))
else:
report('COULD NOT CHECK', label + ' query', 'Default connection failed; service/target not established: ' + problem(code, out, err))
if not checked:
report('COULD NOT CHECK', 'Database query', 'No supported local database client or active service detected')
def hardware_logs():
if not available('journalctl'):
report('COULD NOT CHECK', 'Hardware/resource logs', 'journalctl unavailable on this machine')
return
# A non-root journal view may omit the kernel log without returning an error.
if os.geteuid() != 0 and not {'adm', 'systemd-journal'}.intersection(group_names()):
report('COULD NOT CHECK', 'Hardware/resource logs', 'Kernel journal access not established; run with journal-reading permissions')
return
code, out, err = run(['journalctl', '-k', '--since', '24 hours ago', '--no-pager', '-o', 'short-iso'], timeout=20)
if code or not out or 'No entries' in out or 'permission' in err.lower():
report('COULD NOT CHECK', 'Hardware/resource logs', problem(code, out, err))
return
matches = [line for line in out.splitlines() if re.search(
r'out of memory|oom-kill|killed process|I/O error|buffer I/O|filesystem error|EXT4-fs error|XFS.*(?:error|corrupt)|BTRFS.*(?:error|corrupt)|usb.*(?:reset|disconnect|error)|thermal.*throttl', line, re.I)]
if matches:
report('WARN', 'Hardware/resource logs', '{} matching events in 24 hours; showing newest 30'.format(len(matches)))
for line in matches[-30:]:
print(' ' + re.sub(r'[\x00-\x1f\x7f]', ' ', line))
else:
report('OK', 'Hardware/resource logs', 'No matching events in the accessible kernel journal for the last 24 hours')
def group_names():
import grp
return {grp.getgrgid(g).gr_name for g in os.getgroups()}
def clock_status():
if available('timedatectl'):
code, out, err = run(['timedatectl', 'show', '--property=NTPSynchronized', '--value'])
if code == 0 and out in ('yes', 'no'):
report('OK' if out == 'yes' else 'WARN', 'Time synchronization', 'NTPSynchronized=' + out)
return
if available('chronyc'):
code, out, err = run(['chronyc', 'tracking'])
if code == 0 and 'Leap status' in out:
report('OK' if re.search(r'Leap status\s*:\s*Normal', out) else 'WARN', 'Time synchronization', out)
return
report('COULD NOT CHECK', 'Time synchronization', 'No usable timedatectl/chronyc synchronization result')
def tailscale_status():
if not available('tailscale'):
report('COULD NOT CHECK', 'Tailscale', 'Tailscale is not installed or not on PATH')
return
code, out, err = run(['tailscale', 'status', '--json'])
if code:
report('COULD NOT CHECK', 'Tailscale', problem(code, out, err))
return
data = json.loads(out)
state = data.get('BackendState', 'unknown')
health = data.get('Health') or []
if state == 'Running':
report('WARN' if health else 'OK', 'Tailscale local state', '; '.join(health) if health else 'Running; no reported local health warnings')
else:
report('WARN', 'Tailscale local state', state)
peer = os.environ.get('STATUS_TAILSCALE_PEER', '')
if not peer:
report('COULD NOT CHECK', 'Tailscale peer reachability', 'Set STATUS_TAILSCALE_PEER to an expected peer hostname or IP')
elif peer.startswith('-'):
report('COULD NOT CHECK', 'Tailscale peer reachability', 'Invalid peer value')
else:
code, out, err = run(['tailscale', 'ping', '--c', '1', '--timeout', '5s', peer], timeout=8)
report('OK' if code == 0 else 'WARN', 'Tailscale peer reachability', out or err)
def summary():
print('\n===== NEEDS ATTENTION — ADDED HEALTH CHECKS =====')
for level in ('FAIL', 'WARN', 'COULD NOT CHECK'):
items = [r for r in RESULTS if r[0] == level]
if items:
print('\n{} ({})'.format(level, len(items)))
for _, name, detail in items:
print(' • {}: {}'.format(name, detail))
counts = {level: sum(r[0] == level for r in RESULTS) for level in ('OK', 'WARN', 'FAIL', 'COULD NOT CHECK')}
print('\n' + ' | '.join('{}: {}'.format(k, v) for k, v in counts.items()))
if counts['FAIL'] or counts['WARN']:
print('Please review the items above. No automatic repairs were attempted by these added checks.')
elif counts['COULD NOT CHECK']:
print('Completed checks passed, but coverage is incomplete — unchecked items are not a clean bill of health.')
else:
print('All added checks passed for this snapshot.')
print('The original inventory/report remains above; its unstructured output is not included in these totals.')
def main(domains):
domains = list(dict.fromkeys(domains))
jobs = [('Website responses', lambda: websites(domains)),
('Served TLS certificates', lambda: certificates(domains)),
('Backup evidence', backups), ('Storage thresholds', storage),
('Reboot status', reboot), ('Scheduled jobs and failed units', timers),
('Container health', containers), ('Database responsiveness', databases),
('Recent hardware/resource trouble', hardware_logs),
('Clock synchronization', clock_status), ('Tailscale connectivity', tailscale_status)]
for title, check in jobs:
print('\n===== {} ====='.format(title), flush=True)
try:
check()
except Exception as e:
report('COULD NOT CHECK', title, '{}: {}'.format(type(e).__name__, e))
summary()
if __name__ == '__main__':
main(sys.argv[1:])
SERVER_STATUS_HEALTH_PY
then
:
else
echo "[COULD NOT CHECK] Enhanced checker failed; its report may be incomplete."
fi
else
section "NEEDS ATTENTION — ADDED HEALTH CHECKS"
echo "[COULD NOT CHECK] Python 3 is missing; the original inventory ran, but added checks did not."
fi
# --- Final summary banner ---
echo ""
trap - EXIT # disable the EXIT trap so it won't print a second footer
echo "===== DONE ====="
kv "Logfile" "${LOGFILE}"
kv "Version" "${VERSION}"
echo "===== ${NICKNAME} ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
exit 0