Enhanced Server Status

NotesCode

Enhanced server status

Use server-status-enhanced.sh. It is a standalone script; no companion Python file or pip packages are needed. It preserves the previous merged script's checks and host notes. Both original attachments and server-status-merged.sh remain unchanged.

bash server-status-enhanced.sh
bash server-status-enhanced.sh staging
bash server-status-enhanced.sh production
bash server-status-enhanced.sh both

The default is automatic host-profile selection, falling back to both profiles for an unrecognized machine. Profiles determine website targets and labels; local checks always inspect the machine running the script.

Added checks

  • HTTPS GET responses for every site named in the original host notes, including the four staging sites missing from the original DNS-check array. Reports final status, elapsed time, redirect count, and destination. TLS verification stays enabled; redirects to HTTP are refused. Unexpected cross-host redirects, authentication responses, and responses taking 3 seconds or longer receive warnings.
  • Served certificate validation and expiry: warnings at 30 and 14 days; urgent failures at 7 days, expiry, or invalid hostname/trust. Checks the certificate presented on port 443, possibly a CDN/proxy certificate rather than the origin server's local certificate.
  • Backup destination readability, optional expected mount presence, last successful backup age, and latest recorded job result.
  • Disk-space and inode thresholds: warning at 85%, failure at 95%. Immutable squashfs/ISO images are excluded on Linux because they normally appear full.
  • Ubuntu reboot-required marker and requesting packages; RPM needs-restarting fallback when available.
  • Systemd timer schedule, triggered service's last result, and failed units. A default “success” with no execution timestamp is treated as unverified. Cron listings are preserved, but generic cron success cannot be inferred without job-specific evidence.
  • All Docker containers, including stopped containers, health-check status, restart counts, exit codes, and out-of-memory exits. Running containers without a health check are explicitly unverified.
  • MySQL/MariaDB and PostgreSQL SELECT 1 probes with existing client credentials. No password prompt or credential discovery. Access/authentication problems are distinguished from confirmed probe failures.
  • Last 24 hours of accessible kernel journal entries for out-of-memory kills, disk/filesystem errors, USB resets/disconnects, and thermal throttling.
  • Time synchronization via timedatectl or chronyc.
  • Tailscale local status and health warnings, plus an optional configured peer ping.

The final Needs attention summary groups FAIL, WARN, and COULD NOT CHECK, followed by totals including OK. It summarizes the added health checks; the original inventory output remains above it and is not automatically included in those totals. Missing tools, permissions, credentials, or configuration do not count as a pass.

Settings

Edit the settings block near the top, or set these environment variables:

SettingDefault / purpose
STATUS_BACKUP_DESTINATIONEmpty; local or mounted directory containing backups
STATUS_BACKUP_STATUS_FILEEmpty; JSON status written by the backup job
STATUS_BACKUP_MOUNTEmpty; expected mount point, if backups use a separate mounted destination
STATUS_BACKUP_MAX_HOURS36; maximum age of the last reported successful backup
STATUS_TAILSCALE_PEEREmpty; expected peer hostname or Tailscale IP
STATUS_DISK_WARN_PERCENT85
STATUS_DISK_FAIL_PERCENT95
STATUS_SLOW_SECONDS3

Backup settings should describe the machine on which the script runs. Leave them empty until you know the correct locations. Remote backup APIs and object-storage destinations are not inferred or queried automatically; a local job-result file can still provide evidence for those jobs.

The backup job should maintain a status file with this structure:

{
  "last_success_epoch": 1700000000,
  "last_run_status": "success"
}

The timestamp above is an example, not a current backup result. The job must write its actual last-success Unix timestamp and update last_run_status to success, failed, or running for each attempt. Preserve the prior successful timestamp when a later attempt fails. Write the file atomically at a stable location readable by the status script. This script never fabricates or updates backup evidence.

A readable backup directory and recent success report do not prove archive integrity or restoreability. Continue testing restores separately. The script does not write a test file to the destination or verify its write permissions.

PostgreSQL respects existing PGHOST, PGPORT, PGUSER, PGDATABASE, and standard credential configuration; default database is postgres. MySQL/MariaDB uses its usual client option files. Containerized or nondefault database targets may need client configuration before these probes can work.

Dependencies and behavior

Added checks require Python 3 (3.7 or later); HTTPS response checks also need curl. Other checks use the corresponding installed tools: df, systemctl, Docker, database clients, journalctl, timedatectl/chronyc, and Tailscale. Missing tools produce COULD NOT CHECK and the report continues. Nothing is installed automatically.

New external commands have time limits; HTTPS and TLS probes run sequentially. Unreachable sites can make a report take several minutes. The added checks are observational and do not repair services, renew certificates, run backups, restart containers, or change the clock.

The previous script's behavior remains: package metadata refresh attempts, the Certbot renewal-timer enable attempt, and interactive cleanup of old report logs. Therefore the complete script is not strictly read-only, even though the newly added checks are. A completed report exits zero as before; findings are in the summary, not encoded in the process exit status.

Validation

  • Bash syntax and embedded Python syntax passed.
  • 22 focused tests covered website errors/redirects/timeouts, certificate thresholds, backup freshness/mount/result failures, Linux/BSD disk output, container states, denied database access, timer failures, restricted logs, clock status, Tailscale, reboot requests, and summary completeness.
  • Six full-script simulations passed across the original Linux/macOS shell branches and auto/production/both profiles. HTTPS/TLS and machine commands were mocked; no real sites, databases, or services were contacted.
  • All original named report sections, original host fields, and historical update-reference text remained present.
  • As in the previous merge validation, the sandbox did not permit /dev/fd process substitution. The test copy omitted tee transport and supplied an empty file to the legacy enabled-service loop. The delivered script keeps these operations; they passed syntax checking but were not exercised end-to-end here.
  • No live-server execution has been performed.

Implementation references: Python TLS verification, Docker container listing, and Tailscale CLI.


#!/usr/bin/env bash
# server-status — one-shot health snapshot + logfile
# Copy to: /usr/local/bin/server-status && sudo chmod +x /usr/local/bin/server-status
# Philosophy: chatty, human-friendly banners + thorough checks, always-on
# Safe on boxes missing tools (skips gracefully)

set -o errexit
set -o pipefail
set -o nounset

VERSION="2.2.0-health-checks"

FQDN="$(hostname -f 2>/dev/null || hostname)"
echo "===== LOGFILE  ===="
# ---------- logfile ----------
echo "===== SAVE LOG FILE ====="
STAMP="$(date +"%Y-%m-%d-%H%M%S")"
mkdir -p "${HOME}/server-status"
: > "${HOME}/server-status/.write-check-$$"
rm -f "${HOME}/server-status/.write-check-$$"
LOGFILE="${HOME}/server-status/server-status-${FQDN}-${STAMP}.log"
printf "Log file: %s\n" "$LOGFILE"
# tee everything
exec > >(tee -a "${LOGFILE}") 2>&1

echo " +++++++++++++++++++ if you copy a new version  of html builder on this server +++++++++++++++++++ "
echo "sudo chown -R spiffy-root:www-data /var/www/htmlbuilder.shawns-machine.com/"
echo "sudo chmod -R g+w /var/www/htmlbuilder.shawns-machine.com/embeds/"
echo "sudo chmod g+w /var/www/htmlbuilder.shawns-machine.com/"
echo ""
echo "That gives www-data group write access to the embeds directory where PHP needs to create salt/hash/workspace files. "
echo "The rest of the site stays readable by www-data but only writable by spiffy-root (you)."
echo ""
echo  "The .htaccess rewrite rules wont work — youre on nginx, not Apache. Nginx doesnt read .htaccess files."
echo "You need to add rewrite rules to your nginx config for this site. On the Mac Mini:"
echo ""
echo "sudo nano /etc/nginx/sites-available/htmlbuilder.shawns-machine.com"

echo "Correct stuff is added into the server block"

# ====== ORIGINAL HOST PROFILES (edit these notes per host) ======
# Usage: bash server-status-enhanced.sh [auto|staging|production|both]
# auto matches a recorded IPv4 address or known host label; otherwise shows both.
STAGING_BANNER_IP="192.168.1.9/24 metric 100 fd5f:7f58:cc54:f270:1a7e:b9ff:fe03:3d95/64 fe80::1a7e:b9ff:fe03:3d95/64"
STAGING_NICKNAME="staging-mini - 2018 mac mini staging"
STAGING_BANNER_DESC="WHAT THIS BOX DOES — hosts staging sites:  htmlbuilder.shawns-machine.com  recipes.shawns-machine.com  tagger.shawns-machine.com  vault.shawns-machine.com matomo.shawns-machine.com, flap.shawns-machine.com, wp-flap.shawns-machine.com"
STAGING_EXTRA_DOMAINS=("matomo.shawns-machine.com" "flap.shawns-machine.com" "wp-flap.shawns-machine.com" )
PRODUCTION_BANNER_IP="192.168.1.80/24 metric 100 fd5f:7f58:cc54:f270:6afe:f7ff:fe0f:fbfc/64 fe80::6afe:f7ff:fe0f:fbfc/64"
PRODUCTION_MAC="68:FE:F7:0F:FB:FC"
PRODUCTION_NICKNAME="mimi-prod"
PRODUCTION_BANNER_DESC="PRODUCTION LIVE SITE SERVER - tracker.ferociousbutterfly.com, wp-flap-flap.ferociousbutterfly.com, flap.ferociousbutterfly.com, wp-flap.ferociousbutterfly.com"
PRODUCTION_EXTRA_DOMAINS=("tracker.ferociousbutterfly.com" "wp-flap-flap.ferociousbutterfly.com" "flap.ferociousbutterfly.com" "wp-flap.ferociousbutterfly.com" )

BANNER_NAME="$(hostname -s 2>/dev/null || hostname)"
PROFILE="${1:-auto}"
if [[ "$PROFILE" == auto ]]; then
  host_addresses="$(hostname -I 2>/dev/null || true)"
  case " $host_addresses " in
    *" 192.168.1.9 "*) PROFILE=staging ;;
    *" 192.168.1.80 "*) PROFILE=production ;;
    *) case "$BANNER_NAME" in
         staging-mini|mini-stage) PROFILE=staging ;;
         mimi-prod|mini-prod) PROFILE=production ;;
         *) PROFILE=both ;;
       esac ;;
  esac
fi
MAC=""
case "$PROFILE" in
  staging)
    BANNER_IP="$STAGING_BANNER_IP"; NICKNAME="$STAGING_NICKNAME"
    BANNER_DESC="$STAGING_BANNER_DESC"
    EXTRA_DOMAINS=("${STAGING_EXTRA_DOMAINS[@]}") ;;
  production)
    BANNER_IP="$PRODUCTION_BANNER_IP"; NICKNAME="$PRODUCTION_NICKNAME"
    BANNER_DESC="$PRODUCTION_BANNER_DESC"; MAC="$PRODUCTION_MAC"
    EXTRA_DOMAINS=("${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
  both)
    BANNER_IP="(see recorded host profiles below)"; NICKNAME="$BANNER_NAME"
    BANNER_DESC="Combined staging and production reference; live checks inspect this machine"
    EXTRA_DOMAINS=("${STAGING_EXTRA_DOMAINS[@]}" "${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
  *) echo "Usage: $0 [auto|staging|production|both]" >&2; exit 2 ;;
esac
# ================================================================

# ====== ADDED HEALTH-CHECK SETTINGS ======
# Edit these values or provide environment variables when running the script.
# Backup status is evidence supplied by YOUR backup job (see health-check-notes.md).
export STATUS_BACKUP_DESTINATION="${STATUS_BACKUP_DESTINATION:-}"
export STATUS_BACKUP_STATUS_FILE="${STATUS_BACKUP_STATUS_FILE:-}"
export STATUS_BACKUP_MOUNT="${STATUS_BACKUP_MOUNT:-}"
export STATUS_BACKUP_MAX_HOURS="${STATUS_BACKUP_MAX_HOURS:-36}"
export STATUS_TAILSCALE_PEER="${STATUS_TAILSCALE_PEER:-}"
export STATUS_DISK_WARN_PERCENT="${STATUS_DISK_WARN_PERCENT:-85}"
export STATUS_DISK_FAIL_PERCENT="${STATUS_DISK_FAIL_PERCENT:-95}"
export STATUS_SLOW_SECONDS="${STATUS_SLOW_SECONDS:-3}"
# Check every site named in the original notes, including those absent from DNS arrays.
STAGING_HEALTH_DOMAINS=("htmlbuilder.shawns-machine.com" "recipes.shawns-machine.com"
  "tagger.shawns-machine.com" "vault.shawns-machine.com" "${STAGING_EXTRA_DOMAINS[@]}")
case "$PROFILE" in
  staging) HEALTH_DOMAINS=("${STAGING_HEALTH_DOMAINS[@]}") ;;
  production) HEALTH_DOMAINS=("${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
  both) HEALTH_DOMAINS=("${STAGING_HEALTH_DOMAINS[@]}" "${PRODUCTION_EXTRA_DOMAINS[@]}") ;;
esac
# ================================================================

echo "===== Helpers  ===="
# ---------- helpers ----------
have() { command -v "$1" >/dev/null 2>&1 ; }
ts() { date +"%Y-%m-%d %H:%M:%S %Z"; }
section() { printf "\n===== %s =====\n" "$*"; }
kv() { printf "%-26s %s\n" "$1" "$2"; }

# Always show footer/log path even if we exit early
trap '
  ec=$?
  echo
  if (( ec == 0 )); then
    section "DONE (exit $ec)"
  else
    section "⚠️ EARLY EXIT (code $ec)"
    echo "Some part of the script terminated prematurely — review the last few lines above for the failure point."
  fi
  kv "Logfile" "${LOGFILE:-}"
  kv "Version" "${VERSION}"
  echo "===== ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
' EXIT

echo "===== ${NICKNAME} ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
echo "===== install: sudo cp server-status /usr/local/bin/ && sudo chmod +x /usr/local/bin/server-status ====="
echo "===== NOTES (edit in script header) ====="

section "RECORDED HOST NOTES — labels from the original files"
kv "Selected profile" "$PROFILE"
kv "Staging nickname" "$STAGING_NICKNAME"
kv "Staging addresses" "$STAGING_BANNER_IP"
kv "Staging purpose" "$STAGING_BANNER_DESC"
kv "Production nickname" "$PRODUCTION_NICKNAME"
kv "Production addresses" "$PRODUCTION_BANNER_IP"
kv "Production MAC" "$PRODUCTION_MAC"
kv "Production purpose" "$PRODUCTION_BANNER_DESC"
echo "HTML builder/nginx notes above describe the staging host."

section "HELPERS LOADED"
kv "Now" "$(ts)"

echo "===== OS Identity  ===="
# ---------- OS identity ----------
section "OS, HOST, FQDN, KERNEL"
OS="$(uname -s)"
HOST="$(hostname -s 2>/dev/null || hostname)"
FQDN="$(hostname -f 2>/dev/null || hostname)"
KERNEL="$(uname -srmo 2>/dev/null || uname -a)"
kv "Host"   "${HOST}"
kv "FQDN"   "${FQDN}"
kv "Kernel" "${KERNEL}"
if [[ "${OS}" == "Darwin" ]]; then
  kv "OS" "$(sw_vers 2>/dev/null | paste -sd ' ' -)"
else
  if have lsb_release; then
    kv "OS" "$(lsb_release -d | cut -f2-)"
  elif [[ -f /etc/os-release ]]; then
    . /etc/os-release; kv "OS" "${PRETTY_NAME:-unknown}"
  else
    kv "OS" "unknown"
  fi
fi

echo "===== PACKAGE MANAGER  ===="
# ---------- package manager detect ----------
section "DETECT PACKAGE MANAGER"
echo "Checks for:"
echo "  • apt     → Debian / Ubuntu / Mint"
echo "  • dnf     → Fedora / RHEL / AlmaLinux / Rocky (modern)"
echo "  • yum     → CentOS / RHEL (legacy)"
echo "  • pacman  → Arch / Manjaro"
echo "  • zypper  → openSUSE / SLES"
echo "  • apk     → Alpine Linux"
echo "  • brew    → macOS (Homebrew)"

PM="unknown"; PKG_LIST_CMD=""; PKG_CHECK_CMD=""; PKG_INSTALL_CMD=""; NEEDS_RESTARTING=""
if have apt-get; then
  PM="apt";  PKG_LIST_CMD="apt-mark showmanual"
  PKG_CHECK_CMD="dpkg -s"; PKG_INSTALL_CMD="apt install -y"; NEEDS_RESTARTING=""
elif have dnf; then
  PM="dnf";  PKG_LIST_CMD="dnf history userinstalled"
  PKG_CHECK_CMD="rpm -q";  PKG_INSTALL_CMD="dnf install -y"; NEEDS_RESTARTING="needs-restarting"
elif have yum; then
  PM="yum";  PKG_LIST_CMD="yum history userinstalled"
  PKG_CHECK_CMD="rpm -q";  PKG_INSTALL_CMD="yum install -y";  NEEDS_RESTARTING="needs-restarting"
elif have pacman; then
  PM="pacman"; PKG_LIST_CMD="comm -23 <(pacman -Qqe | sort) <(pacman -Qqg base base-devel | sort)"
  PKG_CHECK_CMD="pacman -Qi"; PKG_INSTALL_CMD="pacman -S --noconfirm"
elif have zypper; then
  PM="zypper"; PKG_LIST_CMD="zypper se -i -s"
  PKG_CHECK_CMD="rpm -q";    PKG_INSTALL_CMD="zypper in -y"
elif have apk; then
  PM="apk"; PKG_LIST_CMD="apk info -v"
  PKG_CHECK_CMD="apk info -e"; PKG_INSTALL_CMD="apk add"
elif [[ "${OS}" == "Darwin" ]] && have brew; then
  PM="brew"; PKG_LIST_CMD="brew list"
  PKG_CHECK_CMD="brew ls --versions"; PKG_INSTALL_CMD="brew install"
fi
kv "Package manager" "${PM}"

echo "===== DEPENDENCIES  ===="
# ---------- dependencies (soft) ----------
section "MISSING DEPENDENCIES (script helpers only)"
missing=0
need_pkgs=(dmidecode smartmontools lm_sensors bind-utils)
[[ "${PM}" == "apt"   ]] && need_pkgs=(dmidecode smartmontools lm-sensors dnsutils)
[[ "${PM}" == "brew"  ]] && need_pkgs=(smartmontools coreutils gnu-sed)
# ensure needs-restarting helper on rpm families (often from yum-utils)
[[ "${PM}" == "dnf" || "${PM}" == "yum" ]] && need_pkgs+=("yum-utils")
for pkg in "${need_pkgs[@]}"; do
  if ! ${PKG_CHECK_CMD:-echo} "$pkg" >/dev/null 2>&1; then
    echo "⚠️  Not installed: ${pkg}   (install: sudo ${PKG_INSTALL_CMD:-} ${pkg})"
    missing=1
  fi
done
[[ $missing -eq 0 ]] && echo "All helper packages present."
if have sensors && have systemctl; then
  systemctl is-enabled --quiet lm_sensors 2>/dev/null || echo "ℹ️  sensors present; enable: sudo systemctl enable --now lm_sensors"
fi
if [[ -n "${NEEDS_RESTARTING}" ]] && have "${NEEDS_RESTARTING}"; then
  ${NEEDS_RESTARTING} -r >/dev/null 2>&1 || true
fi

echo "===== SYSTEM INFO  ===="
# ---------- system identity ----------
echo "===== Hostname  ===="
section "HOSTNAME INFO"
if have hostnamectl; then hostnamectl || echo "Hostname details unavailable."; else
  kv "Timestamp" "$(ts)"; kv "Host" "${HOST}"; kv "FQDN" "${FQDN}"; kv "Kernel" "${KERNEL}"
fi

echo "===== CPU  ===="
# ---------- CPU ----------
section "CPU INFO"
if have lscpu; then
  lscpu | grep -E 'Model name|Socket|Core|Thread|CPU\(s\)|MHz|Vendor|Flags' || true
else
  kv "CPU" "$(sysctl -n machdep.cpu.brand_string 2>/dev/null || true)"
fi

echo "===== MEMORY  ===="
# ---------- memory ----------
section "PHYSICAL MEMORY"
if [[ "${OS}" == "Darwin" ]]; then
  kv "Memory (GB)" "$(sysctl -n hw.memsize | awk '{printf "%.1f",$1/1024/1024/1024}')"
else
  if have dmidecode; then
    sudo dmidecode --type 17 | grep -E 'Size:|Speed:|Locator:' | grep -v "No Module Installed" || true
  fi
fi

echo "===== DISKS  ===="
# ---------- disks ----------
section "PHYSICAL DISKS"
have lsblk && lsblk -o NAME,SIZE,TYPE,MOUNTPOINT,MODEL,FSTYPE,TRAN || echo "Disk inventory unavailable."

echo "===== SMART STATUS ===="
section "SMART STATUS"
smart_one() {
  local dev="$1"
  if have smartctl; then
    echo "Disk: ${dev}"
    sudo smartctl -H "${dev}" 2>/dev/null | grep -E 'SMART overall|SMART Health|result' || echo "No SMART data available"
    echo
  fi
}
for d in /dev/sd? /dev/nvme?n1; do [[ -e "$d" ]] && smart_one "$d"; done

# ---------- temps ----------
section "CPU TEMPERATURE"
if have sensors; then
  sensors | grep -E 'Core|temp1' || true
else
  echo "sensors not found (install lm_sensors and run: sudo sensors-detect --auto)"
fi

# ---------- network ----------
section "NETWORK INTERFACES"
if have ip; then
  ip addr show | grep -E '^[0-9]+:|inet ' || true
else
  ifconfig 2>/dev/null | grep -E 'flags|inet ' || true
fi

# ===== USERS & ACCESS AUDIT =====
section "USERS & ACCESS AUDIT"

OS="$(uname -s)"

# Small allow-list for macOS; edit per box
# (Linux path below doesn't use this flagging)
if [[ "$OS" == "Darwin" ]]; then
  ALLOW_USERS=("spiffy" "encrypto" "Guest")
fi

# --- Human users (login-able)
section "HUMAN USERS (login shells)"
unknown_users=0
if [[ "$OS" == "Darwin" ]]; then
  dscl . list /Users | while read -r u; do
    uid=$(dscl . -read "/Users/$u" UniqueID 2>/dev/null | awk '{print $2}' || true)
    shell=$(dscl . -read "/Users/$u" UserShell 2>/dev/null | awk '{print $2}' || true)
    [[ -z "$uid" || "$uid" -lt 500 ]] && continue
    [[ "$shell" == "/usr/bin/false" || "$shell" == "/usr/sbin/nologin" ]] && continue
    flag="  **UNKNOWN**"
    for allowed_user in "${ALLOW_USERS[@]}"; do
      if [[ "$u" == "$allowed_user" ]]; then flag=""; break; fi
    done
    if [[ -n "$flag" ]]; then unknown_users=1; fi
    printf "%-18s uid=%-6s shell=%s%s\n" "$u" "$uid" "$shell" "$flag"
  done
else
  if have getent; then
    getent passwd | awk -F: '($7~/bash|zsh|sh|fish/) && ($3>=1000){printf "%-18s uid=%-6s shell=%s\n",$1,$3,$7}'
  else
    grep -E '/bin/(bash|zsh|sh|fish)' /etc/passwd | awk -F: '($3>=1000){printf "%-18s uid=%-6s shell=%s\n",$1,$3,$7}'
  fi
fi

# --- Who is admin/sudo
section "ADMIN/SUDO GROUP MEMBERS"
if [[ "$OS" == "Darwin" ]]; then
  echo "macOS admin group:"
  dscl . read /Groups/admin GroupMembership 2>/dev/null || echo "(no info)"
else
  if getent group sudo >/dev/null 2>&1; then
    echo "sudo group: $(getent group sudo | awk -F: '{print $4}')"
  fi
  if getent group wheel >/dev/null 2>&1; then
    echo "wheel group: $(getent group wheel | awk -F: '{print $4}')"
  fi
fi

# --- Sudoers presence (filenames only)
section "SUDOERS FILES (filenames only)"
[[ -r /etc/sudoers ]] && echo "/etc/sudoers present"
if [[ -d /etc/sudoers.d ]]; then
  echo "/etc/sudoers.d entries:"
  ls -1 /etc/sudoers.d 2>/dev/null || true
else
  echo "no /etc/sudoers.d directory"
fi

# --- Current & recent logins
section "CURRENT SESSIONS"
who || true
section "RECENT LOGINS (last 10)"
last -a | head -10 2>/dev/null || echo "no 'last' data"


# --- Cron overview (safe listing)
echo "===== CRON OVERVIEW  ===="
section "CRON OVERVIEW"
echo "/etc/cron.* directories:"
ls -1 /etc/cron.* 2>/dev/null || echo "(no cron.* dirs)"
echo "/var/spool/cron (filenames):"
ls -1 /var/spool/cron 2>/dev/null || true
echo "/var/spool/cron/crontabs (filenames):"
ls -1 /var/spool/cron/crontabs 2>/dev/null || true

echo "===== DEFAULT ROUTE  ===="
section "DEFAULT ROUTE"
if have ip; then ip route show default || true
elif [[ "${OS}" == "Darwin" ]]; then route -n get default 2>/dev/null | awk '/gateway|interface/{print}' || true; fi

echo "===== DNS RESOLVERS  ===="
section "DNS RESOLVERS"
[[ -f /etc/resolv.conf ]] && awk '/^nameserver/{print $2}' /etc/resolv.conf | paste -sd ',' - | sed 's/^/nameservers: /'

echo "===== LISTENING SERVICES  ===="
section "LISTENING SERVICES"
if have ss; then
  ss -tulpn 2>/dev/null | sed -n '1,120p' || true
elif have lsof; then
  lsof -nP -iTCP -sTCP:LISTEN 2>/dev/null | sed -n '1,120p' || true
else
  netstat -an 2>/dev/null | grep -i LISTEN | sed -n '1,120p' || true
fi
set +e
echo "--- External listeners (0.0.0.0 / :::) ---"
if command -v ss >/dev/null 2>&1; then
  ss -tulpn 2>/dev/null \
  | awk '$0 ~ /LISTEN/ && $0 ~ /(0\.0\.0\.0:|:::)/ { print }'
else
  netstat -tulpn 2>/dev/null \
  | awk '$0 ~ /LISTEN/ && $0 ~ /(0\.0\.0\.0:|:::)/ { print }'
fi
set -e

echo "===== SYSTEM UPDATES  ===="
# ---------- updates ----------
section "CHECKING FOR SYSTEM UPDATES"
case "$PM" in
  apt)   sudo -n apt-get update >/dev/null 2>&1 || true; apt list --upgradable 2>/dev/null | sed -n '1,120p' || true ;;
  dnf|yum) ${PM} check-update || true ;;
  pacman) pacman -Qu || true ;;
  zypper) zypper list-updates || true ;;
  apk)   apk update >/dev/null 2>&1 || true; apk upgrade -s || true ;;
  brew)  brew update >/dev/null 2>&1 || true; brew outdated || true ;;
  *)     echo "Unknown package manager." ;;
esac

echo "===== FIREWALL  ===="
# ---------- firewall ----------
section "FIREWALL STATUS"
if command -v ufw &> /dev/null; then
  echo "UFW detected"
  sudo ufw status verbose || echo "Could not get UFW status"
elif command -v firewall-cmd &> /dev/null; then
  echo "firewalld detected"
  echo "--- Active zones ---"; sudo firewall-cmd --get-active-zones || true
  echo "--- Ports ---";        sudo firewall-cmd --list-ports || true
  echo "--- Services ---";     sudo firewall-cmd --list-services || true
  echo "--- Rich rules ---";   sudo firewall-cmd --list-rich-rules || true
elif command -v iptables &> /dev/null; then
  echo "iptables detected"
  sudo iptables -L -n -v | sed -n '1,20p' || true
elif command -v nft &> /dev/null; then
  echo "nftables detected"
  sudo nft list ruleset | sed -n '1,40p' || true
else
  echo "No firewall detected (ufw/firewalld/iptables/nftables)"
fi

echo "===== SYSTEM STATUS  ===="
# ---------- system health ----------
section "SYSTEM STATUS"
uptime || true
if have free; then section "MEMORY USAGE"; free -h; fi
if have df; then section "DISK USAGE"; df -hT 2>/dev/null || df -h; fi

section "SERVICES (RUNNING)"
have systemctl && systemctl list-units --type=service --state=running | sed -n '1,200p' || true
section "FAILED SERVICES"
have systemctl && systemctl --failed || true
section "ENABLED AT BOOT"
have systemctl && systemctl list-unit-files --state=enabled --type=service | sed -n '1,200p' || true

section "RECENT CRITICAL LOGS"
have journalctl && journalctl -p 3 -xb | sed -n '1,200p' || true

echo "===== WEB SERVERS  ===="
# ---------- web servers ----------
section "WEB SERVER CHECK"
if have systemctl; then
  if systemctl is-active --quiet nginx; then
    echo "NGINX is active ✅"
    (ls -la /etc/nginx/sites-enabled/ 2>/dev/null || echo "No sites-enabled")
    (ls -la /etc/nginx/sites-available/ 2>/dev/null || echo "No sites-available")
    nginx -v 2>&1 || true
  else
    echo "NGINX is NOT running ❌"
  fi

  if systemctl is-active --quiet httpd; then
    echo "Apache (httpd) is active ✅"
    ls -la /etc/httpd/conf.d/ 2>/dev/null || true
    httpd -v 2>/dev/null | head -1 || true
  elif systemctl is-active --quiet apache2; then
    echo "Apache (apache2) is active ✅"
    ls -la /etc/apache2/sites-enabled/ 2>/dev/null || true
    apache2 -v 2>/dev/null | head -1 || true
  else
    echo "Apache is NOT running ❌"
  fi
fi

echo "===== AUTOSTART SERVICES (DETAILED) ====="
section "AUTOSTART SERVICES - What runs on boot?"
if have systemctl; then
  echo ""
  echo "📋 ALL ENABLED SERVICES (auto-start on boot)"
  echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
  systemctl list-unit-files --state=enabled --type=service --no-pager --no-legend | \
      awk '{printf "%-50s %s\n", $1, $2}' || echo "Service inventory unavailable."

  echo ""
  echo "⚠️  ENABLED BUT NOT RUNNING (might be an issue)"
  echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
  while read -r svc; do
      if ! systemctl is-active --quiet "$svc" 2>/dev/null; then
        status=$(systemctl is-active "$svc" 2>/dev/null || true)
        printf "%-50s [%s]\n" "$svc" "$status"
      fi
  done < <(systemctl list-unit-files --state=enabled --type=service --no-pager --no-legend | awk '{print $1}')

  echo ""
  echo "🔧 CUSTOM/USER SERVICES (likely added by you)"
  echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
  echo "Services in /etc/systemd/system/ (custom):"
  if ls /etc/systemd/system/*.service 2>/dev/null | grep -v '@' >/dev/null; then
    for svc_file in /etc/systemd/system/*.service; do
      [[ -f "$svc_file" ]] || continue
      svc_name=$(basename "$svc_file")
      enabled=$(systemctl is-enabled "$svc_name" 2>/dev/null || true)
      active=$(systemctl is-active "$svc_name" 2>/dev/null || true)
      printf "%-50s [enabled: %-8s active: %-8s]\n" "$svc_name" "$enabled" "$active"

      # Show ExecStart line (shows what command runs)
      if [[ -f "$svc_file" ]]; then
        exec_line=$(grep -E '^ExecStart=' "$svc_file" 2>/dev/null | sed -n '1p' || true)
        [[ -n "$exec_line" ]] && echo "    └─ $exec_line"
      fi
    done
  else
    echo "  (no custom services found)"
  fi

  echo ""
  echo "📖 QUICK REFERENCE - Managing services"
  echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
  cat <<'SVCHELP'
  sudo systemctl status      # check status
  sudo systemctl enable      # auto-start on boot
  sudo systemctl disable     # disable auto-start
  sudo systemctl restart     # restart service
  sudo journalctl -u  -f     # view logs (live)
  sudo systemctl daemon-reload             # after editing .service files
SVCHELP
else
  echo "systemctl not available"
fi

echo "===== PHP  ===="
# ---------- php ----------
section "PHP CHECK"
if have php; then php -v | sed -n '1,2p' || true; else echo "PHP not installed."; fi
for svc in php-fpm php8.3-fpm php8.2-fpm php8.1-fpm; do
  have systemctl && systemctl is-active --quiet "$svc" 2>/dev/null && echo "Service running: $svc"
done

echo "===== WEB ROOTS  ===="
# ---------- web roots ----------
section "WEB CONTENT"
if [[ -d /var/www ]]; then
  echo "/var/www:"; ls -la /var/www/ || true; echo
  [[ -d /var/www/html ]] && { echo "/var/www/html:"; ls -la /var/www/html || true; }
elif [[ -d /usr/share/nginx/html ]]; then
  echo "/usr/share/nginx/html:"; ls -la /usr/share/nginx/html || true
else
  echo "No standard web roots found."
fi

echo "===== CERTIFICATES  ===="
# ---------- certs ----------
section "SSL CERTIFICATES"
if have certbot; then
  certbot --version 2>/dev/null || true
  sudo certbot certificates 2>/dev/null || echo "No certs or insufficient permissions."
  systemctl enable --now certbot-renew.timer 2>/dev/null || true
else
  echo "Certbot not installed."
fi

  echo "===== ACCID SUBDOMAINS  ===="
  # ---------- accid subdomain inventory ----------
  section "ACCID SUBDOMAINS (*.shawns-machine.com)"

  echo "Per-host nginx configs (excluding wildcard):"
  nginx_found=0
  for d in /etc/nginx/conf.d /etc/nginx/sites-enabled /etc/nginx/sites-available; do
    [[ -d "$d" ]] || continue
    for f in "$d"/*shawns-machine.com*; do
      [[ -e "$f" ]] || continue
      bn="$(basename "$f")"
      [[ "$bn" == *wildcard* ]] && continue
      echo "  $f"
      nginx_found=1
    done
  done
  [[ $nginx_found -eq 0 ]] && echo "  (none)"

  echo
  echo "Let's Encrypt certs for *.shawns-machine.com:"
  cert_found=0
  if [[ -d /etc/letsencrypt/live ]]; then
    for dir in /etc/letsencrypt/live/*shawns-machine.com*; do
      [[ -d "$dir" ]] || continue
      name="$(basename "$dir")"
      expiry="$(sudo openssl x509 -enddate -noout -in "$dir/cert.pem" 2>/dev/null | cut -d= -f2- || true)"
      printf "  %-50s expires %s\n" "$name" "${expiry:-unknown}"
      cert_found=1
    done
  fi
  [[ $cert_found -eq 0 ]] && echo "  (none)"

  echo
  echo "Per-host Apache vhosts (filename-matched):"
  apache_found=0
  for d in /etc/apache2/sites-enabled /etc/apache2/sites-available /etc/httpd/conf.d; do
    [[ -d "$d" ]] || continue
    for f in "$d"/*shawns-machine.com*; do
      [[ -e "$f" ]] || continue
      bn="$(basename "$f")"
      [[ "$bn" == *wildcard* ]] && continue
      echo "  $f"
      apache_found=1
    done
  done
  [[ $apache_found -eq 0 ]] && echo "  (none)"

  echo
  echo "/var/www/ dirs (full-hostname style):"
  www_found=0
  for d in /var/www/*shawns-machine.com*; do
    [[ -d "$d" ]] || continue
    echo "  $d"
    www_found=1
  done
  [[ $www_found -eq 0 ]] && echo "  (none)"

echo "===== DATABASES  ===="
# ---------- databases ----------
section "DATABASE CHECK"
if have systemctl; then
  if systemctl is-active --quiet mariadb || systemctl is-active --quiet mysql; then
    echo "MySQL/MariaDB active ✅"; mysql --version 2>/dev/null || true
  elif systemctl is-active --quiet postgresql; then
    echo "PostgreSQL active ✅"; psql --version 2>/dev/null || true
  else
    echo "No active DB service detected ❌"
  fi
fi

echo "===== DOCKER  ===="
# ---------- docker ----------
section "DOCKER"
if have docker; then
  docker --version
  docker ps --format "table {{.ID}}\t{{.Image}}\t{{.Status}}\t{{.Ports}}" || true
else
  echo "Docker not installed."
fi

echo "===== USERS / LAST / CRON  ===="
# ---------- users / last / cron ----------
section "USER ACCOUNTS"
echo "Users with login shells:"
grep -E '/bin/(bash|zsh|sh|fish)' /etc/passwd | cut -d: -f1 || true

section "LAST LOGINS"
last | head -10 || true

section "CRON JOBS"
echo "System crontabs:"
ls -la /etc/cron.d/ /etc/cron.daily/ /etc/cron.hourly/ /etc/cron.monthly/ /etc/cron.weekly/ 2>/dev/null || true
echo
echo "User crontabs:"
for user in $(cut -f1 -d: /etc/passwd); do
  [[ -f "/var/spool/cron/${user}" ]] && echo "crontab for ${user}"
  [[ -f "/var/spool/cron/crontabs/${user}" ]] && echo "crontab for ${user}"
done

echo "===== MANUALLY INSTALLED PACKAGES  ===="
# ---------- manually installed ----------
section "MANUALLY INSTALLED PACKAGES (first page)"
if [[ -n "${PKG_LIST_CMD}" ]]; then
  # shellcheck disable=SC2086
  eval "$PKG_LIST_CMD" 2>/dev/null | head -50 || true
else
  echo "(no list command for ${PM})"
fi
echo "(The first 50 lines are shown here and saved in the log.)"

echo "===== DNSMASQ  ===="
# ---------- dnsmasq ----------
section "DNSMASQ"
if have systemctl && systemctl is-active --quiet dnsmasq; then
  echo "dnsmasq is active ✅"
else
  echo "dnsmasq is NOT running ❌"
fi
echo
echo "/etc/dnsmasq.d/*.conf:"
if ls /etc/dnsmasq.d/*.conf 1>/dev/null 2>&1; then
  grep -H 'address=' /etc/dnsmasq.d/*.conf || echo "No 'address=' entries found."
else
  echo "No dnsmasq config files found."
fi

echo "===== DNS RESOLUTION TEST  ===="
section "DNS RESOLUTION TEST"

# Only use the domains that are set at the top
EXTRA_DOMAINS_USE=("${EXTRA_DOMAINS[@]}")


resolve_one_domain() {
  local d="$1" iplist=""
  if have dig; then
    iplist="$(dig +short A "$d" 2>/dev/null | grep -E '^[0-9.]+$' | sort -u | paste -sd ',' - || true)"
    [[ -z "$iplist" ]] && iplist="$(dig +short AAAA "$d" 2>/dev/null | grep -E '^[0-9a-fA-F:]+$' | sort -u | paste -sd ',' - || true)"
  elif have getent; then
    iplist="$(getent ahostsv4 "$d" 2>/dev/null | awk '{print $1}' | sort -u | paste -sd ',' - || true)"
    [[ -z "$iplist" ]] && iplist="$(getent ahostsv6 "$d" 2>/dev/null | awk '{print $1}' | sort -u | paste -sd ',' - || true)"
  elif have host; then
    iplist="$(host -t A "$d" 2>/dev/null | awk '/has address/{print $NF}' | sort -u | paste -sd ',' - || true)"
    [[ -z "$iplist" ]] && iplist="$(host -t AAAA "$d" 2>/dev/null | awk '/has IPv6 address/{print $NF}' | sort -u | paste -sd ',' - || true)"
  else
    echo "No dig/getent/host available to test DNS."
    return 0
  fi
  if [[ -n "$iplist" ]]; then
    echo "$d -> $iplist"
  else
    echo "$d -> (no answer)"
  fi
}

for domain in "${EXTRA_DOMAINS_USE[@]}"; do
  resolve_one_domain "$domain"
done


echo "===== TOP PROCESSES  ===="
# ---------- top processes ----------
section "TOP PROCESSES (by RAM)"
if [[ "$OS" == Darwin ]]; then
  ps -amr | sed -n '1,11p' || true
else
  have ps && ps aux --sort=-%mem | awk 'NR==1||NR<=11{print}' || true
fi

echo ""
echo "===== CLEAR LOGS ====="
if [[ -t 0 ]]; then
  ans=""
  read -rp "Clear old logs in ${HOME}/server-status/ (keep this run)? [y/N]: " ans || ans=""
  if [[ "$ans" =~ ^[Yy]$ ]]; then
    for old_log in "${HOME}"/server-status/server-status-*.log; do
      [[ -f "$old_log" ]] || continue
      [[ "$old_log" == "$LOGFILE" ]] && continue
      rm -f -- "$old_log"
    done
    echo "Logs cleared ✅"
  else
    echo "Logs kept ✅"
  fi
else
  echo "(non-interactive run; skipping clear prompt)"
fi
echo ""
# --- AlmaLinux update reference (pure info, not executed) ---
echo ""
echo "===== ALMALINUX SYSTEM UPDATE REFERENCE ====="
echo "Historical notes retained verbatim; version/current-state claims are not live checks."
cat <<'EOF'
📋 Current state & support
--------------------------
- AlmaLinux OS 9 is under active support until **31 May 2027**, and security support until **31 May 2032**.
- AlmaLinux 9.6 (current minor release) includes updated OpenSSL, kernel, and other security packages.
- There are ongoing kernel and library security updates for AlmaLinux 9 (e.g. buffer overflow, use-after-free fixes).
- Stay on a stable minor version (9.5/9.6) and apply regular security updates.

🔍 What to check on your machine before updating
------------------------------------------------
cat /etc/almalinux-release
uname -r

# (Shows AlmaLinux version and kernel version — you want something like “AlmaLinux release 9.x”.)

sudo dnf updateinfo list security     # lists pending security updates
sudo dnf check-update                 # lists all available updates

🛡 What to aim for / avoid
--------------------------
- Stick to official AlmaLinux repos (no beta/testing).
- Always reboot after kernel updates.
- Backup configs before updating (nginx, firewall, etc).
- Schedule updates during a maintenance window.

🧩 Step-by-Step: Safe Update Routine
-----------------------------------
sudo dnf clean all
sudo dnf makecache
sudo dnf check-update
sudo dnf update -y

# After updating:
sudo systemctl daemon-reexec
sudo systemctl daemon-reload
sudo systemctl restart firewalld
sudo systemctl restart nginx
# (or Apache/Redis/etc as appropriate)

sudo dnf autoremove -y
sudo dnf history list

🧭 Notes
--------
- If you saw a “dnf-makecache failed” warning earlier, it usually fixes itself after cleaning cache.
- Kernel updates may appear — always reboot after:
  sudo reboot
- 2012 Intel CPUs are fully supported; AlmaLinux 9 packages use x86_64-v2 baseline.

⚙️ Optional: Security-only updates
----------------------------------
sudo dnf update --security
EOF


echo "===== UBUNTU SYSTEM UPDATE REFERENCE ====="
echo "Historical notes retained verbatim; version/current-state claims are not live checks."
cat <<'EOF'
🐧 Ubuntu System Update Reference
=================================

📋 Current state & support
--------------------------
- Ubuntu **24.04 LTS (Noble Numbat)** is under *standard support until April 2029* and *ESM security support until April 2034*.
- Your 6.8 kernel is current and supported; kernel 6.11 is expected in 2026.
- Firmware (Aug 2025) is recent — no urgent action required unless fwupd prompts an upgrade.

🔍 What to check on your machine before updating
------------------------------------------------
lsb_release -a
uname -r

# Show available updates:
sudo apt list --upgradable

# (Optional) Security-focused dry run:
sudo unattended-upgrade --dry-run | grep Inst

🛡 What to aim for / avoid
--------------------------
- Stick with **official Ubuntu repos only** (no “proposed” or unstable PPAs).
- Always **reboot after kernel or libc updates**.
- Backup critical configs (nginx, Apache, etc.) before large upgrades.
- Run updates via terminal for better visibility.

🧩 Step-by-Step: Safe Update Routine
-----------------------------------
sudo apt update
sudo apt upgrade -y
sudo apt full-upgrade -y        # optional: ensures kernel & meta packages are current
sudo apt autoremove -y
sudo apt autoclean

# After updates:
sudo systemctl daemon-reexec
sudo systemctl daemon-reload
sudo systemctl restart networking
sudo systemctl restart nginx
# (or apache2, redis, etc. as appropriate)

# If a new kernel is installed:
sudo reboot

🧭 Notes
--------
- Ubuntu backports all critical CVEs to the current LTS kernel; staying on 6.8.x is safe.
- T2-based Mac minis benefit from kernel updates for USB/Fan control.
- If prompted to upgrade firmware:
  sudo fwupdmgr refresh
  sudo fwupdmgr get-updates
  sudo fwupdmgr update

⚙️ Optional: Security-only updates
----------------------------------
sudo apt update
sudo apt install unattended-upgrades
sudo unattended-upgrade
EOF

section "SCRAPER REFERENCE — printed examples, not executed"
cat <<'SCRAPER_HELP'
cd /var/www/htmlbuilder.shawns-machine.com/ultimate_site_parser
source ../venv/bin/activate

# Detailed parsing check:
python - <<'PYTHON'
from core import UltimateSiteParser
import requests
url = 'https://scribbled.space/new-normal-post/'
parser = UltimateSiteParser('https://scribbled.space')
print('Testing parsing after structural element fix...')
print('='*60)
response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10)
modules = parser.parse_page(response.text)
print(f'Framework detected: {parser.framework}')
print(f'Modules extracted: {len(modules)}')
if modules:
    print(f'\nFirst 10 modules:')
    for idx, module in enumerate(modules[:10], 1):
        mod_type = module['type']
        layout = module.get('layoutClass', '')
        print(f'{idx}. {mod_type:8s} [{layout:15s}]', end='')
        if mod_type == 'text':
            content = module['content'][:50].replace('\n', ' ')
            print(f' - "{content}..."')
        elif mod_type == 'image':
            src = module['src'].split('/')[-1]
            print(f' - {src}')
        elif mod_type == 'hero':
            print(f' - Hero section')
    print(f'\n✅ Successfully extracted {len(modules)} modules!')
else:
    print('❌ No modules extracted')
PYTHON

# Quick module-count check (same directory and virtual environment):
python - <<'PYTHON'
from core import UltimateSiteParser
import requests
url = 'https://scribbled.space/new-normal-post/'
parser = UltimateSiteParser('https://scribbled.space')
response = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, timeout=10)
modules = parser.parse_page(response.text)
print(f'Modules: {len(modules)}')
PYTHON
SCRAPER_HELP

# ---------- added health checks (Python 3 standard library, no pip packages) ----------
section "ADDED HEALTH CHECKS"
echo "Checking selected sites plus this machine. Added checks do not perform repairs."
if have python3; then
  if python3 - "${HEALTH_DOMAINS[@]}" <<'SERVER_STATUS_HEALTH_PY'
# Embedded standard-library checks. Commands are read-only and have time limits.
import datetime
import json
import math
import os
from pathlib import Path
import re
import shutil
import subprocess
import sys
import time
from urllib.parse import urlsplit

RESULTS = []
ENV = dict(os.environ, LC_ALL='C', SYSTEMD_PAGER='', SYSTEMD_COLORS='0')


def report(level, name, detail):
    # Keep command output and host-supplied strings from adding terminal controls.
    detail = re.sub(r'[\x00-\x1f\x7f]', ' ', str(detail)).strip()
    RESULTS.append((level, name, detail))
    print('[{}] {}: {}'.format(level, name, detail), flush=True)


def run(args, timeout=15, env=None):
    try:
        p = subprocess.run(args, stdin=subprocess.DEVNULL, stdout=subprocess.PIPE,
                           stderr=subprocess.PIPE, universal_newlines=True,
                           timeout=timeout, env=env or ENV)
        return p.returncode, p.stdout.strip(), p.stderr.strip()
    except subprocess.TimeoutExpired:
        return 124, '', 'Timed out after {} seconds'.format(timeout)
    except OSError as e:
        return 127, '', str(e)


def problem(code, out, err):
    return (err or out or 'command returned {}'.format(code))[:350]


def available(name):
    return shutil.which(name) is not None


def number(name, default, minimum=0, maximum=None):
    raw = os.environ.get(name, str(default))
    value = float(raw)
    if not math.isfinite(value) or value < minimum or (maximum is not None and value > maximum):
        raise ValueError('{} is outside the supported range'.format(name))
    return value


def websites(domains):
    if not available('curl'):
        report('COULD NOT CHECK', 'Websites', 'curl is not installed')
        return
    slow = number('STATUS_SLOW_SECONDS', 3, 0.1)
    for domain in domains:
        code, out, err = run(['curl', '--disable', '--silent', '--show-error', '--location',
                             '--max-redirs', '5', '--connect-timeout', '5', '--max-time', '15',
                             '--proto', '=https', '--proto-redir', '=https',
                             '--output', os.devnull, '--write-out',
                             '%{http_code}\t%{time_total}\t%{url_effective}\t%{num_redirects}',
                             'https://' + domain + '/'], timeout=18)
        if code:
            report('FAIL', 'HTTPS ' + domain, problem(code, out, err))
            continue
        try:
            status, elapsed, final, redirects = out.rsplit('\t', 3)
            status, elapsed, redirects = int(status), float(elapsed), int(redirects)
            target = urlsplit(final).hostname
        except (ValueError, TypeError):
            report('COULD NOT CHECK', 'HTTPS ' + domain, 'Unrecognized curl response')
            continue
        detail = 'HTTP {}; {:.2f}s; {} redirect(s); final {}'.format(status, elapsed, redirects, final)
        level = 'OK'
        if status >= 500 or status == 0 or status == 404:
            level = 'FAIL'
        elif not 200 <= status < 300:
            level = 'WARN'
            detail += '; may need authentication or a host-specific expected status'
        if target != domain:
            if level == 'OK':
                level = 'WARN'
            detail += '; redirected to a different hostname — confirm this is intended'
        if elapsed >= slow:
            if level == 'OK':
                level = 'WARN'
            detail += '; slower than {:.1f}s threshold'.format(slow)
        report(level, 'HTTPS ' + domain, detail)
    print('These are HTTPS GET checks from this machine; they do not test login flows or outside-network reachability.')


# Runs in a child so DNS resolution as well as the TLS handshake is time-bounded.
TLS_PROBE = '''import json, socket, ssl, sys
try:
    context = ssl.create_default_context()
    with socket.create_connection((sys.argv[1], 443), timeout=5) as sock:
        with context.wrap_socket(sock, server_hostname=sys.argv[1]) as tls:
            cert = tls.getpeercert()
            print(json.dumps({"expires": ssl.cert_time_to_seconds(cert["notAfter"])}))
except ssl.SSLCertVerificationError as exc:
    print(json.dumps({"verification_error": str(exc)}))
except Exception as exc:
    print(json.dumps({"connection_error": str(exc)}))
'''


def certificates(domains):
    for domain in domains:
        code, out, err = run([sys.executable, '-c', TLS_PROBE, domain], timeout=12)
        if code:
            report('COULD NOT CHECK', 'Certificate ' + domain, problem(code, out, err))
            continue
        data = json.loads(out)
        if 'verification_error' in data:
            report('FAIL', 'Certificate ' + domain, data['verification_error'])
        elif 'connection_error' in data:
            report('COULD NOT CHECK', 'Certificate ' + domain, data['connection_error'])
        else:
            days = (float(data['expires']) - time.time()) / 86400
            expiry = datetime.datetime.fromtimestamp(data['expires'], datetime.timezone.utc).isoformat()
            if days <= 0:
                level, note = 'FAIL', 'expired'
            elif days <= 7:
                level, note = 'FAIL', 'urgent: 7-day expiry threshold'
            elif days <= 14:
                level, note = 'WARN', '14-day expiry threshold'
            elif days <= 30:
                level, note = 'WARN', '30-day expiry threshold'
            else:
                level, note = 'OK', 'chain and hostname validated'
            report(level, 'Certificate ' + domain,
                   '{:.1f} days remaining; expires {}; {}'.format(days, expiry, note))
    print('Certificate checks inspect the certificate served on port 443, which may belong to a proxy/CDN.')


def backups():
    destination = os.environ.get('STATUS_BACKUP_DESTINATION', '')
    status_file = os.environ.get('STATUS_BACKUP_STATUS_FILE', '')
    if not destination:
        report('COULD NOT CHECK', 'Backup destination', 'Not configured: set STATUS_BACKUP_DESTINATION')
    else:
        path = Path(destination).expanduser()
        try:
            # Do not write test files to a backup destination.
            with os.scandir(str(path)) as entries:
                next(entries, None)
            mount = os.environ.get('STATUS_BACKUP_MOUNT', '')
            if mount:
                mount_path = Path(mount).expanduser().resolve()
                if not os.path.ismount(str(mount_path)):
                    report('FAIL', 'Backup destination', '{} is not mounted'.format(mount_path))
                elif os.path.commonpath([str(path.resolve()), str(mount_path)]) != str(mount_path):
                    report('FAIL', 'Backup destination', 'Destination is outside the configured backup mount')
                else:
                    report('OK', 'Backup destination', '{} is readable; expected mount is present'.format(path))
            else:
                report('OK', 'Backup destination', '{} is readable; mount presence not configured'.format(path))
        except FileNotFoundError:
            report('FAIL', 'Backup destination', '{} is missing or disconnected'.format(path))
        except OSError as e:
            report('COULD NOT CHECK', 'Backup destination', str(e))
    if not status_file:
        report('COULD NOT CHECK', 'Backup success/freshness', 'Not configured: set STATUS_BACKUP_STATUS_FILE')
        return
    try:
        data = json.loads(Path(status_file).expanduser().read_text())
        # A report from the backup job is stronger evidence than archive-file mtime.
        last_success = float(data['last_success_epoch'])
        age_hours = (time.time() - last_success) / 3600
        max_hours = number('STATUS_BACKUP_MAX_HOURS', 36, 0.1)
        if not math.isfinite(last_success) or last_success <= 0:
            raise ValueError('last_success_epoch must be a positive Unix timestamp')
        state = data.get('last_run_status')
        if state not in ('success', 'failed', 'running'):
            raise ValueError('last_run_status must be success, failed, or running')
        if age_hours < -0.1:
            report('WARN', 'Backup freshness', 'Success timestamp is in the future; check clock/status file')
        else:
            report('WARN' if age_hours > max_hours else 'OK', 'Backup freshness',
                   'Last reported success {:.1f} hours ago; limit {:.1f} hours'.format(age_hours, max_hours))
        report({'success': 'OK', 'failed': 'FAIL', 'running': 'WARN'}[state],
               'Backup job result', 'Latest recorded run: ' + state)
        print('Backup-job evidence only; archive integrity and a successful restore still need separate verification.')
    except (OSError, ValueError, KeyError, TypeError) as e:
        report('COULD NOT CHECK', 'Backup success/freshness', 'Cannot read valid backup-job status: ' + str(e))


def storage():
    warn = number('STATUS_DISK_WARN_PERCENT', 85, 1, 100)
    critical = number('STATUS_DISK_FAIL_PERCENT', 95, warn, 100)
    for inode, label in [(False, 'Disk space'), (True, 'Inodes')]:
        args = ['df', '-Pi' if inode else '-Pk']
        if sys.platform == 'linux':
            args += ['-x', 'squashfs', '-x', 'iso9660']  # Immutable images are normally 100% full.
        code, out, err = run(args, timeout=15)
        if code:
            report('COULD NOT CHECK', label, problem(code, out, err))
            continue
        count = 0
        for line in out.splitlines()[1:]:
            # Locate the percentage column: BSD df -Pi includes block and inode columns.
            columns = line.split()
            matches = [i for i, v in enumerate(columns) if re.fullmatch(r'\d+%', v)]
            if not matches:
                continue
            index = matches[-1] if inode else matches[0]
            used = int(columns[index][:-1])
            mount = ' '.join(columns[index + 1:])
            count += 1
            level = 'FAIL' if used >= critical else 'WARN' if used >= warn else 'OK'
            report(level, label + ' ' + mount, '{}% used (warn {:.0f}%, fail {:.0f}%)'.format(used, warn, critical))
        if not count:
            report('COULD NOT CHECK', label, 'No usable filesystem percentages returned')


def reboot():
    if sys.platform != 'linux':
        report('COULD NOT CHECK', 'Reboot required', 'Linux reboot markers do not apply on this OS')
        return
    marker = Path('/run/reboot-required')
    if marker.exists():
        report('WARN', 'Reboot required', 'The OS requests a reboot')
        try:
            print('Packages requesting it: ' + Path('/run/reboot-required.pkgs').read_text().strip())
        except OSError:
            print('Package list unavailable.')
    elif available('apt-get'):
        report('OK', 'Reboot required', 'No /run/reboot-required marker; this does not prove all processes use current libraries')
    elif available('needs-restarting'):
        code, out, err = run(['needs-restarting', '-r'])
        report('OK' if code == 0 else 'WARN' if code == 1 else 'COULD NOT CHECK',
               'Reboot required', out or err or ('Not requested' if code == 0 else 'Reboot requested'))
    else:
        report('COULD NOT CHECK', 'Reboot required', 'No supported reboot-status provider')


def properties(unit, fields):
    args = ['systemctl', 'show', unit, '--no-pager']
    for field in fields:
        args += ['--property=' + field]
    code, out, err = run(args)
    if code:
        raise RuntimeError(problem(code, out, err))
    return dict(line.split('=', 1) for line in out.splitlines() if '=' in line)


def timers():
    if not available('systemctl'):
        report('COULD NOT CHECK', 'Scheduled jobs', 'systemctl is not available; cron inventory is above')
        return
    code, out, err = run(['systemctl', 'list-timers', '--all', '--no-pager'])
    if code:
        report('COULD NOT CHECK', 'Scheduled jobs', problem(code, out, err))
        return
    print(out)
    code, out, err = run(['systemctl', 'list-units', '--all', '--type=timer', '--no-legend', '--plain', '--no-pager'])
    if code:
        report('COULD NOT CHECK', 'Timer results', problem(code, out, err))
        return
    units = [line.split()[0] for line in out.splitlines() if line.strip()]
    if not units:
        report('COULD NOT CHECK', 'Timer results', 'No loaded timers found')
    for unit in units:
        try:
            data = properties(unit, ['Triggers', 'ActiveState'])
            if data.get('ActiveState') != 'active':
                report('WARN', unit, 'Timer is {}; may be intentionally disabled'.format(data.get('ActiveState', 'unknown')))
            services = [s for s in data.get('Triggers', '').split() if s.endswith('.service')]
            if not services:
                report('COULD NOT CHECK', unit, 'No triggered service found')
            for service in services:
                result = properties(service, ['Result', 'ExecMainStatus', 'ExecMainStartTimestamp', 'ActiveState'])
                if result.get('Result') and result['Result'] != 'success':
                    report('FAIL', service, 'Last recorded result: {}; exit {}'.format(result['Result'], result.get('ExecMainStatus', '?')))
                elif not result.get('ExecMainStartTimestamp'):
                    report('COULD NOT CHECK', service, 'No recorded execution; default success is not proof the job ran')
                elif result.get('ActiveState') in ('activating', 'deactivating', 'active'):
                    report('WARN', service, 'Still active/in progress; completion is not established')
                elif result.get('Result') == 'success':
                    report('OK', service, 'Last recorded run succeeded; started ' + result['ExecMainStartTimestamp'])
                else:
                    report('COULD NOT CHECK', service, 'No recorded result')
        except RuntimeError as e:
            report('COULD NOT CHECK', unit, str(e))
    code, out, err = run(['systemctl', '--failed', '--no-legend', '--plain', '--no-pager'])
    if code:
        report('COULD NOT CHECK', 'Failed units', problem(code, out, err))
    elif out:
        report('FAIL', 'Failed units', out)
    else:
        report('OK', 'Failed units', 'No failed systemd units reported')
    print('Timer results are the last records retained by systemd, not a complete job history. Cron success needs job-specific evidence.')


def containers():
    if not available('docker'):
        report('COULD NOT CHECK', 'Docker health', 'Docker is not installed')
        return
    code, out, err = run(['docker', 'ps', '-aq'])
    if code:
        report('COULD NOT CHECK', 'Docker health', problem(code, out, err))
        return
    ids = out.split()
    if not ids:
        report('OK', 'Docker inventory', 'No containers found; expected application inventory is not configured')
    for cid in ids:
        code, out, err = run(['docker', 'inspect', '--format', '{{json .}}', cid])
        if code:
            report('COULD NOT CHECK', 'Container ' + cid, problem(code, out, err))
            continue
        item = json.loads(out)
        name = item.get('Name', cid).lstrip('/')
        state = item.get('State', {})
        health = state.get('Health', {}).get('Status', 'not configured')
        status = state.get('Status', 'unknown')
        restarts = item.get('RestartCount', 0)
        detail = 'state={}; health={}; restarts={}; exit={}; OOM={}'.format(
            status, health, restarts, state.get('ExitCode', '?'), state.get('OOMKilled', False))
        if state.get('OOMKilled') or health == 'unhealthy' or status in ('dead', 'restarting'):
            level = 'FAIL'
        elif status != 'running' or restarts or health == 'starting':
            level = 'WARN'
        elif health == 'healthy':
            level = 'OK'
        else:
            level = 'COULD NOT CHECK'
            detail += '; process is running but application health is not established'
        if status == 'exited' and state.get('ExitCode', 0):
            level = 'FAIL'
        report(level, 'Container ' + name, detail)
    print('Stopped containers and historical restarts may be intentional; review before taking action.')


def databases():
    checked = False
    for label, clients, units in [('MySQL/MariaDB', ['mariadb', 'mysql'], ['mariadb', 'mysql']),
                                  ('PostgreSQL', ['psql'], ['postgresql'])]:
        client = next((c for c in clients if available(c)), None)
        active = available('systemctl') and any(run(['systemctl', 'is-active', '--quiet', u])[0] == 0 for u in units)
        if not client and not active:
            continue
        checked = True
        if not client:
            report('COULD NOT CHECK', label + ' query', 'Service detected but database client is missing')
            continue
        env = dict(ENV)
        if client == 'psql':
            args = [client, '-X', '-w', '-At', '-d', os.environ.get('PGDATABASE', 'postgres'), '-c', 'SELECT 1;']
            env['PGCONNECT_TIMEOUT'] = '5'
            env['PGOPTIONS'] = env.get('PGOPTIONS', '') + ' -c statement_timeout=5000'
        else:
            args = [client, '--connect-timeout=5', '--batch', '--skip-column-names', '-e', 'SELECT 1;']
        code, out, err = run(args, timeout=10, env=env)
        if code == 0 and out.strip() == '1':
            report('OK', label + ' query', 'SELECT 1 succeeded using existing client credentials/defaults')
        elif re.search(r'access denied|authentication|password|role .* does not exist|database .* does not exist|permission denied', err, re.I):
            report('COULD NOT CHECK', label + ' query', 'Existing credentials/database selection do not permit the probe; no password requested')
        elif active or code == 124:
            report('FAIL', label + ' query', problem(code, out, err))
        else:
            report('COULD NOT CHECK', label + ' query', 'Default connection failed; service/target not established: ' + problem(code, out, err))
    if not checked:
        report('COULD NOT CHECK', 'Database query', 'No supported local database client or active service detected')


def hardware_logs():
    if not available('journalctl'):
        report('COULD NOT CHECK', 'Hardware/resource logs', 'journalctl unavailable on this machine')
        return
    # A non-root journal view may omit the kernel log without returning an error.
    if os.geteuid() != 0 and not {'adm', 'systemd-journal'}.intersection(group_names()):
        report('COULD NOT CHECK', 'Hardware/resource logs', 'Kernel journal access not established; run with journal-reading permissions')
        return
    code, out, err = run(['journalctl', '-k', '--since', '24 hours ago', '--no-pager', '-o', 'short-iso'], timeout=20)
    if code or not out or 'No entries' in out or 'permission' in err.lower():
        report('COULD NOT CHECK', 'Hardware/resource logs', problem(code, out, err))
        return
    matches = [line for line in out.splitlines() if re.search(
        r'out of memory|oom-kill|killed process|I/O error|buffer I/O|filesystem error|EXT4-fs error|XFS.*(?:error|corrupt)|BTRFS.*(?:error|corrupt)|usb.*(?:reset|disconnect|error)|thermal.*throttl', line, re.I)]
    if matches:
        report('WARN', 'Hardware/resource logs', '{} matching events in 24 hours; showing newest 30'.format(len(matches)))
        for line in matches[-30:]:
            print('  ' + re.sub(r'[\x00-\x1f\x7f]', ' ', line))
    else:
        report('OK', 'Hardware/resource logs', 'No matching events in the accessible kernel journal for the last 24 hours')


def group_names():
    import grp
    return {grp.getgrgid(g).gr_name for g in os.getgroups()}


def clock_status():
    if available('timedatectl'):
        code, out, err = run(['timedatectl', 'show', '--property=NTPSynchronized', '--value'])
        if code == 0 and out in ('yes', 'no'):
            report('OK' if out == 'yes' else 'WARN', 'Time synchronization', 'NTPSynchronized=' + out)
            return
    if available('chronyc'):
        code, out, err = run(['chronyc', 'tracking'])
        if code == 0 and 'Leap status' in out:
            report('OK' if re.search(r'Leap status\s*:\s*Normal', out) else 'WARN', 'Time synchronization', out)
            return
    report('COULD NOT CHECK', 'Time synchronization', 'No usable timedatectl/chronyc synchronization result')


def tailscale_status():
    if not available('tailscale'):
        report('COULD NOT CHECK', 'Tailscale', 'Tailscale is not installed or not on PATH')
        return
    code, out, err = run(['tailscale', 'status', '--json'])
    if code:
        report('COULD NOT CHECK', 'Tailscale', problem(code, out, err))
        return
    data = json.loads(out)
    state = data.get('BackendState', 'unknown')
    health = data.get('Health') or []
    if state == 'Running':
        report('WARN' if health else 'OK', 'Tailscale local state', '; '.join(health) if health else 'Running; no reported local health warnings')
    else:
        report('WARN', 'Tailscale local state', state)
    peer = os.environ.get('STATUS_TAILSCALE_PEER', '')
    if not peer:
        report('COULD NOT CHECK', 'Tailscale peer reachability', 'Set STATUS_TAILSCALE_PEER to an expected peer hostname or IP')
    elif peer.startswith('-'):
        report('COULD NOT CHECK', 'Tailscale peer reachability', 'Invalid peer value')
    else:
        code, out, err = run(['tailscale', 'ping', '--c', '1', '--timeout', '5s', peer], timeout=8)
        report('OK' if code == 0 else 'WARN', 'Tailscale peer reachability', out or err)


def summary():
    print('\n===== NEEDS ATTENTION — ADDED HEALTH CHECKS =====')
    for level in ('FAIL', 'WARN', 'COULD NOT CHECK'):
        items = [r for r in RESULTS if r[0] == level]
        if items:
            print('\n{} ({})'.format(level, len(items)))
            for _, name, detail in items:
                print('  • {}: {}'.format(name, detail))
    counts = {level: sum(r[0] == level for r in RESULTS) for level in ('OK', 'WARN', 'FAIL', 'COULD NOT CHECK')}
    print('\n' + ' | '.join('{}: {}'.format(k, v) for k, v in counts.items()))
    if counts['FAIL'] or counts['WARN']:
        print('Please review the items above. No automatic repairs were attempted by these added checks.')
    elif counts['COULD NOT CHECK']:
        print('Completed checks passed, but coverage is incomplete — unchecked items are not a clean bill of health.')
    else:
        print('All added checks passed for this snapshot.')
    print('The original inventory/report remains above; its unstructured output is not included in these totals.')


def main(domains):
    domains = list(dict.fromkeys(domains))
    jobs = [('Website responses', lambda: websites(domains)),
            ('Served TLS certificates', lambda: certificates(domains)),
            ('Backup evidence', backups), ('Storage thresholds', storage),
            ('Reboot status', reboot), ('Scheduled jobs and failed units', timers),
            ('Container health', containers), ('Database responsiveness', databases),
            ('Recent hardware/resource trouble', hardware_logs),
            ('Clock synchronization', clock_status), ('Tailscale connectivity', tailscale_status)]
    for title, check in jobs:
        print('\n===== {} ====='.format(title), flush=True)
        try:
            check()
        except Exception as e:
            report('COULD NOT CHECK', title, '{}: {}'.format(type(e).__name__, e))
    summary()


if __name__ == '__main__':
    main(sys.argv[1:])
SERVER_STATUS_HEALTH_PY
  then
    :
  else
    echo "[COULD NOT CHECK] Enhanced checker failed; its report may be incomplete."
  fi
else
  section "NEEDS ATTENTION — ADDED HEALTH CHECKS"
  echo "[COULD NOT CHECK] Python 3 is missing; the original inventory ran, but added checks did not."
fi

# --- Final summary banner ---
echo ""
trap - EXIT    # disable the EXIT trap so it won't print a second footer
echo "===== DONE ====="
kv "Logfile" "${LOGFILE}"
kv "Version" "${VERSION}"
echo "===== ${NICKNAME} ${BANNER_IP} ${BANNER_NAME} — ${BANNER_DESC} ====="
exit 0