Let somebody extend this without forking it, and say when things die
Two of Section F. Hooks are the pressure valve. "Can Panama also do X when the theme changes" is now a five-line file in ~/.config/panama/hooks rather than a fork, a feature request, or a patch somebody rebases forever. Each name takes a single file and a .d directory so several things can react without fighting over one, and a broken hook is reported and stepped over: somebody's script must never cost a theme change, an upgrade or a login. Wired at theme-set, post-upgrade and post-migrate. This is the thirty-line version of the plugin host the upstream ledger defers, and it has no API to keep stable beyond "we will run your script and tell you what happened". Testing it caught a real bug the reading would not have: run_one captured the script path but never shifted it off, so every hook got its own filename as $1 and the real arguments arrived one place late. A hook reading $1 as the colour scheme got a path. The crash watcher notices when a program dumps core and says so. Under GNOME, ABRT does this; here nothing did, and applications died silently, which is most of how "Linux is flaky" gets earned. Once per program per session is the entire design, not a nicety. This machine's portal backend crashes between eleven and sixty times a day, and a notification per crash would be one every few minutes for something nobody can act on. The first is news; the fortieth is why people turn notifications off. The health page keeps the running count. It waits for the notification server before reporting, because the crash most worth hearing about is the one that took the shell with it, and it names the executable rather than the kernel's comm field, which truncates at fifteen characters. Verified against real segfaults.
This commit is contained in:
Executable
+73
@@ -0,0 +1,73 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
# Tell somebody when a program crashes.
|
||||
#
|
||||
# On GNOME, ABRT says so. Under a hand-assembled Hyprland desktop nothing does,
|
||||
# and applications die silently -- which is most of how "Linux is flaky" gets
|
||||
# earned. Fedora ships systemd-coredump by default, so the information is
|
||||
# already there; nobody is reading it.
|
||||
#
|
||||
# Follows the journal for systemd-coredump's own message id and reports each
|
||||
# program once per session.
|
||||
#
|
||||
# ONCE PER SESSION IS THE WHOLE DESIGN. This machine's portal backend crashes
|
||||
# between eleven and sixty times a day -- see the portal-stability check in
|
||||
# panama-doctor -- and a notification per crash would be a notification every
|
||||
# few minutes for something the user can do nothing about. The first one is
|
||||
# news; the fortieth is why people turn notifications off. The health page
|
||||
# carries the running count for anyone who wants it.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
PANAMA_PATH="${PANAMA_PATH:-$HOME/.local/share/Panama}"
|
||||
|
||||
# systemd-coredump's MESSAGE_ID. Matching on this rather than on text keeps
|
||||
# working when the wording changes and never matches a program that merely
|
||||
# mentions the word "crash" in its own logs.
|
||||
readonly COREDUMP_MESSAGE_ID='fc2e22bc6ee647b6b90729ab34a250b1'
|
||||
|
||||
command -v journalctl >/dev/null 2>&1 || exit 0
|
||||
command -v notify-send >/dev/null 2>&1 || exit 0
|
||||
|
||||
# The shell owns org.freedesktop.Notifications, and the crash most worth
|
||||
# reporting is the one that took the shell with it. Waiting means that report
|
||||
# arrives rather than vanishing into a bus nobody is serving.
|
||||
for _ in $(seq 1 60); do
|
||||
busctl --user status org.freedesktop.Notifications >/dev/null 2>&1 && break
|
||||
sleep 1
|
||||
done
|
||||
|
||||
declare -A reported=()
|
||||
|
||||
# -f from now, not from the boot: a session that starts after a crash should
|
||||
# not open with a notification about something the user has already lived
|
||||
# through and cannot act on.
|
||||
journalctl --user -f -n 0 --output=json MESSAGE_ID="$COREDUMP_MESSAGE_ID" 2>/dev/null \
|
||||
| while IFS= read -r line; do
|
||||
[[ -n "$line" ]] || continue
|
||||
|
||||
uid="$(jq -r '.COREDUMP_UID // empty' <<<"$line" 2>/dev/null)"
|
||||
exe="$(jq -r '.COREDUMP_EXE // empty' <<<"$line" 2>/dev/null)"
|
||||
comm="$(jq -r '.COREDUMP_COMM // empty' <<<"$line" 2>/dev/null)"
|
||||
|
||||
# Another user's crash is not this session's business, and reporting it
|
||||
# would leak what they are running.
|
||||
[[ "$uid" == "$(id -u)" ]] || continue
|
||||
[[ -n "$exe" || -n "$comm" ]] || continue
|
||||
|
||||
# The executable name first: COREDUMP_COMM is the kernel's comm field
|
||||
# and is truncated to fifteen characters, so it reports
|
||||
# "panama-test-cra" for a program called panama-test-crasher.
|
||||
if [[ -n "$exe" ]]; then
|
||||
program="$(basename "$exe")"
|
||||
else
|
||||
program="$comm"
|
||||
fi
|
||||
[[ -z "${reported[$program]:-}" ]] || continue
|
||||
reported[$program]=1
|
||||
|
||||
notify-send --icon=dialog-error-symbolic --app-name=Panama \
|
||||
"$program stopped unexpectedly" \
|
||||
"It crashed and was not able to recover. System Health has the details." \
|
||||
2>/dev/null || true
|
||||
done
|
||||
Reference in New Issue
Block a user