110 lines
4.8 KiB
Markdown
110 lines
4.8 KiB
Markdown
---
|
|
name: diagnose-crash
|
|
description: >
|
|
Diagnose why a program crashed on this Panama machine, from a systemd-coredump core dump.
|
|
Use when a process has segfaulted, aborted, or otherwise dumped core, when asked why an
|
|
application crashed or disappeared, when launched by panama-agent-crash, or when a
|
|
"stopped unexpectedly" desktop notification is acted on. Triggers: crash, segfault,
|
|
SIGSEGV, SIGABRT, core dump, coredumpctl, "why did X crash", "X keeps crashing",
|
|
backtrace symbolization.
|
|
---
|
|
|
|
<!--
|
|
Adapted from Omarchy's diagnose-crash skill (default/agents/skills/diagnose-crash/SKILL.md).
|
|
Copyright (c) David Heinemeier Hansson. MIT License: permission is hereby granted, free of
|
|
charge, to any person obtaining a copy of this software and associated documentation files,
|
|
to deal in the Software without restriction; the above copyright notice and this permission
|
|
notice shall be included in all copies or substantial portions of the Software.
|
|
-->
|
|
|
|
# Diagnosing a Crash
|
|
|
|
Work from evidence. The goal is an honest account of what happened, not a
|
|
plausible-sounding story.
|
|
|
|
## Establish the facts
|
|
|
|
`coredumpctl info <pid>` is the starting point. Beyond the backtrace, note the
|
|
**command line** the process was started with — it usually reveals what the
|
|
program was working on when it died, which is often the whole answer.
|
|
|
|
`coredumpctl list` shows whether this crash is a one-off or a pattern. Repeated
|
|
crashes of the same program, or several programs dying together, point somewhere
|
|
different than a single failure does. On this machine, Panama's System Health
|
|
page keeps a crash count too — `panama doctor check desktop.portal-stability`
|
|
knows whether the portal backend has been dying, which it chronically does.
|
|
|
|
## Rule out the boring causes first
|
|
|
|
Check resource exhaustion before blaming the program: `free -h`, and the journal
|
|
for OOM kills. A process killed by the OOM killer is not a bug in that process.
|
|
|
|
## Correlate against the timeline
|
|
|
|
The crash timestamp is the most underused piece of evidence. Compare it against:
|
|
|
|
- **Filesystem mtimes.** A file whose mtime lands on the same second as the
|
|
crash strongly suggests what triggered it.
|
|
- **The journal** around that moment, for related warnings from the same or
|
|
neighbouring processes.
|
|
- **Recent package updates.** `rpm -qa --last | head` — a crash that starts
|
|
right after an update points at the update. `dnf history` shows what a recent
|
|
transaction actually changed.
|
|
|
|
## Read the whole core, not just frame 0
|
|
|
|
Thread stacks other than the crashing one show what work was **in flight**.
|
|
That context often explains the trigger even when the crashing frame itself
|
|
cannot be symbolized. Note any third-party code in the address space —
|
|
extensions, plugins, out-of-tree drivers — but do not pin blame on it without
|
|
evidence that it is actually implicated.
|
|
|
|
## Symbolize when you can
|
|
|
|
This is Fedora, which runs a public debuginfod server:
|
|
|
|
```bash
|
|
core=$(mktemp -t crash-XXXXXX.core)
|
|
trap 'rm -f "$core"' EXIT
|
|
coredumpctl dump <pid> --output="$core"
|
|
DEBUGINFOD_URLS="https://debuginfod.fedoraproject.org/" \
|
|
gdb -q <executable> "$core" \
|
|
-batch -ex 'set debuginfod enabled on' -ex 'bt'
|
|
```
|
|
|
|
A core is a verbatim copy of the process's memory and can hold passwords,
|
|
tokens, and private documents. Write it to a fresh `mktemp` path rather than a
|
|
predictable shared one, and delete it when you are done — never leave it in
|
|
`/tmp`.
|
|
|
|
Many packages publish no debug symbols. When frames stay unresolved, say so —
|
|
**never invent function names to fill the gap.** An unsymbolized stack still
|
|
has shape: which library each frame belongs to, and whether the crash came from
|
|
a signal handler, a main loop, or a worker thread.
|
|
|
|
## Report
|
|
|
|
1. What crashed, and what it was doing at the time.
|
|
2. The most likely mechanism — separating clearly what the evidence **proves**
|
|
from what you are **inferring**.
|
|
3. Whether any user data was lost, and where it can be recovered from. Check
|
|
the trash before concluding anything is gone.
|
|
4. Whether it is likely to recur, and what would avoid or fix it.
|
|
|
|
Be straight about the limits of the evidence. If the cause is genuinely
|
|
ambiguous, say so rather than assembling confidence out of guesswork.
|
|
|
|
**Leave the system as you found it.** Diagnosis reads; it does not fix, tidy,
|
|
or reconfigure. The one thing to clean up is your own: delete the core you
|
|
extracted above. If the user wants the fix applied, that is a second,
|
|
explicitly requested step — and anything needing root goes through
|
|
`panama-sudo --reason` (load the `panama-sudo` skill).
|
|
|
|
## If the crash is Panama's
|
|
|
|
The shell, the compositor config, and everything under `~/.local/share/Panama`
|
|
are the user's own repository — a crash there is fixable in place, not an
|
|
upstream report. Load the `panama` skill before touching the repo: it carries
|
|
the live-desktop rules that keep an investigation from becoming a second
|
|
outage.
|