Files
LEDMatrix/docs/SSH_UNAVAILABLE_AFTER_INSTALL.md
T
0c5b9c57d3 fix: keep low-memory boards reachable under load (#464)
* fix(service): survive corrupt health cache and clean exits

Three independent failure modes that each end with a dark panel and no
automatic recovery.

1. PluginHealthTracker._load_health_state returned the cached value
   verbatim. If that value is not a dict, every caller raises
   AttributeError: 'list' object has no attribute 'get' — during
   DisplayController.__init__, so the process dies before the display
   loop starts. systemd restarts it, the same bad entry is read back
   from disk, and it dies again: an unattended restart loop that
   survives reboots because the cause is persisted. Observed in the
   field with plugin_health:<id> holding an unrelated plugin's list
   payload. Now non-dict entries are discarded with a warning and the
   defaults are rebuilt.

2. ledmatrix.service used Restart=on-failure, so any exit with status 0
   left the unit stopped and the panel dark indefinitely — systemd
   treats it as success and never brings it back. Restart=always.

3. ledmatrix-wifi-monitor.service used StandardOutput=syslog, which
   systemd has marked obsolete; it warns and rewrites it to journal on
   every load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(memory): size the cache to the board and stop reinstalling deps

On a 1GB Pi 3B+ the display process settles around 600MB RSS of 905MB
total. When the remaining headroom runs out the failure is not a clean
crash: fork() starts returning ENOMEM, so sshd accepts connections and
closes them before its banner, timer jobs stop running, and the panel
goes dark, while already-resident processes keep serving normally. The
board looks healthy from outside and cannot be logged into. Only a power
cycle clears it.

Three contributing causes:

- MemoryCache had a fixed 1000-entry ceiling. Entries are parsed API
  payloads of tens of KB, so one ceiling cannot serve both a 512MB Zero
  2 W and an 8GB Pi 5. Now scaled from MemTotal (150 entries at <=1GB,
  1500 at >=8GB), overridable with LEDMATRIX_CACHE_MAX_ENTRIES.

- requirements_are_satisfied() returned False for any requirement with
  extras, so a plugin depending on python-socketio[client] re-ran pip on
  every single start: ~8s, a network dependency, and a 100-200MB spike
  at the least convenient moment. During a restart loop it repeats for
  each restart. Extras are now resolved one level deep against installed
  metadata, keeping the conservative "anything unverifiable falls
  through to pip" contract.

- ledmatrix.service had no memory ceiling. MemoryMax=85% expressed as a
  percentage so one unit file suits every board. Note this needs the
  memory cgroup controller, which Pi firmware disables by default;
  first_time_install.sh now adds cgroup_enable=memory to cmdline.txt,
  and the unit file documents how to verify it took effect.

first_time_install.sh also enables persistent journald storage (capped
at 64M). Default storage is volatile, so every reboot destroys the logs
that would explain why the board rebooted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: guidance for 512MB and 1GB boards

Documents the memory ceiling on small boards and, more usefully, what
running into it actually looks like: sshd accepting connections and
closing them before the banner, the web UI still responding normally,
clean ping, a dark panel, and a wrong clock after the next boot. None of
those read as "out of memory", which makes the failure hard to identify
from the symptoms.

Cross-referenced from SSH_UNAVAILABLE_AFTER_INSTALL.md, since "I can't
SSH in any more" is how most people will first meet this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address review findings on the low-memory work

Nine CodeRabbit findings, five in code.

**Health state (the one that matters).** The non-dict guard did not cover a
dict missing fields the callers index directly, which is the shape actually
seen in the wild: a record carrying only circuit_state produced
`plugin clock-simple operation failed: 'circuit_state'` about fifty times a
minute with the panel frozen. The record is now completed against the
defaults per field rather than trusted or discarded wholesale. Per field
matters: a first pass rejected any incomplete record outright, which reset a
tripped breaker and real failure counts to healthy because one optional
field was absent -- an existing test caught it. Values of the wrong type
(a counter persisted as a string, an unknown circuit_state) fall back
individually, valid neighbours survive, and newer fields the schema has
grown since (degraded, degraded_reason) are carried through untouched.

**Cache ceiling.** MemoryCache.set() accepted entries without bound between
cleanup sweeps, which run every 300s by default, so a burst could take the
cache far past max_size -- the unbounded growth the limit exists to stop.
Eviction now runs under the same lock on every write, sharing one helper
with the periodic sweep so the two cannot drift.

**Installer, cgroups.** Only cgroup_enable=memory was checked, so a board
carrying that without cgroup_memory=1 reported success and got no change,
leaving MemoryMax= inert. Each parameter is now checked and appended
independently; verified against all four combinations, single line preserved.

**Installer, journald.** Persistence was inferred from /var/log/journal being
non-empty, which proves neither Storage=persistent nor a size cap -- the
directory survives a switch back to volatile. The effective configuration is
read instead (systemd-analyze cat-config, falling back to the conf files),
and an explicitly configured SystemMaxUse is preserved rather than
overwritten. Verified across volatile, persistent-without-cap,
persistent-with-user-cap, cap-without-storage, and commented-only configs.

**Dependency extras.** _extras_are_satisfied stopped at one level, so a
gated dependency that itself requests an extra (requests[socks]) passed on
the base distribution's version while the extra's own dependency was
missing, and pip was skipped. It now recurses, with a visited
(distribution, extras) set so a cycle terminates.

Docs: both kernel command-line paths documented (the installer falls back to
/boot/cmdline.txt), daemon-reload and restart added after the systemd
override example, memory exhaustion added to the SSH summary with its
power-cycle-only recovery, and a language on the fenced block for MD040.

Tests: five for the health-state repair including the exact wild shape and
that record_failure/record_success no longer raise against it, and one for
the cache ceiling. Both mutation-checked. Full suite 2927 passed, with the
one pre-existing tmpfs failure that also fails on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix: harden the health-state repair and confirm journald took effect

Second review round; all three findings were valid and two were bugs in the
repair added last commit.

The repair could raise out of itself. An unhashable circuit_state (a list or
dict on disk) hit `value in {...}` and raised TypeError -- from the code
whose whole job is to stop a malformed record crashing the caller. It now
requires a str before the membership test.

bool is a subclass of int, so True passed the timestamp check and then
compared as 1.0: enough to expire a cooldown the instant the breaker opened,
while False would stop the elapsed check firing at all. Timestamps now
exclude bool explicitly.

The regression test for the original crash was seeded with a record that
*contained* circuit_state, so it passed against the old raw-return behaviour
too -- the counters are read with .get(), so circuit_state is the only field
whose absence used to raise. Reseeded to omit it, and it now fails against
raw-return as intended.

journald: drop-ins apply in lexical order, so a local file sorting after
ledmatrix-persistent.conf still wins and writing ours proves nothing. The
effective Storage is re-read afterwards and a warning naming the diagnostic
command is printed if persistence is still not active, rather than reporting
a success that was not verified.

Full suite 2934 passed, same single pre-existing tmpfs failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 12:28:22 -04:00

235 lines
6.8 KiB
Markdown

# SSH Unavailable After Installation - Troubleshooting Guide
## Why SSH Becomes Unavailable
After running `first_time_install.sh`, SSH may become unavailable for the following reasons:
### 1. WiFi Monitor Service Enables AP Mode
**Primary Cause**: The WiFi monitor service (`ledmatrix-wifi-monitor`) automatically enables Access Point (AP) mode when it detects that the Raspberry Pi is not connected to WiFi. When AP mode is active:
- The Pi creates its own WiFi network: **LEDMatrix-Setup** (password: `ledmatrix123`)
- The Pi's WiFi interface (`wlan0`) switches from client mode to AP mode
- **This disconnects the Pi from your original WiFi network**
- SSH becomes unavailable because the Pi is no longer on your network
### 2. Network Configuration Changes
The installation script:
- Installs and configures `hostapd` (Access Point daemon)
- Installs and configures `dnsmasq` (DHCP server for AP mode)
- These services can interfere with normal WiFi client mode
### 3. The Board Ran Out of Memory
On a 512MB or 1GB board, memory exhaustion stops `sshd` being able to fork a
session process. The connection is accepted and then closed immediately, before
any banner:
```text
kex_exchange_identification: Connection closed by remote host
```
The giveaway is that the board is otherwise healthy — ping is clean and the web
UI still responds — but nothing that needs to start a new process works, and
the panel is usually dark. Only a power cycle clears it. See
[LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md).
### 4. Reboot After Installation
If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode.
## How to Regain SSH Access
### Option 1: Connect to AP Mode (Recommended for Initial Setup)
1. **Find the AP Network**:
- Look for a WiFi network named **LEDMatrix-Setup** on your phone/computer
- Default password: `ledmatrix123`
2. **Connect to the AP**:
- Connect your device to the **LEDMatrix-Setup** network
- The Pi will have IP address: `192.168.4.1`
3. **SSH via AP Mode**:
```bash
ssh devpi@192.168.4.1
```
4. **Disable AP Mode and Reconnect to WiFi**:
Once connected via SSH:
```bash
# Check WiFi status
nmcli device status
# Disable AP mode manually
sudo systemctl stop hostapd
sudo systemctl stop dnsmasq
# Connect to your WiFi network (replace with your SSID and password)
sudo nmcli device wifi connect "YourWiFiSSID" password "YourPassword"
# Or use the web interface at http://192.168.4.1:5000
# Navigate to WiFi tab and connect to your network
```
### Option 2: Disable WiFi Monitor Service Temporarily
If you have physical access to the Pi or can connect via AP mode:
```bash
# Stop the WiFi monitor service
sudo systemctl stop ledmatrix-wifi-monitor
# Disable it from starting on boot (optional)
sudo systemctl disable ledmatrix-wifi-monitor
# Stop AP mode services
sudo systemctl stop hostapd
sudo systemctl stop dnsmasq
# Reconnect to your WiFi network
sudo nmcli device wifi connect "YourWiFiSSID" password "YourPassword"
```
### Option 3: Use Ethernet Connection
If your Pi is connected via Ethernet:
- SSH should remain available via Ethernet even if WiFi is in AP mode
- Connect via: `ssh devpi@<pi-ip-address>`
### Option 4: Physical Access
If you have physical access to the Pi:
1. Connect a keyboard and monitor
2. Log in locally
3. Follow Option 2 to disable AP mode and reconnect to WiFi
## Preventing SSH Loss in the Future
### Method 1: Configure WiFi Before Installation
Before running `first_time_install.sh`, ensure WiFi is properly configured and connected:
```bash
# Check WiFi status
nmcli device status
# If not connected, connect to WiFi
sudo nmcli device wifi connect "YourWiFiSSID" password "YourPassword"
# Verify connection
ping -c 3 8.8.8.8
```
### Method 2: Disable WiFi Monitor Service
If you don't need the WiFi setup feature:
```bash
# After installation, disable the WiFi monitor service
sudo systemctl stop ledmatrix-wifi-monitor
sudo systemctl disable ledmatrix-wifi-monitor
```
### Method 3: Configure WiFi Monitor to Not Auto-Enable AP
Edit the WiFi monitor configuration to prevent automatic AP mode:
```bash
# Edit the WiFi config (if it exists)
nano /home/devpi/LEDMatrix/config/wifi_config.json
# Or modify the WiFi monitor daemon behavior
# (requires code changes to wifi_monitor_daemon.py)
```
## Verification Steps
After regaining SSH access, verify your installation:
```bash
cd /home/devpi/LEDMatrix
./scripts/verify_installation.sh
```
This script will check:
- Systemd services status
- Python dependencies
- Configuration files
- File permissions
- Web interface availability
- Network connectivity
## Quick Reference Commands
```bash
# Check WiFi status
nmcli device status
nmcli device wifi list
# Check AP mode status
sudo systemctl status hostapd
sudo systemctl status dnsmasq
# Check WiFi monitor service
sudo systemctl status ledmatrix-wifi-monitor
# View WiFi monitor logs
sudo journalctl -u ledmatrix-wifi-monitor -f
# Connect to WiFi
sudo nmcli device wifi connect "SSID" password "password"
# Disable AP mode
sudo systemctl stop hostapd dnsmasq
# Restart network services
sudo systemctl restart NetworkManager
```
## Web Interface Access
Even if SSH is unavailable, you can access the web interface:
1. **Via AP Mode**: Connect to **LEDMatrix-Setup** network and visit `http://192.168.4.1:5000`
2. **Via WiFi**: If WiFi is connected, visit `http://<pi-ip-address>:5000`
3. **Via Ethernet**: Visit `http://<pi-ip-address>:5000`
The web interface allows you to:
- Configure WiFi connections
- Enable/disable AP mode
- Check service status
- View logs
- Manage the LED Matrix display
## Summary
**SSH becomes unavailable because** — two unrelated causes, and they need
different responses:
*AP mode (most common):*
- WiFi monitor service enables AP mode when WiFi disconnects
- AP mode switches WiFi from client to access point mode
- Pi loses connection to your original network
*Memory exhaustion (low-memory boards):*
- The board runs out of memory, so `sshd` cannot fork a session process
- The connection is accepted and closed before any banner
- Ping still answers and the web UI still responds, so it looks healthy
- The panel is usually dark and the service cannot restart
- **Only a power cycle clears this** — there is no remote recovery, because
every remote route needs a new process
- Prevention and tuning: [LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md)
**To regain SSH**:
1. Connect to **LEDMatrix-Setup** AP network (password: `ledmatrix123`)
2. SSH to `192.168.4.1`
3. Disable AP mode and reconnect to your WiFi network
4. Or disable the WiFi monitor service if not needed
**To prevent future issues**:
- Ensure WiFi is connected before installation
- Or disable WiFi monitor service if you don't need AP mode feature