Aller au contenu principal

Issues Encountered & Fixes

Real issues hit during this build — documented for future reference.


Issue 1 — SSH Permission Denied​

Symptom:

ubuntu@10.0.0.x: Permission denied (publickey)

Cause: cloud-init users block was overriding MAAS's SSH key injection.

Fix:

✔ Remove the entire users: block from cloud-init
✔ Let MAAS inject the SSH key from your profile
✔ Redeploy the node

Issue 2 — IPv6 Conflicts (Wrong Subnet Selected)​

Symptom: Node receives an IPv6 address instead of 10.0.0.x, or MAAS deploys to wrong subnet.

Cause: MAAS was selecting the IPv6 subnet (2a02:...) over the intended 10.0.0.0/24.

Fix:

✔ Delete the IPv6 subnet from MAAS UI (Subnets → delete)
✔ Disable IPv6 via cloud-init sysctl
✔ Verify only 10.0.0.0/24 has DHCP enabled

Issue 3 — Alias Interface (enp0s31f6:1)​

Symptom: MAAS shows two interfaces for one NIC, causing IP conflicts or failed commissioning.

Fix:

✔ Delete the alias interface in MAAS machine network config
✔ Keep only the primary interface (enp0s31f6)
✔ Recommission the node

Issue 4 — MAAS 502 Error​

Symptom: Accessing http://10.0.0.1:5240/MAAS returns 502 Bad Gateway.

Cause: MAAS URL binding was pointing to wrong address after installation.

Fix:

sudo snap set maas url=http://10.0.0.1:5240/MAAS
sudo snap restart maas

Issue 5 — Node Stuck at "Disk Erasing"​

Symptom: Node deployment hangs indefinitely at the disk erasing phase.

Fix:

1. Abort the deployment from MAAS UI
2. Mark node as Broken
3. Mark node as Ready (via Actions)
4. Redeploy

The node will go through commissioning again cleanly.


Issue 6 — PXE Boot Loop (dhcpd Missing After Controller Reboot)​

Symptom: Node powers on, shows Lenovo logo, attempts "PXE boot over IPv4", then resets and loops endlessly — never reaches Ubuntu.

Diagnose:

# Run on the MAAS controller (10.0.0.1)
pgrep -af dhcpd

If this returns no output, dhcpd is not running. Nodes send DHCP DISCOVER on boot but receive no response, so PXE times out and the machine resets.

Cause: This was originally diagnosed as a dhcpd "crash", but log analysis showed the real cause is a boot-time startup race inside the MAAS snap:

  1. On controller boot, pebble (MAAS's internal service supervisor) starts regiond, apiserver, and rackd in parallel.
  2. rackd calls regiond's HTTP endpoint at http://10.0.0.1:5240/MAAS to fetch the DHCP config.
  3. If regiond isn't yet listening when rackd asks, rackd logs "Region is not advertising RPC endpoints", retries a few times, and gives up without ever telling pebble to start dhcpd.
  4. From the user's perspective the MAAS UI works (regiond + http are up), but the cluster nodes can't PXE-boot.

You can confirm this in the journal — look for these lines around boot time:

journalctl --since "<controller boot time>" | grep -E "(rackd.*Region|dhcpd)"

A failed boot shows Region not available: Connection refused and no dhcpd start lines. A successful boot shows Service "dhcpd" starting.

Manual fix (still useful for ad-hoc situations):

sudo snap restart maas

Wait ~30 seconds, then power-cycle the affected nodes. By the time rackd asks regiond for RPC info on a clean restart, regiond is already listening, so dhcpd starts cleanly.

Verify dhcpd is back:

pgrep -af dhcpd
# Should show two lines: one with `-f -4` (IPv4) and one with `-f -6` (IPv6)

Permanent fix — boot reconciler​

A small systemd timer fires 120 s after every boot, checks whether dhcpd is running, and runs snap restart maas automatically if it isn't. This makes the cluster self-healing on cold boot — no manual intervention needed.

Three files:

/usr/local/sbin/maas-dhcpd-reconciler

#!/bin/bash
# Restart the MAAS snap once if dhcpd didn't come up at boot.
# Triggered by maas-dhcpd-reconciler.timer ~120s after boot.

set -euo pipefail
LOG_TAG="maas-dhcpd-reconciler"

if pgrep -f '/snap/maas/.*/usr/sbin/dhcpd -f -4' >/dev/null; then
logger -t "$LOG_TAG" "dhcpd is running; nothing to do"
exit 0
fi

logger -t "$LOG_TAG" "dhcpd not running 120s after boot; restarting MAAS snap"
/usr/bin/snap restart maas
logger -t "$LOG_TAG" "MAAS snap restart complete"

/etc/systemd/system/maas-dhcpd-reconciler.service

[Unit]
Description=Restart MAAS snap if dhcpd did not start at boot
After=snap.maas.pebble.service network-online.target
Wants=network-online.target

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/maas-dhcpd-reconciler
StandardOutput=journal
StandardError=journal

/etc/systemd/system/maas-dhcpd-reconciler.timer

[Unit]
Description=Reconcile MAAS dhcpd 120s after boot

[Timer]
OnBootSec=120s
Unit=maas-dhcpd-reconciler.service

[Install]
WantedBy=timers.target

Install:

sudo chmod +x /usr/local/sbin/maas-dhcpd-reconciler
sudo systemctl daemon-reload
sudo systemctl enable --now maas-dhcpd-reconciler.timer

Verify it's enabled:

systemctl list-timers --all | grep maas-dhcpd-reconciler
journalctl -t maas-dhcpd-reconciler -n 20

After installation, every boot logs either "dhcpd is running; nothing to do" (happy path) or "dhcpd not running 120s after boot; restarting MAAS snap" (race hit, auto-recovered).

:::tip Boot order still recommended The reconciler removes the requirement to power on the controller before the cluster nodes, but it still adds ~2 minutes of recovery time on a bad boot. Powering on the MAAS controller first and waiting ~30 seconds is still the cleanest sequence. :::


Issue 7 — Node Boots with Wrong Hostname (Auto-Renamed by MAAS)​

Symptom: Node boots successfully but the login screen shows a random adjective-animal hostname (e.g. needed-lion) instead of the correct name (fast-heron, set-hog, etc.).

Cause: During a PXE boot loop, MAAS can accidentally trigger a re-deploy and assign the node a new auto-generated hostname. The OS gets installed with that temporary name.

Fix:

# SSH in using the IP (still correct even if hostname is wrong)
ssh ubuntu@10.0.0.7

# Set the correct hostname
sudo hostnamectl set-hostname fast-heron

# Update /etc/hosts to match
sudo sed -i 's/needed-lion/fast-heron/g' /etc/hosts

# Exit and verify
exit
ssh ubuntu@10.0.0.7 "hostname"

Then update MAAS to stay in sync (run on controller):

maas admin machine update q6m3px hostname=fast-heron

Replace q6m3px with the correct system_id for the affected node:

Nodesystem_id
set-hognbc6cx
fast-skunksby3w7
fast-heronq6m3px

Issue 8 - MAAS Temporal Power Dispatch Fails​

Symptom (MAAS 3.7, webhook power type):

Error: Unexpected exception: UnroutablePowerWorkflowException:
Error determining BMC task queue for machine nbc6cx

Cause:

MAAS 3.7 routes power commands through the Temporal workflow engine. The function get_temporal_task_queue_for_bmc must resolve which Temporal task queue to dispatch the power-on activity to. It does this in two steps:

  1. Extract the BMC IP from power_query_uri via ip_extractor, map it to a subnet → VLAN
  2. Fallback: look up BmcRoutableRackControllerRelationship entries in the DB

If the webhook URIs use 127.0.0.1, MAAS can extract the IP but cannot match loopback to any managed subnet → both paths fail → UnroutablePowerWorkflowException is raised before any workflow is started.

Fix:

Bind the power broker on all interfaces and use the management network IP in MAAS power parameters:

# 1. Update power broker to listen on 0.0.0.0 (edit /home/ktayl/bin/maas-power-broker.py)
# Change: HTTPServer(("127.0.0.1", 5241), Handler)
# To: HTTPServer(("0.0.0.0", 5241), Handler)

# 2. Restart the broker
systemctl --user restart maas-power-broker.service

# 3. Update power parameters for each machine to use 10.0.0.1 instead of 127.0.0.1
maas ktayl machine update <system_id> power_type=webhook power_parameters='{
"power_on_uri": "http://10.0.0.1:5241/power/on/<hostname>",
"power_off_uri": "http://10.0.0.1:5241/power/off/<hostname>",
"power_query_uri":"http://10.0.0.1:5241/power/query/<hostname>",
"power_on_regex": "status.*\\:.*running",
"power_off_regex":"status.*\\:.*stopped",
"power_verify_ssl":"n",
"power_user":"", "power_pass":"", "power_token":""
}'

After the update, MAAS maps 10.0.0.1 → subnet 10.0.0.0/24 → VLAN 5001, creates a BMC routing relationship, and dispatches power activities to agent:power@vlan-5001. All maas ktayl machine power-on/query-power-state commands work.

Why 10.0.0.1: That is the controller's IP on enx606d3cdf7160 (the cluster management interface). MAAS manages 10.0.0.0/24 (VLAN 5001, DHCP enabled), so it can map this address to a VLAN.

Architecture note: The Go maas-agent binary registers Temporal activity workers on two task queues: <system_id>@agent:main (DHCP, config) and agent:power@vlan-<id> (power activities). The Python maas-temporal-worker runs the PowerOnWorkflow on the shared region queue and dispatches the power-on activity to the agent's power queue. The webhook driver activity executes the HTTP call to the power broker.