Skip to main content

Troubleshooting

Solutions to common problems and issues with Overdeck.

PR stuck with stale failed checks after main recovers

A pull request may remain UNSTABLE with a failed test check and ready_for_merge=0 even though its review, test, and verification stages passed and main is green. ready_for_merge=0 means the merge gate is deliberately holding the pull request out of the merge queue because GitHub still reports a required check failure. This happens when the pull request’s check ran while main was failing and GitHub did not run it again after main recovered. The dashboard’s stale-check re-trigger service detects these inherited failures and re-runs them once. A red window is the period from a completed failing run on main until the next completed successful run of the same workflow. The service only acts on a pull request run created inside a closed red window, never while that workflow is still failing on main. The service also requires the GitHub workflow run to have attempt: 1. This once-only attempt guard means a run that Overdeck already retriggered, including one that failed again, is not put into a retry loop. Decisions and actions are logged with the [stale-check-retrigger] prefix. If the dashboard is down, the run predates the available main history, or the run is already on attempt 2 or later, re-run the failed jobs manually:
After GitHub reports the new result, the existing merge-blocker reconciliation poll clears the stale blocker within about 10 minutes.

pan start refuses: “references files that no longer exist”

pan start runs a plan-freshness preflight right before it spawns a work agent: it refuses a missing files_scope path only when a commit on the workspace’s HEAD deleted that path after the plan was created. If the listed files really did move or get deleted since planning, re-plan against the current tree with pan plan <ID>. If the files are ones the plan legitimately creates or restores, the preflight should already let it through — --skip-freshness spawns the agent anyway without waiting for a re-plan.

Local main diverged after a backlog or order-book write

pan backlog write-sequence and pan orders commit their .pan/ write and push it so local main doesn’t drift ahead of origin. When that push runs into trouble, you’ll see a Sequence push: ... or Order book push: ... warning — printed to the CLI’s stderr and, since PAN-4224, also filed as a plan-artifacts warning in the dashboard activity feed. Two causes produce this warning:
  • An unrelated local commit. Local main holds a commit outside .pan/ that the push refuses to carry along. Push it yourself, then re-run the .pan/ write.
  • A colliding untracked file. Origin gained a .pan/ file (through a merge elsewhere) that this checkout already has as an untracked file at the same path — most often a planning-promotion draft. The push moves that file aside before it can move local main onto the new tip: if its bytes matched what’s now tracked it’s deleted, otherwise it’s kept under .overdeck/plan-artifact-backups/<stamp>/<path> and named in the warning.
To recover:
  1. If the warning names a backed-up file, compare .overdeck/plan-artifact-backups/<stamp>/<path> against the file now tracked at <path> and merge in whatever local edit is worth keeping.
  2. Run git pull --rebase in the plan home to bring local main back in sync with origin.
  3. Re-run the .pan/ write (pan backlog write-sequence or the pan orders command) if it didn’t land.

WSL2 Stability Issues (Windows Users)

WSL2 can experience crashes and networking issues, especially under heavy AI agent workloads. Here are recommended .wslconfig settings to improve stability. Create/edit C:\Users\<username>\.wslconfig:
After changing .wslconfig:
Verify settings applied:

Windows 10 Limitations

Windows 10 users: Most advanced WSL2 features require Windows 11. On Windows 10, only the basic resource limits and localhostForwarding/guiApplications settings are supported. Other settings will be silently ignored or may cause instability. If you experience frequent WSL2 crashes on Windows 10, consider:
  • Using only the basic settings shown above (memory, processors, swap, localhostForwarding, guiApplications)
  • Reducing memory allocation if system is under pressure
  • Upgrading to Windows 11 for full WSL2 feature support
  • Checking Windows Event Viewer for specific crash causes

Additional Windows 10 Workarounds

If NAT networking is unstable on Windows 10:
Common conflict sources:
  • VPN clients (especially corporate VPNs)
  • Docker Desktop (can conflict with WSL networking)
  • Third-party firewalls
  • Hyper-V virtual switch issues
If NAT fails completely, WSL 2.3.25+ automatically falls back to VirtioProxy mode. This is less performant but more stable. You’ll see: "Failed to configure network (networkingMode Nat), falling back to networkingMode VirtioProxy." References:

Slow Vite/React Frontend with Multiple Workspaces

If running multiple containerized workspaces with Vite/React frontends, you may notice CPU spikes and slow HMR. This is because Vite’s default file watching polls every 100ms, which compounds with multiple instances. Fix: Increase the polling interval in your vite.config.mjs:
A 3000ms interval supports 4-5 concurrent workspaces comfortably while maintaining acceptable HMR responsiveness.

Dev server dies with ENOSPC — inotify watch exhaustion

If a workspace’s frontend container crash-loops at startup with ENOSPC: System limit for number of file watchers reached while the rest of the host looks healthy, the per-user inotify watch budget (fs.inotify.max_user_watches) is exhausted. This budget is a kernel limit shared by every process and container the user runs — Docker does not isolate it — so a few heavy file watchers can starve every other workspace on the machine. The dashboard shows a File watchers running low / exhausted banner when usage crosses 80% / 90% of the limit, and pan doctor reports usage, the top consumers, and whether the configured limit survives a reboot. Fix, in order of leverage:
  1. Shrink the watchers. The usual culprit is a huge directory inside the watched project root that the watcher does not ignore by default — e.g. pnpm’s .pnpm-store/ (pnpm places it inside the project when installing into a Docker bind mount, and Vite ignores node_modules but not .pnpm-store). Add it to the watch ignore list:
    In one real incident this cut each dev server from ~157k watches to ~13k — a 12× reduction.
  2. Raise and persist the limit (requires sudo; Overdeck never runs sudo itself):
    A sysctl -w alone does not survive a reboot; the /etc/sysctl.d file does. pan doctor warns when the live limit is higher than anything persisted.

Corrupted Workspaces

A workspace can become “corrupted” when it exists as a directory but is no longer a valid git worktree. The dashboard will show a yellow “Workspace Corrupted” warning with an option to clean and recreate.

Symptoms

  • Dashboard shows “Workspace Corrupted” warning
  • git status in the workspace fails with “not a git repository”
  • The .git file is missing from the workspace directory

Common Causes

Resolution

Via Dashboard (recommended):
  1. Click on the issue to open the detail panel
  2. Click “Clean & Recreate” button
  3. Review the files that will be deleted
  4. Check “Create backup” to preserve your work (recommended)
  5. Click “Backup & Recreate”
Via CLI:

Prevention

  • Don’t interrupt pan workspace create commands
  • Don’t run git worktree prune in the main repo without checking for active workspaces
  • Ensure adequate disk space before creating workspaces

Docker Issues

Container won’t start

”No such network: overdeck”

Permission denied on mounted volumes

If containers run as root and create files, you won’t be able to delete them:

Network Issues

HTTPS not working

  1. Check certificates exist:
  2. Regenerate if missing:
  3. Install the CA:

Can’t reach workspace URLs

  1. Check Traefik is running:
  2. Check DNS resolution:
  3. Check Traefik dashboard (http://localhost:8080) for routing rules

Agent Issues

Agent stuck / not responding

To attach from a shell, run the Attach: command that pan start MIN-123 prints (it prints it again when the agent is already running), or open the agent in the dashboard. See “Reaching an agent from a shell” in docs/TERMINAL-BACKENDS.md.

Agent keeps failing

Check the handoff count in state.json. If it’s high, the task may be too complex:
Consider:
  • Breaking the issue into smaller tasks
  • Adding more context to the issue description
  • Manually handling complex parts

Messages not reaching agent

Use the proper messaging API:

Command Deck Issues

Command Deck shows “Unknown project”

A URL such as /command-deck/<slug> shows an Unknown project state when <slug> matches neither a registered project key nor its display name. This usually means the URL is stale, the project was renamed, or the project is not registered on this machine. The page does not open a functional project deck for that slug. Use the registered-project buttons in the recovery panel to navigate to a valid deck, or select Back to Command Deck to return to /command-deck. To inspect the registered keys from the CLI, run:

Conversation creation fails from a launcher

When a launcher cannot create a conversation, the failure appears inline beside the composer and the typed query remains available for correction or retry. The launcher blocks duplicate keyboard and pointer submissions while creation is in progress, and it opens a conversation pane only after the server reports a successful creation. If the inline error reports an unknown project, use the recovery steps above to open a registered project deck before retrying.

Workspace Issues

Workspace creation fails

Can’t delete workspace

If containers created root-owned files:

Performance Issues

Dashboard slow to load

High CPU usage

  • Check number of concurrent workspaces
  • Increase Vite polling interval (see above)
  • Run docker stats to identify resource-heavy containers

High memory usage

Getting Help

Health page

The dashboard’s Health page reports live host, admission, agent, and optional-service health. Use it when the running fleet looks unhealthy: it shows current pressure evidence, spawn headroom, agent/session state, and service state without turning unavailable measurements into zeroes.
These surfaces answer different questions. pan health reports runtime health of Overdeck services, while pan doctor checks dependencies, installation, configuration, and broader system diagnostics. They overlap, but neither CLI command is a textual equivalent of the live Health page. pan doctor also reports a Tiered execution row that warns when a tier’s model class does not fit the difficulties it owns. Its Plan home .pan/ tracking row warns when a registered project’s plan home git-ignores .pan/ (so planning artifacts cannot be committed), with the rule’s file:line; Overdeck’s legacy .pan/ line is removed and committed by pan admin migrate-plan-home <key> --commit.

Diagnostic Information

When reporting issues, include:

Resources