OpenClaw Agent Troubleshooting Checklist: Sessions, Subagents, Permissions, and Network Failures
OpenClaw News Editorial Desk
Even the most advanced agents can run into trouble. Whether it's a connectivity issue or a permission error, knowing how to troubleshoot is key to maintaining a healthy OpenClaw environment.
1. Checking Logs
The first step in any investigation is checking the logs. OpenClaw provides a command to view recent activity:
openclaw logs --tail 100
Look for ERROR or WARN messages. Common culprits include:
- API Rate Limits: Check your provider dashboard.
- Network Timeouts: Verify your Gateway connection.
2. Common Errors & Fixes
"web_search needs a Brave Search API key"
If your agent fails to search the web, ensure your API key is configured:
openclaw configure --section web
# Enter your Brave Search API key
Or verify the environment variable BRAVE_API_KEY is set in your .env file.
"Permission denied" on exec
By default, OpenClaw runs commands with restricted permissions. To run privileged commands, use sudo within exec or adjust the security.exec policy in openclaw.toml.
Note: Always be cautious when granting elevated permissions.
Subagent Failures
If subagents are not responding:
- Check the main agent's session status:
openclaw session status. - Ensure the subagent process hasn't been terminated prematurely.
- Review the inter-agent communication logs.
3. Advanced Diagnostics
Monitoring Resource Usage
Agents can consume significant memory during long sessions. Use top or htop to monitor the openclaw-agent process.
Network Debugging
If the agent cannot connect to external services:
- Verify DNS settings.
- Check firewall rules blocking outgoing traffic on ports 80/443.
- Test connectivity with
curl -I https://api.openclaw.ai.
Who should use this troubleshooting checklist first
This page is the right starting point if you are in one of these situations:
- your agent used to work, but now fails after a config, model, or gateway change
- a subagent hangs, exits early, or never returns usable output
exec, web, or browser tools fail and you are not sure whether the blocker is permissions, credentials, or connectivity- you need to decide whether to restart, reconfigure, or verify an upstream provider first
If your issue is clearly limited to a first-time install on Mac, the installation or quick-install guides are usually a better first stop than a general troubleshooting checklist.
90-second routing for search visitors
If you only need the next step, choose the entry point by symptom first:
- The CLI is slow or frozen: check the CLI regression page before editing agent config.
- A tool call fails: identify whether it is exec, browser, web, or an external API before changing permissions or credentials.
- A subagent returns nothing: confirm the parent session is alive, then check whether the child run exited or is waiting on approval.
- The first task after installation fails: return to install validation before diving into provider routing.
The goal is to turn “OpenClaw does not work” into four actionable paths.
Common troubleshooting judgment questions
When should you restart first, and when should you inspect logs first
Restart first only when you already know the previous state was healthy and the failure looks transient. If the same symptom repeats, inspect logs before touching config, because repeated restarts hide whether the real cause is rate limiting, approval policy, or broken connectivity.
When is the real problem the provider or network, not the local agent
If sessions launch but stall on model calls, or web and external API tools fail together, suspect upstream API health, DNS, firewall, or provider limits before changing local agent logic. When multiple tools fail at the same outbound layer, local prompt changes rarely fix it.
If your symptom is really one of these two high-frequency failures, where should you go next?
If you already know this is not a generic agent failure, do not stay on the checklist longer than necessary. Go directly to the narrower page that matches the symptom:
- The CLI hangs for 20 to 40 seconds after hooks load, but Gateway health still looks normal: read OpenClaw CLI may hang for 20 to 40 seconds after v4.5+: why the Gateway looks healthy while the CLI feels dead to separate a CLI or RPC regression from a real machine, network, or gateway slowdown.
- Codex OAuth looks successful, but runtime still falls back to the default provider order or throws a No API key style failure: read Codex OAuth login succeeds but runtime still falls back to the default provider order to check whether per-agent auth order never landed or provider routing state stayed incomplete.
That split matters because a broad troubleshooting page should pass high-intent readers into the exact failure page quickly, instead of forcing them to reopen the same logs and settings from scratch.
If you have already landed on this checklist, what is the next judgment you should add first?
The biggest risk on a general troubleshooting page is not lack of information. It is that the reader knows something is broken, but still does not know which page to open next.
So if you have already landed here, the next useful judgment is usually:
- The CLI is obviously slow right after hooks load, while Gateway health still returns fast: go straight to the CLI regression troubleshooting page.
- Codex OAuth appears complete, but real requests still hit the wrong provider or provider order never takes effect: go straight to the Codex OAuth troubleshooting page.
- You just migrated from Moltbot or Clawdbot and suspect an old command, path, image tag, or remote is still in use: go back to the 1-minute migration guide.
The clearer that routing layer becomes, the more likely this checklist page can keep handing traffic into narrower, higher-intent troubleshooting pages instead of losing readers in a generic loop.
Add four evidence items before handing the issue to the next operator
If this checklist helped narrow the direction but the issue is not fully fixed, do not hand off a vague “still broken.” Add four evidence items first:
- Which layer is currently most suspicious: installation, gateway, message entrypoint, tool call, or model/provider. Name at least one likely layer.
- The last reproducible command or user input: keep the original command, session input, or smallest trigger steps.
- A real log slice: avoid screenshot-only handoffs. Keep 20-50 lines around the error so the next operator can grep and compare.
- What has already been ruled out: list restarts, model swaps, channel swaps, and permission checks that were already tried.
That turns a broad troubleshooting page into a handoff-ready incident entry point and reduces bouncing between overlapping issue pages.
Route from the checklist to the right incident page first
This checklist is useful for narrowing the scope, but high-intent search visitors should not stay on a generic page longer than necessary. Use this three-step routing rule before continuing:
- Identify the failure layer first: decide whether the agent failed to start, a channel failed to initialize, a tool call failed, or a provider / runtime was selected incorrectly.
- Check whether a dedicated incident page already exists: if the symptom points to Codex OAuth, custom provider keys, session resets, Feishu schema compatibility, or Telegram initialization, open that incident page before staying in the general checklist.
- Return here only to complete evidence: use the checklist again when the symptom spans multiple layers or you still need logs, config diffs, and reproduction steps.
A simple rule works well: if the user can name the exact error or channel, route to the dedicated page; if they can only say “the agent does not work,” keep using this checklist.
If you arrived from setup, first decide whether this is post-install verification
The latest traffic pattern shows the setup page and this troubleshooting checklist appearing together, which usually means some readers are not debugging a single bug yet. They are trying to answer: “is this OpenClaw install actually usable now?” For that path, do not start by changing agent config. Use this order instead:
- Run the smallest usability check first: confirm the CLI starts, the Gateway reports healthy, and the default model can complete one short reply.
- Locate the step that failed: if the setup command itself has not succeeded, return to the installation guide; if install succeeded but the first real task fails, use this checklist.
- Move to a dedicated incident page last: only jump to a provider, channel, schema, or session bug page when the error message already points there.
That split keeps “new install, not sure if it works” readers separate from “production task hit a known failure mode” readers, so setup traffic does not get lost in an overly broad troubleshooting checklist.
Decide in five minutes: stay on the checklist or jump to a dedicated incident page
If you arrived from /, setup, or search, do not run every checklist item from the top. Spend five minutes classifying the entry point first:
- There is an exact error string: search this page and the news directory for the dedicated incident page, such as
No API key,/reset,Telegram initialize, orschema. - There is a specific channel: triage by channel first, because Feishu, Telegram, Discord, Kimi, and ACP usually fail at different log boundaries.
- The report only says “the agent does not work”: stay on this checklist, identify the failing layer, then capture the smallest reproduction.
- This started right after install or migration: return to the install or migration guide first and check for stale commands, paths, remotes, or provider names.
The goal is not to make readers consume one more article. It is to route high-intent traffic to the page that can resolve the issue fastest: if they can name the error and channel, send them to the dedicated page; if they cannot, keep narrowing scope here.
Turn checklist traffic into the next measurable action
When GA4 sends readers to this general troubleshooting page, do not leave them with a generic checklist. Convert the visit into one of three measurable next actions:
- Incident handoff: if the reader already has a failing command, capture the command, channel, model/provider, and first error line before opening a dedicated bug page.
- Setup validation: if the reader came from
/setup, ask whether the first agent task, browser action, and channel delivery have each passed once. If not, send them back to installation validation instead of deeper debugging. - Release regression check: if the failure appeared after an update, compare the symptom with the release note and the CLI performance regression path before changing model or provider settings.
This makes the page work as a routing hub: each high-intent visit should either collect evidence, validate setup, or move to a dedicated incident article.
Before handoff, fill this four-layer evidence table
If you land here during an incident, the most expensive mistake is restarting the agent before preserving evidence. Capture the failure in four layers first:
| Layer | Capture first | Escalate when |
|---|---|---|
| Runtime | agent name, model, provider, latest config change | the same task fails after switching model/runtime |
| Tooling | failing tool name, input summary, error timestamp | multiple tools show the same permission pattern |
| Channel | Discord, Telegram, Feishu, browser, or CLI path | the same agent fails across multiple channels |
| Permission | token, OAuth, file path, external API write scope | reads work but writes fail |
This turns “my agent is broken” into transferable evidence and helps high-intent readers choose the next route: the CLI regression page, the Skills decision page, or this broader Agents checklist.
When should you stop self-triage and escalate?
The point of a general troubleshooting checklist is not endless self-checking. Escalate to the maintainer or internal platform owner when any of these signals appears:
- Cross-channel reproduction: the same agent fails through CLI, Feishu, Discord, or the browser entrypoint.
- Cross-model reproduction: switching provider / model still fails at the same layer, so it no longer looks like one model issue.
- Reads work but writes fail: permissions, tokens, paths, or external API write boundaries may be broken.
- The issue started after an upgrade: the latest version, gateway, plugin, or auth config change lines up with the failure window.
- Delivery is blocked: automation, customer work, or production on-call is now blocked, so personal trial-and-error is too slow.
Escalate with the four-layer evidence table, the last reproducible input, and the actions already ruled out. That gives maintainers an incident packet instead of a vague “the agent is broken” report.
Route vague agent failures to a specific incident page
Use this page as the triage hub, not the final destination. After the first five minutes, route the visitor to the most specific incident class so the next click matches the failure they actually have.
- Session or history looks empty: send them to the ACP transcript /
sessions_historyarticle before changing models. - The chat pane shows a warning triangle: send them to the main-session warning-triangle recovery guide before clearing all session storage.
- A reusable automation times out: send them to the
sessions_sendtimeout recovery guide and compare old-session versus fresh-session behavior. - A provider returns a blank assistant message: send them to the Vertex Gemini blank-reply triage path and preserve raw provider response fields.
This turns broad “OpenClaw agent not working” traffic into sharper troubleshooting intent. The reader either leaves with a scoped recovery path or brings back a better evidence packet for maintainers.
Related reading
- OpenClaw Complete Installation Guide
- OpenClaw Quick Install Guide for Mac
- How to verify an OpenClaw installation after setup
- Safety and cost
Need More Help?
Join our community Discord or check the official documentation for more detailed guides.
Happy coding!
