Back to News
Troubleshooting
OpenClaw Agent Sessions Can Get Stuck After Idle Timeout: Why `sessions_send` Keeps Timing Out Until a Gateway Restart

OpenClaw Agent Sessions Can Get Stuck After Idle Timeout: Why `sessions_send` Keeps Timing Out Until a Gateway Restart

OpenClaw News Editorial

OpenClaw News Editorial

A newly reported OpenClaw bug describes a painful automation failure mode: once an agent session goes idle, hits timeout, and becomes ended, later calls to sessions_send may keep returning timeout instead of waking the session back up.

Source issue: openclaw/openclaw#60948

For operators, the important detail is not the idle timeout by itself. The real problem is that the system may stop behaving like a recoverable message-driven session and start behaving like a dead endpoint that only comes back after a gateway restart.

What is being reported

According to the issue report, the sequence looks like this:

  1. an agent session runs normally
  2. it stays idle long enough to time out
  3. the session enters an ended state
  4. a later sessions_send(sessionId, message) call returns timeout
  5. automation does not recover until the gateway is restarted manually

That matters because many operators treat sessions_send as a safe way to re-engage a known session for scheduled work, checkpoints, or downstream delegation.

Why this hurts real automation

This is more than a cosmetic timeout bug.

The failure cuts directly into common production-style patterns such as:

  • scheduled checkpoint messages
  • recurring dispatch into existing agent sessions
  • follow-up steps triggered by cron or orchestration
  • background workflows that assume a session can be reused after inactivity

If an ended session cannot be reactivated cleanly, teams usually get pulled into one of three bad outcomes:

  • silent automation drift — jobs stop progressing, but nothing clearly says the workflow model changed
  • manual babysitting — someone has to restart the gateway just to restore expected behavior
  • unsafe retries — operators resend work without knowing whether prior state is partially preserved or completely dead

In other words, the real damage is not “one timeout.” It is loss of trust in session reuse as an automation primitive.

How this differs from a normal timeout

A normal timeout is a policy event.

You may not like it, but the system is still behaving coherently: the session becomes inactive because its lifecycle rules say it should.

This reported bug is different. The complaint is that after timeout has already happened, the recovery path is broken or missing.

That puts the issue in the boundary between:

  • session lifecycle state
  • message-triggered reactivation
  • runtime/session reattachment logic

If sessions_send is meant to act like a valid re-entry path, then a persistent timeout after the session has ended suggests the system no longer knows how to turn stored session identity back into an active processing path.

The most likely root-cause boundary

Based on the report alone, the problem does not primarily look like “the original session timed out too aggressively.”

It looks more like one of these deeper failures:

  • the session record remains addressable, but not actually resumable
  • the ended state is terminal in practice even when the operator expects message-driven reactivation
  • sessions_send targets a session handle that no longer maps cleanly to a live worker/runtime path
  • restart rebuilds enough internal routing or state attachment to make sends work again

That last point is especially telling. When a gateway restart restores behavior, the bug often lives in recovery wiring or runtime attachment, not in the user’s prompt or schedule itself.

How to confirm you hit this exact failure mode

You are probably looking at the same bug if all of these are true:

  1. the session previously accepted sessions_send normally
  2. the session later idled into an ended state
  3. the next send attempts fail with timeout
  4. creating unrelated activity does not restore the old session path
  5. restarting the gateway makes the workflow usable again

That combination separates this issue from simpler cases like:

  • wrong session id
  • downstream tool timeout inside the agent turn
  • cron misfire
  • network loss to a remote dependency
  • a run that was expected to be one-shot rather than reusable

Practical mitigations before upstream fixes it

Until upstream behavior is clarified or fixed, operators should treat ended agent sessions as potentially non-recoverable in long-running automation.

The safer operating pattern is:

1) Do not assume sessions_send can always revive an idle-ended session

If the workflow is business-critical, explicitly design for the possibility that an ended session must be replaced rather than resumed.

2) Add session health checks before enqueueing important follow-up work

Before sending the next scheduled message, verify whether the target session is still expected to accept new work.

3) Prefer recreatable workflows over stateful reuse when the task is high-value

If losing a session would break revenue, publishing, moderation, or deployment chains, prefer a design that can start a fresh session cleanly instead of depending on indefinite reuse.

4) Log the boundary between “timed out” and “manually restored”

If restart is your current workaround, record when the session entered the failed state and what was replayed after recovery. That reduces duplicate side effects and makes debugging easier later.

What operators should watch for next

When upstream clarifies the intended behavior, the key question is not just whether the bug gets patched.

The deeper question is: what is the contract of sessions_send against ended sessions?

Operators need a reliable answer to one of these models:

  • ended sessions should be auto-reactivated
  • ended sessions should fail fast with a clear terminal error
  • ended sessions should transparently create a replacement session

Any of those can be workable if clearly documented. The real operational problem is the current ambiguous middle: the session looks targetable, but behaves like a dead route until a full gateway restart.

Bottom line

This report is worth watching because it hits a high-intent operational use case: scheduled and reusable agent automation.

If your setup depends on sessions_send to keep existing agent sessions alive across idle gaps, do not assume timeout is a harmless state transition. In the currently reported behavior, it can become a persistent recovery failure that breaks automation until the gateway is restarted.

If you have confirmed ended-session recovery is the failure boundary

If the problem is no longer a single timeout and the whole session reuse path is unreliable after the session ends, add two routing checks before you restart everything:

  1. Start with the OpenClaw Agents troubleshooting guide: what to check first when tasks stall, tools stop responding, or results look wrong, and confirm whether provider errors, tool execution, or message delivery are also involved.
  2. If the first failure actually followed a crash or restart rather than an idle timeout, compare it with Gateway does not resume orphaned sessions after crash/restart so you do not mix session reuse semantics with restart recovery.

That distinction matters because operators often collapse “send fails after ended” and “restart did not recover orphaned work” into one incident. Separating them keeps the investigation short and routes high-intent readers to the right fix path.

If GA4 brought you from setup or troubleshooting, separate idle recovery from launch failure

This article now sits in the same traffic cluster as the setup guide, the troubleshooting hub, and ACP transcript-history incidents. That means many readers are probably trying to decide whether a stuck sessions_send is a bad agent launch, a stale API view, or the idle-session recovery bug described here.

Use this quick split before restarting the Gateway:

  1. The session never accepted a first message: go back to setup validation, authentication, sandbox, and runtime startup logs. This page is probably not the primary fix yet.
  2. The run produced transcript evidence but history looks empty: compare with the ACP transcript-history article first, because the work may exist even if the session API projection is stale.
  3. The session worked before idle time, then every follow-up sessions_send times out: treat it as the ended-session recovery boundary, preserve timestamps, and only then plan a controlled Gateway restart.

That routing keeps this page aligned with high-intent operators who arrive from broader troubleshooting paths and need to choose the least disruptive recovery step.

Recovery checklist before you restart the Gateway

Use this checklist when the incident is live and you need to decide whether to resume, replace, or restart. It turns the bug report into an operator-facing recovery path:

  1. Capture the last successful sessions_send timestamp and the first timeout timestamp.
  2. Confirm whether the target session is merely idle, explicitly ended, or missing from the session index.
  3. Send one low-risk diagnostic message only; do not replay the business task yet.
  4. If the diagnostic send also times out, create a replacement session for new work and mark the old session as unsafe to reuse.
  5. Restart the Gateway only after preserving the failed session id, logs, and queued follow-up message, so the upstream report can distinguish routing repair from data loss.

The key SEO and operations point is that the workaround should not be “restart first.” Restarting too early destroys the evidence that proves whether the real bug is ended-session reactivation, orphaned runtime attachment, or a caller-side retry loop.

Related reading

Search-entry triage: what to check first when sessions_send times out

If you landed here from searches like “sessions_send timeout”, “agent session timeout”, or “OpenClaw subagent not returning”, do not start by blaming a slow model. Split the failure in this order:

  1. Check whether the target session is still alive: if the target session exited or was cleaned up, more retries only create more waiting.
  2. Check the tool boundary: separate a send failure, a long-running child task, and a result that was produced but lost in routing.
  3. Check whether the timeout matches the task shape: long work should become a background task or an explicit waiting point, not one synchronous call that waits forever.
  4. Preserve the target session id and timestamp: for log review, the session id, call time, and timeout value are more useful than “it got stuck”.

The fastest way to debug this class of incident is to split “the model did not answer” into four checks: session liveness, tool delivery, task execution, and result return.

Related reading

Alerting pattern for reusable session automation

For high-intent automation, the useful alert is not just “a send timed out.” Track the sequence that proves the reusable session path is unsafe:

  1. the same target session accepted a previous sessions_send
  2. the session later became idle or ended
  3. a diagnostic follow-up send timed out before business work was replayed
  4. creating a fresh session works while the old target still fails

That four-part pattern separates an ended-session recovery bug from a slow model, a long tool call, or a caller-side timeout that simply needs a larger deadline.

If this pattern appears in scheduled publishing, moderation, inbox triage, or deployment workflows, route new work to a replacement session first. Treat the old session as evidence to preserve, not as the safest place to keep retrying.

Add a replacement-session route to orchestration

If this timeout pattern shows up in scheduled jobs or background orchestration, do not alert only on “old session failed.” Add a replacement-session route at the orchestration layer: when the old session fails two diagnostic sends in a row, move new business work to a fresh session and keep the old session only for evidence capture and upstream reproduction.

That route is especially useful for publishing, moderation, and data-sync workflows that cannot stay blocked for long. It turns the search intent behind “sessions_send timeout fix” into a concrete operator move: isolate the old session, continue work in a new one, and then decide whether a controlled Gateway restart is still required.

FAQ: sessions_send timeout after an idle session

Should I keep retrying the same ended session?

Not as the first recovery move. If the same ended session times out twice on low-risk diagnostic sends, route new business work to a replacement session and preserve the old session for logs. Replaying production work into the old target can create duplicate side effects without proving that recovery works.

When is a Gateway restart justified?

Use restart as a controlled recovery step after evidence capture, not as the first debug step. Restart becomes more justified when a fresh session cannot accept work either, or when logs show the Gateway cannot attach message delivery to any live runtime path.

What should I record for an upstream bug report?

Keep the target session id, the last successful send time, the first timeout time, whether the session was idle or explicitly ended, and whether a fresh replacement session worked. Those details separate lifecycle semantics from routing failure.

Recovery playbook for cron and background jobs

If the timeout appears inside a recurring job, treat it as a workflow-continuity problem rather than a chat UX problem. A practical runbook is:

  1. mark the old session as unsafe after two diagnostic sessions_send timeouts;
  2. create or select a replacement session before replaying any write-capable task;
  3. copy only the minimal context needed to continue the job, not the whole failed transcript;
  4. record whether the replacement session completed the same business step;
  5. inspect Gateway logs only after the new path has protected the user-facing workflow.

This sequence is useful for teams searching for “OpenClaw cron sessions_send timeout” because it keeps publishing, deployment, and triage jobs moving while preserving enough evidence for a real upstream bug report.

Source

© 2025 OpenClawNews.org
All rights reserved.
This is an independent news site. Not affiliated with, endorsed by, or connected to OpenClaw. OpenClaw is a trademark of its respective owner.
Join the waitlist:

OC NEWS