Back to News
Tutorial
OpenClaw Compaction Stuck? How to Recover a Frozen Session and Reduce Repeats

OpenClaw Compaction Stuck? How to Recover a Frozen Session and Reduce Repeats

OpenClaw News 编辑部

OpenClaw News 编辑部

If your OpenClaw session suddenly stops responding after a long conversation, and the logs keep repeating:

cancelling compaction with no real conversation messages to summarize

this is the practical guide you want.

This is not written as a release-note recap. It is a field troubleshooting memo turned into a user-facing recovery guide: how to identify the failure mode, how to restore service fast, and what config changes are worth trying if it comes back.

TL;DR

If all of these line up:

  • a session freezes after long context growth,
  • logs show compaction / compact / summarize,
  • deleting lock files does not help,
  • restarting gateway restores the session path,

then you can usually treat it as a compaction-stuck incident.

Fastest recovery order:

  1. Confirm the logs point to compaction
  2. Restart the gateway first
  3. Do not treat lock-file deletion as the primary fix
  4. If it repeats within 24-48 hours, change compaction settings instead of repeating the same manual cleanup

Important: this article separates observed behavior, operational inference, and recommended configuration. Not every inference below is presented as a source-proven engine internals claim.

When this guide fits

This guide fits when:

  • one or a few sessions freeze after long conversations,
  • gateway is still alive,
  • logs clearly mention compaction-related keywords,
  • the session path recovers after a gateway restart.

This guide does not fit when:

  • gateway itself fails to start,
  • all providers are timing out globally,
  • all sessions are broken with no compaction signals,
  • openclaw.json is simply invalid or permission-blocked.

What the symptom usually looks like

Common pattern:

  • a previously active session stops replying,
  • logs start showing compaction-related lines,
  • one especially important line may appear repeatedly:
Compaction safeguard: cancelling compaction with no real conversation messages to summarize.
  • removing a lock file does not revive the session,
  • restarting the gateway clears the stuck state.

Observed behavior

From an ops perspective, the strongest signal is not any single log line by itself. It is the combination of:

  • compaction-related logs,
  • session no longer responding,
  • gateway restart restoring service.

Operational inference

When those signals appear together, the problem often looks less like “a lock file is blocking everything” and more like “the session entered a compaction-related state that did not unwind cleanly”.

That is an operational inference, not a claim that the exact internal root cause has already been fully proven from source code.

Why restarting gateway is usually the first move

Many users try lock-file cleanup first. That instinct makes sense, but in this failure pattern it often wastes time.

More practical recovery order:

openclaw gateway restart

Or, if your deployment uses systemd:

sudo systemctl restart openclaw-gateway

Why this often works better

In this class of incidents, the lock file is often only the visible artifact. The more important stuck state appears to live in the running gateway/session path.

That explains the common pattern:

  • the lock looks suspicious,
  • deleting it changes nothing,
  • restarting gateway restores normal behavior.

So the safer phrasing is:

Do not treat lock-file deletion as the primary repair path.

Quick checks to run

1) Check overall status

openclaw status

Look for:

  • gateway service still running,
  • abnormal session behavior despite gateway being alive,
  • tight contextTokens settings,
  • whether the problem is isolated to one session or broader.

2) Search compaction-related logs

If gateway is managed by systemd:

journalctl -u openclaw-gateway -n 300 --no-pager | grep -iE 'compaction|compact|summar|timeout|stuck'

If you are unsure about the service name:

systemctl list-units | grep -i openclaw

3) Search the key failure lines directly

journalctl -u openclaw-gateway -n 500 --no-pager | grep -F 'cancelling compaction with no real conversation messages to summarize'
journalctl -u openclaw-gateway -n 500 --no-pager | grep -i 'compaction'
journalctl -u openclaw-gateway -n 500 --no-pager | grep -iE 'timeout|timed out'

When to move from “recover” to “repair”

If it happens only once, restart-and-watch may be enough.

But if any of these are true, stop treating it as a one-off:

  1. it repeats within a day or two,
  2. the same logs show up again soon after recovery,
  3. multiple sessions are affected,
  4. compaction triggers too often near the context edge,
  5. the exact line changes, but the session still times out and stalls.

Which settings are worth checking

Focus on these:

agents.defaults.contextTokens
agents.defaults.compaction.mode
agents.defaults.compaction.maxHistoryShare
agents.defaults.compaction.reserveTokensFloor
agents.defaults.compaction.memoryFlush.enabled
agents.defaults.contextPruning

Typical risk pattern

Higher-risk combinations often include:

  • compaction.mode = "safeguard",
  • context headroom that is too tight,
  • history share set too aggressively,
  • repeated automatic compaction near the boundary,
  • sessions that seem unable to recover cleanly once compaction begins.

A safer baseline to try after recurrence

If you have already seen this failure mode for real, this is a pragmatic baseline to test:

"agents": {
  "defaults": {
    "contextTokens": 100000,
    "compaction": {
      "mode": "default",
      "maxHistoryShare": 0.25,
      "reserveTokensFloor": 30000,
      "memoryFlush": {
        "enabled": false,
        "softThresholdTokens": 4000
      }
    }
  }
}

Why these changes are worth trying

safeguard -> default

If you are already seeing:

  • empty-summary style behavior,
  • no real conversation messages,
  • compaction loops or timeouts,

then moving off the safeguard path is the first meaningful config change to test.

contextTokens -> 100000

The goal is not “bigger is always better”. The goal is:

  • less edge-triggered compaction,
  • more working headroom,
  • fewer forced compactions right at the boundary.

reserveTokensFloor -> 30000

This gives more stable room for:

  • system prompts,
  • tool calls,
  • final response generation.

That extra buffer matters when a session is already near the edge.

maxHistoryShare -> 0.25

This is not about keeping everything forever. It is about keeping compressed sessions coherent enough that they still behave like an ongoing conversation.

Standard config-change flow

Back up first:

cp -p ~/.openclaw/openclaw.json ~/.openclaw/openclaw.json.bak.$(date -u +%Y%m%dT%H%M%SZ)

Validate JSON after editing:

python3 -m json.tool ~/.openclaw/openclaw.json >/dev/null

Then restart gateway:

openclaw gateway restart

Then watch for 24-48 hours:

  • Does compaction still fire too often?
  • Can sessions keep responding after compaction?
  • Does the same no real conversation messages line come back?

What not to do

1) Do not call lock-file deletion the fix

It may help in some environments, but in this failure pattern it is often not the real repair.

2) Do not spam new test messages into the stuck session

That usually makes the state noisier, not clearer.

3) Do not change a pile of unrelated settings at once

If you do, you will not know which change actually mattered.

4) Do not make “upgrade immediately” your first reflex

Upgrade is sensible when you are validating an upstream fix. Otherwise, a small recovery-first config repair is usually faster.

When to escalate to upstream or deeper source debugging

Escalate if any of these become true:

  1. it still repeats after switching to default,
  2. it still repeats after raising contextTokens and reserveTokensFloor,
  3. multiple sessions are hit in a short window,
  4. restart only helps briefly,
  5. the exact log line changes but the same timeout behavior remains.

Keep a short incident record

Every incident should leave behind a compact record like this:

## Compaction incident
- Time:
- Session:
- Symptom:
- Key log:
- Recovery:
- Config at time:
- Reproduced again?: yes/no

That record will matter later when deciding whether you should:

  • keep tuning locally,
  • open an upstream issue,
  • or test a newer build.

Common recovery questions

When should you treat this as a compaction problem instead of a generic timeout

If these four signals line up, compaction should move to the top of your troubleshooting tree:

  • the session froze after a long conversation, not immediately on startup
  • logs contain compaction, compact, or summarize
  • deleting lock files did not recover the session
  • restarting gateway restored the session path

If those signals are missing, for example the whole provider stack is timing out, gateway itself will not start, or all sessions fail without compaction logs, do not spend your first hour on this playbook.

When should you move from restart-first recovery to config repair

A practical threshold is simple:

  • it repeats again within 24 to 48 hours
  • the same class of logs returns quickly
  • multiple sessions get hit
  • restart only gives you a short-lived recovery

At that point, stop treating gateway restart as the full answer and move straight into compaction.mode, contextTokens, and reserveTokensFloor tuning.

Copy this handoff note before escalating to maintainers

If restart and config tightening only recover the session briefly, stop describing it as “the session got stuck again.” Compress the scene into this handoff so maintainers can tell whether the root is compaction, provider timeout, or gateway recovery:

Compaction stuck handoff
- OpenClaw version:
- Install method:
- Session age / approximate turns:
- Last user-visible symptom:
- Exact compaction-related log lines:
- Config: compaction.mode / contextTokens / reserveTokensFloor:
- Tried: gateway restart / lock cleanup / config change:
- Result after second restart:
- Reproduces in a new session?: yes/no

This handoff turns “a long conversation froze” into three answerable questions: whether the compaction strategy is wrong, whether the context budget is too tight, and whether recovery only masks the failure briefly. For high-intent visitors from search, that is more useful than another blind restart.

On the second recurrence, use this 15-minute prevention card

If the same team hits compaction stuck for a second time, stop treating it as a one-off incident. Spend 15 minutes turning recovery into recurrence prevention:

  1. Freeze the scene: keep the current logs, session age, last three turns before failure, and config at the time.
  2. Change only one knob: adjust either compaction.mode or the context budget first, not provider, model, and gateway settings all together.
  3. Replay the same long-session path: continue the task that froze for 10 minutes, so the result is not just a clean new session looking healthy.
  4. Update the default template: write the working config into the team install template or runbook, so the next machine does not inherit the old threshold.
  5. Schedule a review: check for returning compaction logs after 24 hours instead of closing the incident on same-day recovery.

This card is for high-intent readers who already saw recurrence: the goal is not to rescue the session one more time, but to turn a temporary fix into a repeatable stable configuration.

If you came from search after a second freeze, choose the next branch quickly

Search visitors often land here after they already restarted once and the same long session froze again. Do not repeat the same recovery step without changing the decision tree.

Use this split before touching unrelated provider or model settings:

  1. Restart fixed it once, then it returned: move from emergency recovery to config prevention, especially compaction.mode, contextTokens, and reserveTokensFloor.
  2. Restart did not help at all: stop treating this as normal compaction drift and capture gateway logs, provider timeout evidence, and whether new sessions also fail.
  3. Only one old session fails: preserve the transcript and session age, then compare it with a fresh session so the repair targets long-context pressure rather than the whole install.
  4. Multiple sessions fail together: escalate beyond this playbook and check provider health, gateway process health, and recent deployment changes first.

This keeps high-intent readers from looping on gateway restart when the real goal is either recurrence prevention or upstream incident isolation.

Related reading

Final one-card emergency summary

If you only remember one version, remember this one:

  1. See a compaction-stuck symptom
  2. Confirm it in logs
  3. Restart gateway
  4. Do not over-trust lock-file cleanup
  5. If it comes back, switch compaction.mode to default, raise contextTokens, and add more reserve via reserveTokensFloor

That is not the whole story, but it is the fastest practical playbook for most users who just need the session path working again.

© 2025 OpenClawNews.org
All rights reserved.
This is an independent news site. Not affiliated with, endorsed by, or connected to OpenClaw. OpenClaw is a trademark of its respective owner.
Join the waitlist:

OC NEWS