Back to News
Troubleshooting
OpenClaw v2026.4.8 Can Peg a CPU Core in Headless Docker: What the Bonjour Sidecar Bug Breaks and How to Work Around It

OpenClaw v2026.4.8 Can Peg a CPU Core in Headless Docker: What the Bonjour Sidecar Bug Breaks and How to Work Around It

OpenClaw News Editorial Desk

OpenClaw News Editorial Desk

A newly opened bug report says OpenClaw v2026.4.8 can enter a runaway bonjour/mDNS advertise loop inside a normal Docker bridge network, driving one child process to 99% CPU and, after several minutes, causing the container to exit.

Source issue: openclaw/openclaw#63248

For operators who run OpenClaw on VPS hosts, CI boxes, or disposable test containers, this is worth paying attention to. The reported failure is not a vague "Docker is unstable" complaint. It points to a specific pattern:

  • OpenClaw starts normally
  • the gateway becomes ready
  • a child openclaw process keeps one CPU core busy
  • logs show bonjour watchdog re-advertise attempts repeating
  • the container may exit later without a clean explanation

What appears to be happening

According to the report, the problem shows up when OpenClaw runs in a headless container without --network=host.

That matters because bonjour/mDNS service discovery is built for local network announcement. In a typical Docker bridge network, that announcement path may be ineffective or partially broken from the point of view of the container. The bug report suggests OpenClaw keeps trying to re-advertise the service instead of backing off cleanly.

The log signature cited in the issue looks like this:

[bonjour] watchdog detected non-announced service; attempting re-advertise
[bonjour] restarting advertiser (service stuck in announcing ...)

If that loop never settles, the cost is not just noisy logs. The container can lose real capacity to do useful work.

Why this hurts in production

1) It wastes the very CPU budget small VPS deployments depend on

Many OpenClaw installs run on modest hosts. Burning one full core on failed service advertisement means less room for:

  • agent turns
  • CLI invocations
  • browser or tool work
  • normal background jobs

2) It creates a misleading symptom chain

The first symptom you notice may not be "bonjour is broken." It may be:

  • the container exits after 9 to 18 minutes
  • docker exec commands feel hung
  • helper CLI commands never return
  • the gateway looks alive at first, then degrades

That makes the issue easy to misdiagnose as a generic container or resource problem.

3) It is especially painful in CI and test environments

Disposable containers usually run on bridge networking by default. If the report holds broadly, a clean, isolated Docker test environment can become the exact place where this bug shows up fastest.

Reported side effect: CLI starvation

The same issue report claims that once the bonjour loop pegs a core, follow-up commands such as:

  • openclaw cron list
  • openclaw message --help
  • openclaw channels list

can enter their own busy-wait behavior and fail to return useful output.

That may turn out to be a separate bug, but from an operator perspective the distinction does not matter much. Once the runtime is CPU-starved, diagnosis becomes harder and the system looks less trustworthy.

Fast decision tree for operators

If you need a fast go or no-go call, use this split:

  • CPU spikes begin right after advertise startup and logs keep repeating bonjour watchdog lines
    • treat Bonjour as the primary fault boundary and test OPENCLAW_DISABLE_BONJOUR=1 first.
  • CPU is high only when document throughput is also clearly high
    • do not assume the sidecar is the root cause yet; compare queue depth, job volume, and worker activity first.
  • You rely on LAN discovery for nearby devices
    • isolate in staging before disabling Bonjour globally, because the workaround may remove a feature you actually need.
  • You run on a VPS, CI box, or behind a reverse proxy
    • prioritize stability over discovery and treat disabling Bonjour as the fastest low-risk mitigation.

What you can do right now

The workaround reported in the issue is to disable bonjour inside the container:

docker run -d \
  -e OPENCLAW_DISABLE_BONJOUR=1 \
  -p 28789:18789 \
  -v <state>:/home/node/.openclaw \
  ghcr.io/openclaw/openclaw:2026.4.8

The reporter says this stops both the CPU runaway process and the bonjour log spam.

If you operate OpenClaw in headless Docker, a practical short checklist is:

  1. Check docker top for an unexpected child openclaw process sitting near 99% CPU.
  2. Check container logs for repeated bonjour watchdog or re-advertise messages.
  3. If you do not need LAN service discovery, test with OPENCLAW_DISABLE_BONJOUR=1.
  4. Re-run the same container workload and compare CPU behavior and uptime.

Preserve five evidence points before disabling bonjour

This failure can easily get reduced to “CPU is at 99%,” which is not enough for a good fix, rollback, or upstream issue. Before changing the environment, preserve five evidence points:

  1. Container environment: whether this is Docker, headless, a VPS, or a CI runner, and whether LAN discovery is needed at all.
  2. CPU timeline: record how long after startup CPU spikes, and whether it happens immediately or after an announce/retry loop begins.
  3. Bonjour log sample: keep repeated advertise, mDNS, socket, and network-interface logs instead of only a top screenshot.
  4. Before/after comparison: with the same image and entry command, compare CPU and CLI responsiveness with and without OPENCLAW_DISABLE_BONJOUR=1.
  5. Blast radius: note whether agent runs, helper commands, and gateway health checks are slowed or still normal.

This turns “the host is stuck” into a reproducible headless-discovery failure, instead of pushing readers toward blind upgrades, reinstalls, or overly broad rollbacks.

What to verify before calling the incident closed

After the CPU spike disappears, verify the parts that actually protect user-facing reliability:

  1. One real document or headless job finishes normally.
  2. One representative agent workflow completes without hanging helper commands.
  3. Container uptime survives past the earlier crash window.
  4. Logs stop showing repeated bonjour re-advertise churn.

That matters because lower CPU alone is not enough. The real success condition is that the high-intent workload path is healthy again.

If you arrived from safety/cost, release evaluation, or CLI latency

This Bonjour CPU page should not serve only readers who already found bonjour watchdog logs. GA4 shows /zh/safety-and-cost, /zh/news/openclaw-2026-3-13, and the CLI regression page appearing in the same daily cluster, which means readers are likely deciding whether they are seeing cost runaway, release risk, or a real command-path slowdown.

Use this split before applying the workaround:

  1. Start with safety and cost: if the primary issue is API spend, exposed port 8080, credential rotation, or budget guardrails, go back to Safety & Cost Control instead of blaming every resource symptom on Bonjour.
  2. Then check the release note: if you are evaluating whether OpenClaw can become a team baseline, return to the 2026.3.13 release note and treat Bonjour as a deployment-shape risk to validate separately.
  3. Finally compare the CLI regression path: if commands hang for 20-40 seconds after hook loading, rather than after mDNS advertisement pegs a core at container startup, read the CLI performance regression triage.

That keeps headless Docker operators from collapsing four different intents into one fix: CPU burn, cost control, release readiness, and command latency are different paths. Put OPENCLAW_DISABLE_BONJOUR=1 first only when the log pattern and timeline match Bonjour re-advertise churn.

Turn mDNS CPU traffic into a headless-deployment guardrail checklist

If you arrived because the Bonjour or mDNS sidecar pegs 99% CPU in a headless environment, do not only restart the process. Split the issue into three evidence layers: whether service discovery is spinning on missing network interfaces, whether the advertise sidecar registers the same service repeatedly, and whether health checks misclassify high CPU as recoverable.

The minimal guardrail checklist includes the headless host network interfaces, the mDNS advertise toggle state, process-level CPU curves, and a comparison load after disabling the sidecar. That routes mDNS incident traffic toward headless deployment, resource protection, and automatic fallback intent instead of leaving it at a single CPU screenshot.

FAQ: Bonjour / mDNS CPU spikes in headless OpenClaw deployments

Should every VPS deployment disable Bonjour immediately?

Not automatically. Disable it first when the host is headless, LAN discovery is not part of the user path, and CPU graphs or logs show mDNS advertise churn. If local device discovery matters, collect evidence and test a scoped canary before making it a default.

How do I prove this is not just a slow model or CLI regression?

Compare timing. Bonjour incidents usually start around service advertisement or container startup and show process-level CPU burn. Model or CLI regressions usually correlate with provider calls, hook loading, or command execution. Keep those timelines separate before applying the workaround.

What is the safest temporary mitigation?

Set the Bonjour disable flag for the affected headless deployment, restart during a controlled window, and compare CPU load plus command latency before and after. Keep the original logs so the workaround does not erase the evidence needed for an upstream fix.

Related reading

When disabling bonjour is a reasonable tradeoff

For many server-side deployments, bonjour is not mission-critical.

If your OpenClaw instance is:

  • running on a remote VPS
  • accessed through a domain, reverse proxy, tunnel, or fixed port
  • used for automation rather than local LAN discovery

then disabling bonjour may be the right short-term move until upstream behavior is hardened.

In other words, this looks like one of those cases where desktop-friendly discovery defaults can become a liability in headless infrastructure.

What maintainers may need to change

The issue proposes several sensible directions:

  • document OPENCLAW_DISABLE_BONJOUR=1
  • auto-disable bonjour in headless/containerized environments
  • add exponential backoff instead of a hot retry loop
  • stop treating repeated announce failure like something that should run forever at full speed

That last point matters most. Even if announcement fails, the failure mode should be bounded.

Bottom line

If your OpenClaw v2026.4.8 container seems healthy at boot but later burns CPU, hangs on helper commands, or exits unexpectedly, bonjour re-advertise churn is now a credible root cause to check first.

For operators, the immediate move is simple:

  • confirm whether the log pattern matches
  • test OPENCLAW_DISABLE_BONJOUR=1
  • avoid assuming the whole stack is broken when the fault may be in service advertisement behavior inside Docker

Source

Quick answer

If OpenClaw headless-doc pegs CPU after Bonjour or mDNS advertise starts, treat it as a sidecar runaway condition, not normal background warm-up. First confirm whether CPU spikes begin immediately after service advertisement, then isolate the sidecar path before blaming the whole host.

  • •What to assume first: If the CPU jump correlates with advertise startup and headless-doc stays hot while the rest of the stack remains mostly responsive, the probable fault boundary is the sidecar advertise loop, not generic site traffic.
  • •Fastest low-risk triage: Check whether disabling or bypassing Bonjour advertisement drops CPU immediately, then compare behavior across headless and non-headless runs before changing unrelated worker settings.
  • •When it becomes an incident: Escalate it as an operational incident once the sidecar pins CPU long enough to delay document jobs, starve other services, or keep recurring after restarts on the same build.

Frequently asked questions

If CPU drops when Bonjour advertisement is disabled, is the root cause confirmed?

It is strong boundary evidence, but not absolute proof of the exact line of failure. It confirms the advertise path is the best isolation point and that operators should inspect service-discovery sidecar behavior before chasing generic host tuning or unrelated document logic.

How can operators quickly tell this is a sidecar runaway, not simply heavy document workload?

Look for a timing split. In a runaway sidecar case, CPU spikes start with advertise initialization and stay hot even without proportionate document throughput. A real workload spike should correlate with visible job volume, queue growth, or active document processing.

What is the safest immediate mitigation while waiting for an upstream fix?

Reduce blast radius first: disable the advertise path if possible, restart only the affected component, and verify document jobs recover before making broader platform changes. The goal is to stop wasteful CPU burn without masking the trigger you still need to document.

When should operators keep Bonjour enabled instead of disabling it immediately?

Keep it enabled only if your deployment truly depends on LAN discovery, such as a local-network setup where nearby devices need to find the service automatically. For headless VPS, CI, or reverse-proxy deployments, Bonjour usually is not part of the user acquisition path, so disabling it is the lower-risk default while you isolate the bug.

What should operators verify after applying OPENCLAW_DISABLE_BONJOUR=1?

Do not stop at lower CPU usage. Verify one real document job, one agent turn, and container uptime over the same window where crashes previously happened. That confirms you removed the high-cost failure mode without introducing a quieter regression in the actual production path.

© 2025 OpenClawNews.org
All rights reserved.
This is an independent news site. Not affiliated with, endorsed by, or connected to OpenClaw. OpenClaw is a trademark of its respective owner.
Join the waitlist:

OC NEWS