Groq Audio Transcription in OpenClaw Can Fail With ‘request Content-Type isn't multipart/form-data’
OpenClaw News Editorial
A newly reported OpenClaw bug can make Groq-based audio transcription fail even when the surrounding voice-message pipeline is otherwise healthy.
The visible error is specific:
invalid_request_error: request Content-Type isn't multipart/form-data
That matters because Groq’s speech-to-text endpoint expects either a direct file upload or a URL input. For direct uploads, the request must be sent as proper multipart form data. If that encoding is wrong, Groq rejects the call before transcription ever starts.
Sources:
- OpenClaw issue: openclaw/openclaw#62605
- Groq Speech-to-Text docs: Groq Docs — Speech to Text
What is breaking
According to the report, the failing path looks like this:
- Telegram receives a voice message normally
- OpenClaw hands the audio into
tools.media.audio - Groq is configured as the transcription provider
- Groq rejects the request with
request Content-Type isn't multipart/form-data
The important detail is that the same voice-note flow reportedly works with Deepgram.
That makes this much less likely to be:
- a Telegram delivery problem
- a bad voice-note file from the sender
- a general audio-ingest failure inside OpenClaw
Instead, the failure appears to sit in the Groq transcription request path itself.
Why operators should take this seriously
If you use voice notes as a serious input channel, this is not a cosmetic error.
The practical impact is that users can keep sending audio successfully while transcripts never appear, which creates a confusing half-working system:
- the channel looks alive
- voice delivery succeeds
- another provider may work
- but Groq-backed transcription fails at request time
That kind of partial failure is easy to misdiagnose as a chat-platform glitch unless you inspect provider-specific logs.
Why this looks like a request-format bug, not a generic provider outage
The issue report gives three strong clues.
1) Deepgram works on the same incoming Telegram voice notes
If one provider succeeds on the same audio objects and only Groq fails, the bug is probably not the voice-note input itself.
2) The error is about HTTP request shape, not model quality or authentication
Groq is not returning a transcript-quality issue or a model refusal. It is rejecting the upload format before processing.
3) Groq’s own docs describe file transcription as a file-or-URL input flow
Groq’s speech-to-text docs list a file parameter for direct upload and describe the endpoint as OpenAI-compatible. In practice, that means the upload request needs the correct multipart encoding. A plain JSON body or a malformed upload wrapper can trigger exactly the error shown in the issue.
How to confirm you are hitting this exact failure mode
You are likely seeing the same bug if most of the following are true:
- you enabled audio transcription under
tools.media.audio - Groq is selected as primary or active transcription provider
- Telegram voice notes arrive normally
- Deepgram or another provider works when substituted
- the failing Groq log includes:
invalid_request_error: request Content-Type isn't multipart/form-data
That pattern is much narrower than a generic “audio transcription is broken” complaint.
What this is probably not
Based on the report, operators should avoid jumping to the wrong fix first.
This issue does not initially look like:
- a Telegram webhook or polling problem
- an unsupported audio file type coming from Telegram
- a missing Groq model id alone
- a user-side microphone problem
- a transcript rendering bug after inference
The failure seems to happen earlier: the Groq request is rejected before transcription work begins.
Immediate workaround
The reporter says Deepgram works as the primary provider on the same setup.
So if voice transcription is production-critical, the safest short-term move is:
- switch Groq out of the primary path
- use Deepgram for live transcription
- keep Groq only after upstream confirms the upload request is encoded correctly
That is not elegant, but it is the fastest way to restore user-visible reliability.
What maintainers likely need to inspect
The useful repair direction is fairly concrete.
Maintainers likely need to verify that the Groq transcription integration:
- constructs a real multipart form body for file uploads
- sets the matching
Content-Typeboundary correctly - does not accidentally serialize the audio request as JSON or another non-upload format
- handles Telegram voice-note inputs the same way as other accepted audio sources
Because the error is so specific, this looks like the kind of bug that should be reproducible with a small provider-level integration test.
What to capture if you escalate it internally
If you need to file a follow-up or hand this to another engineer, keep these details together:
- OpenClaw version
- Groq model used for transcription
- whether Telegram voice notes transcribe correctly with Deepgram
- the exact Groq error text
- whether the deployment uses any custom proxy layer
- a sanitized
tools.media.audioblock if possible
That will help separate this bug from unrelated transcription failures.
Bottom line
If Groq transcription suddenly fails in OpenClaw with request Content-Type isn't multipart/form-data, do not assume the whole voice pipeline is broken.
The report points to a narrower issue: OpenClaw may be sending Groq an invalid upload request format on the transcription path, even though the same voice-note flow still works with another provider such as Deepgram.
