Skip to content

Transcript Hygiene: Managing Provider-Specific Fixes

I have spent far too much time debugging why an LLM provider rejects a transcript that looks perfectly fine. Usually, it is something small: a tool ID has a character the API hates, or the turn order does not alternate exactly how the provider expects. It is frustrating when these strict rules break your model context.

To solve this, we use “Transcript Hygiene.” These are in-memory adjustments that happen right before we build the model context. They do not change your JSONL files on disk, but they ensure the provider gets exactly what it needs to function.

You can find the core logic for these fixes in these files:

  • src/agents/transcript-policy.ts (Policy selection)
  • src/agents/pi-embedded-runner/google.ts (Sanitization and repair application)
  • src/agents/session-file-repair.ts (Session file repair logic)
  • src/agents/pi-embedded-helpers/images.ts (Image sanitization)

The hygiene process runs automatically in the embedded runner. Here is the 5-minute breakdown of how it works:

  1. Policy Selection: The runner looks at your provider, modelApi, and modelId to decide which rules to apply.
  2. Image Sanitization: This is a global rule. We downscale or recompress oversized base64 images so the provider does not reject them.
  3. Malformed Tool Calls: We drop assistant tool-call blocks that are missing both input and arguments. This prevents errors from partially saved calls.
  4. Provider Fixups: The runner applies specific logic based on the provider matrix.

Different providers have different requirements. Here is how we handle them:

  • OpenAI / OpenAI Codex: We keep it simple. We apply image sanitization and drop orphaned reasoning signatures when you switch models. We do not touch tool IDs or turn ordering here.
  • Google (Gemini/Antigravity): This is the strictest. We enforce alphanumeric tool IDs, repair tool result pairing, and ensure turns alternate correctly. If a history starts with an assistant turn, we add a tiny user bootstrap turn.
  • Anthropic / Minimax: We repair tool result pairing and merge consecutive user turns to satisfy the strict alternation requirement.
  • Mistral: We enforce “strict9” tool call IDs (alphanumeric, length 9).
  • OpenRouter Gemini: We strip thought_signature values unless they are base64.

If your session file itself is malformed, we handle that before the session even loads.

Problem: The session file has invalid lines or is corrupted. Solution: repairSessionFileIfNeeded runs during the load process. It drops invalid lines and creates a backup of your original file alongside the session file. This happens in run/attempt.ts and compact.ts.

Problem: Provider rejects tool call IDs. Solution: Check if you are using a provider like Mistral or Google. We automatically sanitize these to alphanumeric strings (or strict9 for Mistral) in memory.

Managing different provider requirements does not have to be a manual chore. By centralizing these fixups in the runner, we keep the stored data clean while keeping the APIs happy.

If you have specific questions about your setup, check out the AI Setup Assistant.

OpenClaw

OpenClaw Expert

Still stuck?

If this page didn't answer your case, ask OpenClaw Expert for step-by-step guidance.