Some OpenAI-compatible providers (LM Studio's built-in server, certain
strict-mode vLLM / SGLang deployments) reject 400 "System message must
be at the beginning" when SystemMessages appear after user / assistant
/ tool messages. The reasoning loop currently emits four SystemMessage
segments — main prompt at index 0, skill catalog inserted at index 1,
progress-ledger snapshot and stale-reminder appended at the end of
nonHistoryPrefix after the runtime-context UserMessage. The latter two
violate the strict shape, so conversations on LM Studio 400 on the
first turn (reported in #218).
Add MessageNormalizer: collects every SystemMessage in the outbound
prompt regardless of position, joins their text with a blank-line
separator, and emits a single SystemMessage at index 0. Non-system
messages keep their relative order, so AssistantMessage(tool_calls) ↔
ToolResponseMessage adjacency is preserved verbatim (required by strict
pair validators).
Wire it into doStreamCall as the first pre-egress step so every node
(reasoning, step-execution, summarizing, plan-generation, limit-exceeded)
inherits the fix without per-node changes, and any future node that
emits multiple SystemMessages stays compliant.
The transformation is semantically equivalent on permissive providers
(OpenAI, DashScope, Ollama, DeepSeek, Kimi, Doubao, GLM) — the merged
token sequence matches what they would have seen across N SystemMessages
— and safe on non-OpenAI protocols (Anthropic, Vertex / Gemini), whose
adapters already extract SystemMessages into a top-level system field
and receive an identical payload.
Kill switch: -Dmateclaw.llm.message-normalizer.enabled=false reverts to
the prior behavior for emergency rollback.
Tests: 11 unit tests on MessageNormalizer cover empty / no-system /
canonical / mid-list / tail / blanks / tool-pair preservation / Prompt
option-reference preservation / kill switch. 1 wiring test pins the
call site in doStreamCall. Full vip.mate.agent.** suite (504 tests)
stays green.
Closes#218.
Some providers (notably SiliconFlow) return "network connection error" in the response body when their backend is overloaded or the upstream model connection is disrupted. classifyError() had no pattern for this string, so it fell through to UNKNOWN (non-retryable), surfacing the raw error to the user on the first failure instead of running the exponential-backoff recovery. Adds the pattern to the SERVER_ERROR classifier and a friendly message mapping in extractUserFriendlyError(); bumps MAX_RETRIES from 5 to 10 so sustained wiki batch load can ride out provider flaps without surfacing an error to the channel user.
Closes#178
When a user-installed skill (e.g. RedisOps) was bound to an agent, the
model frequently called the skill name directly as a tool, hit
"Tool not found: RedisOps", and either gave up or fell back to shell
guessing. Two compounding causes:
1. The system prompt block injected by SkillRuntimeService listed each
skill as `- **RedisOps** — desc`, which is the same format used for
tool catalogs and primed the model to call the names directly. The
"how to use" instructions referenced `read_skill_file` /
`run_skill_script` — names that don't exist in the tool registry,
so even a compliant LLM couldn't follow them.
2. ToolExecutionExecutor's `callback == null` branches returned a bare
"Tool not found: <name>" string. The model had no recovery signal
and no hint that the name it called was actually a skill.
Fix is two-layered:
- Prompt rewrite (SkillRuntimeService.buildSkillPromptEnhancement): lead
with an explicit warning that skills are NOT directly callable, use the
correct camelCase tool names (readSkillFile / runSkillScript), include
a concrete worked example anchored to the first enabled skill, and
render the listing as a markdown table so it stops looking like a
callable tool list. listAvailableSkills tool description and output
follow the same pattern.
- Runtime safety net (ToolExecutionExecutor): when toolCallbackMap.get
misses, check if the requested name (case-insensitive) matches an
active skill. If so, return a precise hint telling the LLM the right
invocation pattern instead of the bare error. Wired through both the
main execute path and the pre-approved replay path. SkillRuntimeService
is attached via a setter from AgentGraphBuilder so the executor's many
legacy constructors stay untouched, and it's nullable so isolated
tests still work.
Adds 5 unit tests covering: skill match -> hint, case-insensitive match,
no-match -> bare error, no SkillRuntimeService wired -> bare error,
pre-approved replay path -> hint.
Reported and reproduced by @pipima9950-glitch in issue #46.