A bundle of stability fixes that all surfaced together while running
the same long-form generation task across multiple turns. Each one
addresses a distinct way the previous behavior silently dropped
content the user had already seen on screen.
1. Mid-turn narrative persistence (StateGraphReActAgent +
SummarizingNode). Intermediate ReasoningNode rounds and
SummarizingNode broadcast their content_delta directly to the
SSE channel for live display, but the StreamAccumulator only
received the final answer. After refresh the assistant message
showed only tool_call cards with no body text.
StateGraphReActAgent now also forwards STREAMED_CONTENT (already
set per round) as a persistOnly StreamDelta whenever it changes,
so every narrative chunk lands in the accumulator's content
buffer and gets written to mate_message. SummarizingNode now
writes its summary into the same key so summarize narratives
persist too.
2. Follow-up message queue, not dispose (ChatController#interruptStream).
Sending a new message while a turn was running called
requestInterrupt, which dispose()d the active Reactor chain mid
LLM call. That cancelled the in-flight generation, lost partial
tokens, and left the user staring at a half-finished bubble.
The endpoint now uses enqueueMessage in all paths, matching
the "wait for current turn, then run" behavior. The old
requestInterrupt API is kept for any future force-replace UI
but no caller routes to it.
3. Queued user message ordering (ChatStreamTracker.QueuedInput +
ChatController.startQueuedMessage). interruptStream used to save
the queued user message immediately, before the in-flight
assistant message finalized in doOnError. listMessages orders
by create_time ASC, so the queued user message ended up above
the assistant reply it was supposed to follow. QueuedInput now
carries contentParts; persistence is delayed to startQueuedMessage,
which runs only after Asst-N is on disk.
4. JVM shutdown flush (ChatStreamTracker @PreDestroy +
emergencySaveAccumulator). A mvn spring-boot:run restart used to
wipe in-flight turns: SSE emitter timed out, ShutdownHook fired,
HikariPool closed before doOnError could save. ChatStreamTracker
now exposes an emergency-save callback per RunState; ChatController
registers one per stream that snapshots the accumulator and
writes status="interrupted_shutdown". @PreDestroy walks active
runs, invokes the callback, then disposes. Spring's reverse-order
bean teardown keeps ConversationService and Hikari alive long
enough for the save to complete.
5. Observation thresholds for summarize (GraphObservationProperties +
application.yml). The previous total-chars threshold of 12 KB
triggered summarize after one or two RFC reads, costing a 40 to
80 second compaction LLM call per loop. Tuned to: total 200 KB,
single 16 KB, large-result 32 KB, rounds safety net 25. Java
field defaults reverted to the conservative original values so
application.yml stays the source of truth.
6. Frontend thinking segmentation (useChat.ts thinking_delta +
phase). Multi-round ReAct turns merged every reasoning + summarize
round's thinking into one segment, accumulating to 9 KB+ in a
single bubble. thinking_delta now uses findLast(running) so a
tool_call_started or phase transition closes the previous segment
and the next delta opens a fresh one. phase event also closes
running thinking/content segments.
7. Other small things bundled: removed a debug metadata-keys log
that flooded the log file with one line per stream chunk; fixed
three stale tests that didn't compile after earlier constructor
changes (WikiLogServiceTest, WikiOverviewSpliceTest,
WikiProcessingServiceLazyTest); added rfc-066 documenting the
unified message queue + priority refactor as the next logical
step on top of these stabilizations.
Verified end-to-end with multiple full sessions: a four-minute
generation that produced the expected docx and a follow-up enqueue
that ran cleanly after the previous turn naturally completed,
without the old "Disposable unavailable" interrupt path.
Same bug as the prior queue-drop fix in doOnComplete, but in the
sister branch that fires when the agent's reactive stream errors
out (CancellationException from a user stop). The guard
cr.queuedInput() != null && !(isUserStop && !isInterruptFollowup)
mis-classified "user stopped, no interrupt-with-followup, but a
message is in the queue" as an explicit abort and silently dropped
the freshly-typed follow-up.
The frontend's enqueue path never sets interruptType — it just
calls requestStop + offers to messageQueue. Whoever puts a message
in the queue means it; just run it. Aligns with doOnComplete and
the four other queue-launch sites in this controller.
A series of cross-cutting stability fixes that surfaced together
during a long debugging session.
reasoning_content / Claude prefill self-replicating 400:
- ChatController persists typed errors (content starts with '[错误] ')
with status='error', so the failure text stops being re-sent as
multi-turn context — DeepSeek thinking 400 ('reasoning_content
must be passed back') and Claude 400 ('does not support assistant
message prefill') used to recursively re-create themselves every
retry by polluting history.
- BaseAgent.sanitizeForLlm filters status='error' / '[错误] ' prefix
assistant messages from history before LLM dispatch.
- BaseAgent.fetchHistoryMessages defensively drops trailing
AssistantMessages — Claude rejects assistant-tail prompts.
- NodeStreamingChatHelper.dropTrailingAssistant runs the same
defense at every doStreamCall pre-egress, so the in-turn
summarizing→reasoning transition (which leaves an assistant
scaffold at the tail) doesn't trip Claude either.
- AgentGraphBuilder.FallbackPolicy.DEEPSEEK switched (null,true,true)
→ (' ',false,true), aligning with KIMI/OPENAI's tolerant ' '
fallback. The previous 'force explicit 400' design was the
self-replicating loop's prime mover.
narration + tool args truncation:
- ReasoningNode.DEFAULT_MAX_OUTPUT_TOKENS 4096 → 16384. The 4k cap
was decapitating renderDocx tool_call args mid-stream when the
model emitted a long content field on top of thinking content;
the resulting 'invalid JSON' aborted execution silently.
- ReasoningNode appends a hermes-style TOOL_USE_ENFORCEMENT clause
to every system prompt: 'when you say you will perform an action,
call the tool now in the same response — narration is a protocol
violation'. Treats 'now I will generate the docx' (and never
actually calling renderDocx) as a forbidden pattern.
- ToolExecutionExecutor.normalizeToolExecutionError reframes the
JSON-truncated error as actionable instructions: 're-call the
same tool now with shorter content or split into multiple
sequential calls; do NOT describe the result as text'.
side fixes from the same evening:
- ChatController doOnComplete skips completionPublisher.publish
when isError=true, keeping memory extraction off the garbage path.
- ChatController doOnComplete queued-message guard simplified to
'cr.queuedInput() != null', matching the other 4 sites in the
controller. The previous 'isInterruptFollowup || !wasStopped'
guard silently dropped queued messages when the user did
Stop-then-Enqueue (wasStopped=true && interruptType=null), losing
the freshly-typed follow-up message.
- prompts/graph/summarize-system.txt now distinguishes 'single
task' (default; output one cohesive summary) from 'multiple
independent sub-tasks' (use the子任务 N format). Stops the
summarizer from inventing '子任务 1: PRO-027' decomposition for
unitary requests like 'write me a project proposal'.
- ChannelMessageRouter: include assistantMessageId in message_complete
and done broadcasts so ChatConsole observers can reconcile the
streaming placeholder to the persisted DB row by id instead of
falling back to the FIFO 'claim' heuristic (which occasionally
dropped the assistant bubble on external channel conversations).
Capture the id from ConversationService.saveMessage in both the
sync agentService.chat path and the streaming processWithStreaming
path; switch to HashMap since Map.of rejects null values when save
is skipped (e.g. under approval).
- ChatConsole: add a 'running' indicator on the sidebar so users can
tell which conversations have an in-flight agent run. Pulsing amber
dot on the channel icon (both expanded and collapsed modes) plus a
'生成中…' / 'Generating…' pill in expanded mode.
- ChatConsole: don't cancel the previous conversation's streaming run
when switching conversations — let it keep running in the background
and reconcile when the user comes back.
- ConversationWindowManager: cap reserve token at 50% of effective max
to prevent negative historyBudget on small-context models (8K/16K)
- common.security.SecretEquals: new constant-time comparison utility
(MessageDigest.isEqual wrapper) for secrets/tokens/signatures
- WeixinChannelAdapter: migrate context_token comparison to SecretEquals
- FeishuChannelAdapter: fail-fast on empty encrypt_key when connection_mode=webhook
- TelegramChannelAdapter: sanitize attachment captions — strip control bytes
(\p{Cc} except \t\r\n) + format chars (\p{Cf}) + 4096 char cap
- AgentGraphBuilder: fallback Anthropic max_tokens to 4096 on null/0/negative
Tests: SecretEqualsTest (5) + TelegramCaptionSanitizeTest (5) — all green.
Full-stack AI assistant built on Spring AI Alibaba.
Features: ReAct Agent, Plan-and-Execute, MCP Protocol, Multi-Model, Multi-Channel.
Apache-2.0 License