Three small but high-impact fixes that all surfaced together while
verifying the long-form generation flow.
1. ChatConsole onBeforeUnmount no longer kills the backend turn.
Previously, switching tabs / route navigation / any cause that
unmounted the chat view called stopChatGeneration(), which POSTs
/chat/{cid}/stop and aborts the in-flight LLM call. The user
reported a turn dying mid-generation just from switching pages.
Replaced with resetForNewConversation() — front-end SSE disconnect
only, no /stop. Backend keeps running; pollActivity / status probe
reconnects on return. Aligns with the existing comment in
selectConversation: "let A's backend agent run continue running."
2. Agent max_iterations raised 25 → 100 with a hard ceiling.
The previous 25-step ceiling caused LimitExceededNode to fire on
substantive multi-tool tasks (document generation + image conversion
+ retry loops). 100 matches QwenPaw's _MAX_MAX_ITERATIONS upper
bound. New plumbing:
- BaseAgent.MAX_ITERATIONS_HARD_CEILING = 100 public constant
- BaseAgent default field 25 → 100 (Java-side fallback)
- AgentGraphBuilder clamps any per-agent DB override to the
ceiling at runtime; if the row holds 200, runtime sees 100 and
a WARN is logged with the original value.
- V47 migration (h2 + mysql) idempotently bumps the three default
seeded agents (1000000001, 1000000002, 1000000003) only if they
still hold the old defaults (25 / 20). User-customized values
are not touched.
- data-en/zh/-mysql-en/-mysql-zh seed files updated to 100 for
fresh installs.
3. DocxRenderTool tells the LLM not to prepend a host to the URL.
DeepSeek and Claude have both been observed wrapping the
/api/v1/files/generated/{id} relative path returned by renderDocx
into an absolute URL with a hallucinated domain (e.g.
https://ai-tools-system.com/...), breaking the download link in
the rendered chat bubble. The tool's return string now appends an
explicit "must use the relative path verbatim, do not add any
https:// or http:// prefix" instruction, which Claude and
DeepSeek both honor.
A bundle of stability fixes that all surfaced together while running
the same long-form generation task across multiple turns. Each one
addresses a distinct way the previous behavior silently dropped
content the user had already seen on screen.
1. Mid-turn narrative persistence (StateGraphReActAgent +
SummarizingNode). Intermediate ReasoningNode rounds and
SummarizingNode broadcast their content_delta directly to the
SSE channel for live display, but the StreamAccumulator only
received the final answer. After refresh the assistant message
showed only tool_call cards with no body text.
StateGraphReActAgent now also forwards STREAMED_CONTENT (already
set per round) as a persistOnly StreamDelta whenever it changes,
so every narrative chunk lands in the accumulator's content
buffer and gets written to mate_message. SummarizingNode now
writes its summary into the same key so summarize narratives
persist too.
2. Follow-up message queue, not dispose (ChatController#interruptStream).
Sending a new message while a turn was running called
requestInterrupt, which dispose()d the active Reactor chain mid
LLM call. That cancelled the in-flight generation, lost partial
tokens, and left the user staring at a half-finished bubble.
The endpoint now uses enqueueMessage in all paths, matching
the "wait for current turn, then run" behavior. The old
requestInterrupt API is kept for any future force-replace UI
but no caller routes to it.
3. Queued user message ordering (ChatStreamTracker.QueuedInput +
ChatController.startQueuedMessage). interruptStream used to save
the queued user message immediately, before the in-flight
assistant message finalized in doOnError. listMessages orders
by create_time ASC, so the queued user message ended up above
the assistant reply it was supposed to follow. QueuedInput now
carries contentParts; persistence is delayed to startQueuedMessage,
which runs only after Asst-N is on disk.
4. JVM shutdown flush (ChatStreamTracker @PreDestroy +
emergencySaveAccumulator). A mvn spring-boot:run restart used to
wipe in-flight turns: SSE emitter timed out, ShutdownHook fired,
HikariPool closed before doOnError could save. ChatStreamTracker
now exposes an emergency-save callback per RunState; ChatController
registers one per stream that snapshots the accumulator and
writes status="interrupted_shutdown". @PreDestroy walks active
runs, invokes the callback, then disposes. Spring's reverse-order
bean teardown keeps ConversationService and Hikari alive long
enough for the save to complete.
5. Observation thresholds for summarize (GraphObservationProperties +
application.yml). The previous total-chars threshold of 12 KB
triggered summarize after one or two RFC reads, costing a 40 to
80 second compaction LLM call per loop. Tuned to: total 200 KB,
single 16 KB, large-result 32 KB, rounds safety net 25. Java
field defaults reverted to the conservative original values so
application.yml stays the source of truth.
6. Frontend thinking segmentation (useChat.ts thinking_delta +
phase). Multi-round ReAct turns merged every reasoning + summarize
round's thinking into one segment, accumulating to 9 KB+ in a
single bubble. thinking_delta now uses findLast(running) so a
tool_call_started or phase transition closes the previous segment
and the next delta opens a fresh one. phase event also closes
running thinking/content segments.
7. Other small things bundled: removed a debug metadata-keys log
that flooded the log file with one line per stream chunk; fixed
three stale tests that didn't compile after earlier constructor
changes (WikiLogServiceTest, WikiOverviewSpliceTest,
WikiProcessingServiceLazyTest); added rfc-066 documenting the
unified message queue + priority refactor as the next logical
step on top of these stabilizations.
Verified end-to-end with multiple full sessions: a four-minute
generation that produced the expected docx and a follow-up enqueue
that ran cleanly after the previous turn naturally completed,
without the old "Disposable unavailable" interrupt path.
- Drag-over highlight (orange border + shadow + arrow icon) on the upload
zone, using dragCounter to prevent flicker over nested children
- Optimistic list items appear immediately on drop/select with UPLOADING
badge and progress bar, before the HTTP request completes
- Wire axios onUploadProgress through api.uploadRaw → store.uploadRawFile
so the progress bar tracks real byte transfer (0–100%), with an
indeterminate shimmer until the first tick
- try/catch around every upload: on failure, placeholder flips to an error
state with ElMessage.error toast and a × dismiss button
- Upload all dropped/selected files concurrently via Promise.all
- i18n keys added (zh-CN + en-US): dropToUpload, uploading, uploadFailed,
status.uploading, progress.uploading
Three bugs fixed:
1. Badge stays "processing" after job completes: pollJobs() never called
fetchRawMaterials() when a job reached terminal status, so
raw.processingStatus stayed processing in the store. Fix: detect
terminal job status in pollJobs, trigger fetchRawMaterials to sync.
2. "Completed" dot pulses instead of solid: dotClass() treated completed
the same as in-progress (target === cur → active). Fix: add
isTerminal computed (includes completed), return done for all
dots at or before the terminal position — no pulse animation.
3. Stage label stays orange at terminal: same cause — active class
applied regardless of terminal state. Fix: use done class for
terminal labels (green instead of orange).
Also: v-if on JobStageBar now shows for terminal status jobs (not just
stage !== queued), so completed/failed stage bars remain visible.
Root cause: processRawMaterial() created a job record at queued stage
but never called jobService.transition() during processing. The job row
stayed at queued forever, so the stage bar never advanced.
Backend (WikiProcessingService):
- Transition job to ROUTING immediately after creation
- Transition to PHASE_A_RUNNING before chunk processing begins
- Transition to COMPLETED/PARTIAL/FAILED at the end based on finalStatus
- Transition to FAILED in the catch block on unhandled exceptions
Backend (WikiProcessingJobService.transition):
- Handle FAILED, PARTIAL terminal stages (set finishedAt + status)
- Handle non-terminal intermediate stages (set status to running)
Frontend (JobStageBar.vue):
- Add stageMapping for backend stages not shown as dots: phase_a_done →
phase_b_running, failed/partial/cancelled → completed position
- Guard stageIndex() against -1 (unknown stages default to all-pending)
- Terminal failure states show red failed dot instead of pulsing active
Root cause: addFile()/addText() hash dedup only matched rows with
status=completed, so the same file uploaded while in partial/pending/
processing/failed status would create a duplicate row.
Fix:
- Remove .eq(processingStatus, "completed") from dedup queries — match
any non-deleted row with the same content hash in the KB
- On dedup hit: completed/pending/processing → return as-is;
partial/failed → trigger reprocess (partial enters resume branch)
- Clean up the newly uploaded temp file when dedup discards it
- Frontend: uploadRawFile/addRawText check for existing id in the list
before unshift to prevent visual duplicates
Test: WikiRawMaterialDedupTest — 10 cases covering all 5 statuses,
reprocess triggers for partial/failed, no-op for others, insert only
when no match.
Three root causes for broken progress:
1. JobStageBar was shown whenever a job record existed (even at queued
stage), hiding the working SSE-driven progress bar. Fix: only show
JobStageBar when job.stage !== queued.
2. SSE connection only opened when hasProcessing was true (status ===
processing), but reprocess sets status to pending first. Fix:
include pending in the hasProcessing check.
3. After reprocess, if processing finished before SSE connected, the
status badge stayed on pending forever. Fix: immediately set local
status to processing after reprocess API call, clear stale job
entries, and add delayed re-fetches (5s/15s) as safety net.
Also: clear rawJobs entries on raw.completed/raw.failed SSE events to
prevent stale JobStageBar from lingering after processing ends.
Two related issues from the Kimi-401 user report:
1. Backend (NodeStreamingChatHelper): a primary AUTH_ERROR (e.g. Kimi 401
with an invalid API key) returned immediately without trying the
fallback chain — a fallback provider with a different, valid key
never got a chance. Even with DashScope correctly configured as the
fallback, the user chat dead-ended on a 401.
The original assumption ("auth never self-heals so do not retry")
holds for the primary same-model retry loop but is wrong for the
fallback chain — different providers have different keys. Apply the
same break-into-fallback policy that BILLING and MODEL_NOT_FOUND
already use. recordPrimary(false) is preserved so the cooldown
counter still accumulates.
2. Frontend (chatError.ts + i18n): the error-text matching for
/认证|auth|unauthorized|401/i was so broad it matched the substring
"auth" inside URLs like https://api.kimi.com/.../auth, classifying
any model 401 as user "session expired" and rendering the misleading
"页面将自动跳转到登录页" copy. (The redirect itself only fires from
/api/v1/auth/* axios paths and SSE-connection 401s, not from this
payload-text path — but the copy alone is the worst kind of false
alarm.)
Add a new ChatErrorCategory provider_auth_error and split the
pattern matching: narrow auth_expired (HTTP 401 / 登录已过期 /
session expired / 凭证失效) is matched FIRST, then the broad
401-ish pattern routes to provider_auth_error. BACKEND_ERROR_TYPE_MAP
for AUTH_ERROR is also remapped, since structured backend payloads
currently always come from LLM providers — never from our own
/api/v1/auth path.
Tests
- NodeStreamingChatHelperFailoverTest (5 cases): primary 401 →
fallback succeeds; chain skips auth-failing fallback to next healthy
one; whole-chain failure surfaces last AUTH_ERROR (no silent drop);
BILLING regression unchanged; primary-success path does not touch
chain
- Browser preview verified: new i18n keys resolve in en-US, classifier
correctly routes "[错误] 401 from kimi.com" → provider_auth_error
while "[错误] HTTP 401 from /api/v1/auth/ping" stays auth_expired
- 186 tests pass (was 181 + 5 new); vue-tsc clean
Do-not-touch list: handleAuthFailure() in useStream/api/index.ts (real
session-expiry path) is unmodified — only the misclassification
upstream is fixed. auth_expired i18n copy is unchanged.
UI — Failover priority editor
- ProviderConfigRequest + ProviderInfoDTO carry fallbackPriority
- ModelProviderService.updateProviderConfig persists it (null = unchanged);
toProviderInfo exposes the current value to the UI (defaults to 0)
- ProviderConfigModal advanced panel exposes a number input with hint
- ProviderCard shows a "Fallback #N" badge for chain members so the
priority order is visible at a glance without opening the modal
- 5 new i18n keys (zh + en) — verified to resolve at runtime via i18n.global.t
Backend — Per-provider health tracker
- ProviderHealthTracker: ConcurrentHashMap-backed counters; N consecutive
failures (default 3) push the provider into a cooldown window (default
5 min) during which the chain walker skips it. Success resets both
counter and cooldown atomically. Lazy expiry on lookup so dead entries
do not accumulate.
- ProviderHealthProperties exposed under mateclaw.llm.failover.health.*
with sane production defaults
- New FallbackEntry record (providerId + ChatModel) replaces raw
List<ChatModel> in the chain so the walker can correlate cooldown
state to entries; AgentGraphBuilder.buildFallbackChain returns the
new type
- NodeStreamingChatHelper takes the tracker through a new 4-arg
constructor and consults it before each fallback call; records
success/failure on each chain attempt. Legacy 2/3-arg constructors
preserved as @Deprecated wrappers (synthetic providerId means no
health tracking on the legacy path — that path is opt-out anyway)
Tests
- ProviderHealthTrackerTest (9 tests): below/at threshold, success
reset, cooldown expiry (via reflection on the min-clamp setter),
disabled-tracker no-op, null-providerId safety, per-provider
isolation, snapshot output
- NodeStreamingChatHelperFallbackChainTest updated to FallbackEntry
field type — verifies providerId + ChatModel survive the chain
- 168 tests pass (was 159 + 9 new)
Verification
- mvn test green; vue-tsc clean; live UI confirms i18n resolution
Pasted prompts (test cases, structured asks, JSON dumps) currently render
through the same markdown pipeline as assistant output, so '#'/'-'/'**'
characters are processed and long prompts dominate the scrollback.
- New UserMessageContent.vue: plain-text rendering (white-space: pre-wrap
preserves user-typed newlines and indentation), with auto-collapse beyond
8 lines and a "Show more (N more lines) / Show less" toggle. Soft mask
gradient at the collapse boundary instead of a hard cut.
- MessageBubble.vue: route role==='user' messages through the new component;
assistant messages keep the existing markdown pipeline unchanged.
- Add chat.expandLines / chat.collapse i18n keys (zh + en).
Verified end-to-end in browser preview: 15-line content collapses to 8,
toggle expands to full 15 with "Show less" label, raw '#' / '**' / '`' chars
shown literally with no <strong>/<h1>/<li> tags emitted.
- db/migration/mysql: replace ADD COLUMN/CREATE INDEX IF NOT EXISTS with
idempotent checks via information_schema (MySQL 8.0 <8.0.29 and some
forks don't support IF NOT EXISTS for ADD COLUMN). Affects V2/V4/V5/V7
/V8/V9/V11/V12/V13/V14. Fix: gitee#IIYHLJ.
- application-mysql.yml: add createDatabaseIfNotExist=true so MySQL
Connector/J auto-creates the schema on first connection (requires
CREATE privilege — documented fallback for restricted accounts).
- llm/OllamaAutoDiscoveryRunner: rewrite seed tag when fuzzy-matching,
prefer exact tag for default; skip models without tool support when
auto-activating a default (prevents the phantom ':latest' trap when
users pulled a specific size).
- agent/graph/NodeStreamingChatHelper: detect 'does not support tools'
and 'model not found' errors from Ollama and emit actionable Chinese
prompts guiding users to qwen3 / qwen2.5:7b+ / llama3.1:8b+ etc.
- ui/MessageBubble + types/chatError: surface the backend's actionable
rawMessage in the failed-message card instead of a generic '未知错误';
strip redundant prefixes (Bad request: / [错误] / LLM 调用失败:) since
the title already conveys the category.