Five-commit bundle brings the Dream v2 P1 engine layer online, sitting
on top of the lifecycle mediator foundation already merged.
B.1-B.4 · Schema + records
- Flyway V26 (dream_report) + V27 (memory_recall review fields),
both h2 and mysql
- DreamReportEntity + DreamMode + DreamStatus enum + record types
- DreamReportMapper repository layer
B.5-B.8 · Consolidate refactor + focused dream
- MemoryEmergenceService refactored for plug-in dream modes
- MemoryRecallService extended with promoted/rejected review fields
- Focused dream endpoint + prompt template
- MemoryController exposes the review/trigger surface
B.9-B.10 · Monthly archive service
- MemoryArchiveService rolls cold promoted entries into archival rows
and reclaims daily_count storage
- DreamingScheduler runs archive job on its own schedule
B.12-B.14 · Tests
- MemoryArchiveServiceTest
- DreamFlagGuardTest
- DreamV2AcceptanceIT (end-to-end acceptance under feature flag)
Plus a verification script + HTTP e2e kit in the private test/ dir,
used for local staged rollout — not part of the open-source
distribution.
All features stay gated behind the mate.memory.dream.* flags from
Phase 1. Enable per-phase after staging validation.
beforeLlmCall / afterLlmCall / onSessionEnd only logged on failure,
making flag on/off indistinguishable in logs. Add debug lines on the
success path so lifecycle activation is observable.
Wire memory-facing events (turn-started, turn-completed, session-ended,
memory-written) through a single MemoryLifecycleMediator so
MemoryProvider implementations can hook into the agent conversational
flow without spreading side-effects across the runtime.
Ten atomic steps shipped under feat/dream-v2-p1-lifecycle:
- A.1 + A.2: MemoryLifecycleMediator class + TurnContext value object
- A.3: TurnStartedEvent / TurnCompletedEvent domain events
- A.4: MemoryLifecycleEventListener bean for Spring event plumbing
- A.5: MemoryProvider.onMemoryWrite default method (backward compatible)
- A.7: wire the mediator into AgentService at the right hook points
- A.8: LifecycleFlagGuardTest — feature flag must gate every hook
- A.9: MemoryLifecycleMediatorTest — unit coverage per hook
- A.10: LifecycleRecallCountIT — F4 regression across the stack
Feature flags (all default OFF; enable per phase after staging):
- mate.memory.lifecycle-mediator-enabled
- mate.memory.dream.focused-enabled
- mate.memory.dream.archive-enabled
This is Phase 1 foundation only — focused-dream and archive-dream
providers arrive in later phases.
Root cause: processRawMaterial() created a job record at queued stage
but never called jobService.transition() during processing. The job row
stayed at queued forever, so the stage bar never advanced.
Backend (WikiProcessingService):
- Transition job to ROUTING immediately after creation
- Transition to PHASE_A_RUNNING before chunk processing begins
- Transition to COMPLETED/PARTIAL/FAILED at the end based on finalStatus
- Transition to FAILED in the catch block on unhandled exceptions
Backend (WikiProcessingJobService.transition):
- Handle FAILED, PARTIAL terminal stages (set finishedAt + status)
- Handle non-terminal intermediate stages (set status to running)
Frontend (JobStageBar.vue):
- Add stageMapping for backend stages not shown as dots: phase_a_done →
phase_b_running, failed/partial/cancelled → completed position
- Guard stageIndex() against -1 (unknown stages default to all-pending)
- Terminal failure states show red failed dot instead of pulsing active
Two related changes that align buildFallbackChain with how users actually
think about failover.
1) Source = configured providers (was: only providers with fallback_priority > 0)
Earlier the chain was strictly "providers the user explicitly opted in via
fallback_priority > 0". A healthy in-pool provider with priority=0 was
silently excluded — surprising since the pool was supposed to be the source
of truth for "what is usable". After this change:
- Candidates = every configured provider
- Pool gating = same as before (in-pool members only at build time;
runtime walker re-checks)
- Order = agent prefs (PR-3) → fallback_priority asc (>0) →
priority==0 alphabetical
So fallback_priority is now purely an ordering hint, never an exclusion.
2) Per-provider model picker = default OR first-enabled (was: default only)
Previously a provider was skipped if no chat model on it had is_default=true.
That is admin friction with no benefit — every provider had to be visited in
Settings just to mark a default before it could appear in failover. New
pickFallbackModel():
- first try getDefaultModelByProvider — user explicit pick wins
- otherwise take the first enabled chat model on the provider
- skip only if neither exists
User-visible effect on the deployment that surfaced this:
- kimi-code primary fails (401 — real auth issue, separate from this bug)
- Pool short-circuits primary → walker fires
- Walker now sees dashscope (in-pool) AND ollama (in-pool) as candidates,
even though neither has fallback_priority set
- dashscope first enabled qwen model is picked → request succeeds via
dashscope without anyone touching Settings
45 failover-related tests still green (unit-level chain-build behavior is
backward-compatible; only the candidate set and model-selection lookups
changed, both broadening the chain rather than narrowing it).
Two real bugs the user restart surfaced — both turned healthy providers
into HARD-removed false positives.
Bug #1 — URL duplication
OpenAiCompatibleListModelsProbe always concatenated /v1/models, so
providers whose Base URL already includes the version segment got the
wrong URL:
LMStudio http://localhost:1234/v1 → /v1/v1/models → 404
ZhipuAI .../api/paas/v4 → /v4/v1/models → 404
Fix: detect a trailing /vN suffix and append /models instead. Six unit
tests in OpenAiCompatibleListModelsProbeTest lock the rule down.
Bug #2 — 404 false positives
Kimi for Coding API does not expose /v1/models even though chat works
fine, so the probe correctly received a 404 and incorrectly HARD-removed
the provider from the pool. Other vendors will hit the same — listing
is not a universal contract.
Fix: classify HTTP responses semantically.
401 / 403 → HARD remove (real auth failure)
404 / 405 / 410 → fail-open (endpoint missing, server may be alive)
other 4xx / 5xx → fail-open (probe inconclusive — let chat decide)
network errors → fail (unreachable)
This is the same philosophy as ChatGPTOAuthStatusProbe: when we cannot
cheaply confirm health, we do not proactively penalize the provider.
Same logic applied to Anthropic + DashScope probes for consistency.
Net effect on the user deployment after restart:
- kimi-code stays in pool (404 → fail-open) → primary path works again
- lmstudio + zhipu-cn also stay in pool (URL bug fixed)
- dashscope + ollama unchanged (real 200 OK)
Tests: 6 new for resolveModelsPath. The 2 unrelated WikiRawMaterialDedupTest
failures pre-date this commit and live in ba86bea.
Root cause: addFile()/addText() hash dedup only matched rows with
status=completed, so the same file uploaded while in partial/pending/
processing/failed status would create a duplicate row.
Fix:
- Remove .eq(processingStatus, "completed") from dedup queries — match
any non-deleted row with the same content hash in the KB
- On dedup hit: completed/pending/processing → return as-is;
partial/failed → trigger reprocess (partial enters resume branch)
- Clean up the newly uploaded temp file when dedup discards it
- Frontend: uploadRawFile/addRawText check for existing id in the list
before unshift to prevent visual duplicates
Test: WikiRawMaterialDedupTest — 10 cases covering all 5 statuses,
reprocess triggers for partial/failed, no-op for others, insert only
when no match.
PR-0 only installed the strategy seam; the actual ~600 LOC of provider-
specific construction stayed in AgentGraphBuilder as transitional public
helpers. PR-0b moves the DashScope + Anthropic halves into their builders
proper. (OpenAI larger refactor — 5 sub-helpers including Kimi/o-series
special cases — is left for a follow-up PR-0c.)
AgentDashScopeChatModelBuilder now owns:
- buildDashScopeApi (with provider/env/reflection key+url fallback chain)
- buildDashScopeOptions (model/temp/max-tokens/topP + built-in search)
- normalizeDashScopeBaseUrl (strip /compatible-mode/, return null for SDK default)
- readApiKeyFromDefaultChatModel + readBaseUrlFromDefaultChatModel +
readDashScopeApiFromDefaultChatModel (reflection-based final fallback)
- isBuiltinSearchEnabled (renamed from isDashScopeSearchEnabled, called
by AgentGraphBuilder.build via the now-injected dashScopeBuilder ref)
AgentAnthropicChatModelBuilder now owns:
- buildAnthropicApi (key validation, applyHttpTimeouts duplicated locally)
- buildAnthropicOptions (extended-thinking budget mapping low/medium/high/max
→ 4k/8k/16k/32k, temperature=1 enforcement, RFC-014 prompt cache options)
AgentGraphBuilder dropped:
- DashScope: ~120 LOC (api + options + 4 helpers + isDashScopeSearchEnabled)
- Anthropic: ~75 LOC (api + options)
- DashScopeChatModel + DashScopeConnectionProperties fields (unused after move)
- Deprecated single-fallback buildFallbackModel (no callers, superseded
by buildFallbackChain since RFC-009 PR-1)
- 5 imports for moved DashScope/Anthropic types
Net: -154 LOC in AgentGraphBuilder (1721 → 1567), +372 across the two new
builders. Strategy seam is now real for 3 of 4 protocols (ChatGPT was
already standalone, OpenAI is PR-0c). 220/220 tests still green — no
behavior change.
Two related issues from the Kimi-401 user report:
1. Backend (NodeStreamingChatHelper): a primary AUTH_ERROR (e.g. Kimi 401
with an invalid API key) returned immediately without trying the
fallback chain — a fallback provider with a different, valid key
never got a chance. Even with DashScope correctly configured as the
fallback, the user chat dead-ended on a 401.
The original assumption ("auth never self-heals so do not retry")
holds for the primary same-model retry loop but is wrong for the
fallback chain — different providers have different keys. Apply the
same break-into-fallback policy that BILLING and MODEL_NOT_FOUND
already use. recordPrimary(false) is preserved so the cooldown
counter still accumulates.
2. Frontend (chatError.ts + i18n): the error-text matching for
/认证|auth|unauthorized|401/i was so broad it matched the substring
"auth" inside URLs like https://api.kimi.com/.../auth, classifying
any model 401 as user "session expired" and rendering the misleading
"页面将自动跳转到登录页" copy. (The redirect itself only fires from
/api/v1/auth/* axios paths and SSE-connection 401s, not from this
payload-text path — but the copy alone is the worst kind of false
alarm.)
Add a new ChatErrorCategory provider_auth_error and split the
pattern matching: narrow auth_expired (HTTP 401 / 登录已过期 /
session expired / 凭证失效) is matched FIRST, then the broad
401-ish pattern routes to provider_auth_error. BACKEND_ERROR_TYPE_MAP
for AUTH_ERROR is also remapped, since structured backend payloads
currently always come from LLM providers — never from our own
/api/v1/auth path.
Tests
- NodeStreamingChatHelperFailoverTest (5 cases): primary 401 →
fallback succeeds; chain skips auth-failing fallback to next healthy
one; whole-chain failure surfaces last AUTH_ERROR (no silent drop);
BILLING regression unchanged; primary-success path does not touch
chain
- Browser preview verified: new i18n keys resolve in en-US, classifier
correctly routes "[错误] 401 from kimi.com" → provider_auth_error
while "[错误] HTTP 401 from /api/v1/auth/ping" stays auth_expired
- 186 tests pass (was 181 + 5 new); vue-tsc clean
Do-not-touch list: handleAuthFailure() in useStream/api/index.ts (real
session-expiry path) is unmodified — only the misclassification
upstream is fixed. auth_expired i18n copy is unchanged.
Track the primary model health, not just fallback entries
- NodeStreamingChatHelper accepts primaryProviderId via a new 5-arg
constructor; AgentGraphBuilder passes ModelConfigEntity.getProvider()
- Before the 5-retry primary loop, check
healthTracker.isInCooldown(primaryProviderId): if true, log + broadcast
"主模型暂时不可用(冷却中),直接尝试备选模型..." and short-circuit
straight to the fallback chain. Prevents a degraded primary from
burning 30+ seconds of backoff on every conversation turn.
- recordPrimary(success/failure) now fires on every primary verdict —
AUTH, BILLING, MODEL_NOT_FOUND, EMPTY_RESPONSE, generic UNKNOWN, and
the explicit success path. Three consecutive failures push the
primary provider into cooldown automatically.
- Legacy 1/2/3-arg constructors leave primaryProviderId null; tracking
silently disables for them so existing tests/wiring keep working.
Split BILLING and MODEL_NOT_FOUND out of CLIENT_ERROR / AUTH_ERROR
- BILLING (HTTP 402, "insufficient_quota", "credit balance is too low",
"billing_hard_limit_reached", "quota exceeded"): payment failure on
primary does not kill the call — a different provider may have credits.
Skips same-model retries and heads to fallback chain.
- MODEL_NOT_FOUND (HTTP 404, "Model not exist", "model_not_found",
DashScope "[InvalidParameter] url error"): unknown model id will not
start working on retry. Was previously misclassified as CLIENT_ERROR
and terminated the whole call; now routes to fallback so a different
provider can attempt with its default model.
- classifyError ordering matters: BILLING / MODEL_NOT_FOUND are matched
BEFORE the generic 400 / Bad Request branch, otherwise they would be
swallowed by CLIENT_ERROR.
Tests
- ErrorClassificationTest: 11 tests, covers multi-vendor error phrasing
for both new types + regression checks that 401 / 429 / 400 still
classify as before
- NodeStreamingChatHelperFallbackChainTest: +2 tests verifying
primaryProviderId persistence on the new constructor and null on
legacy ones
- 181 tests pass (was 168 + 13 new)