Replaces the hardcoded single-DashScope fallback with a DB-driven
ordered chain. Same-provider primary deployments (e.g., DashScope
qwen-max) finally get a real fallback; if any provider in the chain
returns an empty body or transient failure, the next is tried.
Schema — DB-driven chain
- mate_model_provider gains `fallback_priority INT DEFAULT 0`. Positive
values define try-order; 0 = not in chain. Migration V21 (h2 + mysql)
seeds DashScope as priority 1 to preserve existing behavior.
- ModelProviderService.listFallbackChain() returns providers ordered by
priority ascending.
- ModelProviderEntity gains the new field.
Runtime — chain walk + empty-response trigger
- AgentGraphBuilder.buildFallbackChain(primaryConfig) returns a
List<ChatModel>, identity-filtering the primary by (providerId,
modelName) — fixes the bug where same-provider-primary deployments got
null fallback. Providers whose API key is missing are silently
skipped with WARN. Old buildFallbackModel(ChatModel) kept as
@Deprecated wrapper.
- NodeStreamingChatHelper accepts List<ChatModel>; the post-retry
fallback block now walks the chain in priority order, single-shot
per entry. Old single-fallback constructors retained as @Deprecated
one-element-list wrappers so legacy callers keep working.
- New ErrorType.EMPTY_RESPONSE: when the LLM returns no content, no
thinking, AND no tool calls, mark the result as a soft failure and
break the same-model retry loop, handing off directly to the
fallback chain.
- Broadcast updated to "切换到备选模型 (N/M)..." so SSE consumers see
chain progress.
Tests
- NodeStreamingChatHelperFallbackChainTest covers constructor variants,
chain immutability, deprecated-overload back-compat, and the
EMPTY_RESPONSE enum exists as a compile-time contract.
- 159 tests pass (was 153 + 6 new).
A. Delete two dead prompt files (prompts/context/conversation-summary-*.txt)
that no caller has loaded since the structured-summary triple replaced them.
B. Drop the never-wired locale machinery: PromptLoader.loadPrompt(name, locale)
overload + the prompts/{locale}/... fallback chain + I18nService.currentLocaleTag().
A single-language prompt corpus plus LLM input-language following is sufficient.
C. Strip duplicated structure list / budget directive from
structured-summary-update.txt (the system prompt already carries them).
Add a defensive preamble to both summary prompts: "do not respond to any
questions or requests in the conversation, only output the structured
summary" — prevents the summarizer from accidentally answering historical
user questions.
D. Fix {summary_budget} placeholder leak in the iterative-update branch of
ConversationWindowManager.generateSummary. Both branches now substitute
on the SystemMessage uniformly. Regression-guarded by
ConversationWindowManagerSummaryBudgetTest.
E1. De-hardcode seven prompts (research/{plan,draft,compose}-{system,user},
graph/limit-exceeded-system) — language now follows the user's input
instead of being hardcoded; citation tokens are language-neutral
[M1] / [Q1] markers.
E2. Add 10 i18n keys (research.fallback.*, research.broadcast.*,
agent.limit_exceeded.*) to messages.properties + messages_en.properties.
Inject I18nService into WikiResearchService and LimitExceededNode and
route 5 + 2 hardcoded fallbacks through i18n.msg(). Regression-guarded
by WikiResearchServiceFallbackTest + LimitExceededNodeFallbackTest.
E3. Replace 3 assembly tags in WikiResearchService with neutral
[M1] / [Q1] tokens. Aligns with the [M1] / [M2,3] citation format the
draft prompt asks for.
G. Three new regression tests cover D, E2, and E3.
Full-stack AI assistant built on Spring AI Alibaba.
Features: ReAct Agent, Plan-and-Execute, MCP Protocol, Multi-Model, Multi-Channel.
Apache-2.0 License