Replaces the prior ThreadLocal context plumbing with explicit Spring AI
ToolContext threading carried by an immutable ChatOrigin value object,
so a cron created from inside WeChat (or any IM channel) delivers its
results back to the originating channel.
Architecture
- ChatOrigin / ChannelTarget value objects + per-entry-point factories
(ChannelChatOriginFactory in vip.mate.channel, CronChatOriginFactory
in vip.mate.cron — symmetric, no cyclic deps).
- LocaleAwareToolCallback now forwards call(String, ToolContext) and
getToolMetadata so the decorator chain cannot silently drop the origin.
- AgentService 6-method overhaul + ChatOriginHolder bridge into
StateGraph buildInitialState which writes CHAT_ORIGIN; ActionNode +
StepExecutionNode forward it to ToolExecutionExecutor.
- ToolExecutionExecutor builds ToolContext per call; 8/8 tools migrated
(CronJobTool, WorkspacePathGuard, Video/Image/Browser/ReadFile/Music,
DelegateAgentTool with parent-origin inheritance).
- CronJobRunner + CronJobLifecycleService 3-segment REQUIRES_NEW model
(T1 startRun / no-tx runAgent / T2 finishRunAndPublish); ArchUnit
pins CronJobRunner as @Transactional-free.
- CronResultDelivery Strategy + AbstractCronResultDelivery Template
with SQL CAS idempotency on mate_cron_job_run.delivery_status —
replaces the prior process-local Caffeine TTL, cluster-safe.
- CronJobCompletedEvent + @Async @TransactionalEventListener(AFTER_COMMIT);
cronDeliveryExecutor (core=2, max=4, queue=1000, AbortPolicy + audit).
- CronRunStaleCleanup @Scheduled(5min) sweeps PENDING-15min and
status='running'-30min in one query each.
- CronJobRunner.wrapWithDeliveryGuard prepends a system note for
channel-bound crons to suppress hallucinated 'install CLI to send
WeChat' suggestions.
- ApprovalWorkflowService Memento: persist ChatOrigin snapshot on
create, restore on replay so cross-restart approvals keep channel
binding; ChannelMessageRouter + ChatController web-replay both prefer
the Memento and fall back to fresh-build.
- ChannelManager.sendToChannel 4-arg DeliveryOptions overload;
ChannelAdapter#proactiveSend default 4-arg pass-through; Slack
overrides for thread_ts and Telegram overrides for message_thread_id.
- CronJobs UI: read-only 'last delivery' badge driven by
CronJobMapper.selectListWithDeliveryStatus subquery.
Schema migrations V57/V58/V59 (V56 was already taken by an unrelated
provider migration — Flyway processes versions in order regardless of
gaps):
- V57: mate_cron_job_run delivery_status / target / error + composite
index (delivery_status, started_at) covering the cleanup sweep.
- V58: mate_cron_job channel_id (indexed) + delivery_config TEXT (JSON
via MyBatis Plus JacksonTypeHandler).
- V59: mate_tool_approval chat_origin TEXT (Memento).
All idempotent in both H2 (IF NOT EXISTS) and MySQL (INFORMATION_SCHEMA
guard + PREPARE).
ArchUnit guards (test scope, archunit-junit5 1.3.0):
- every concrete vip.mate.* ToolCallback must override
call(String, ToolContext) — pins the decorator-forward fix.
- CronJobRunner must NOT carry @Transactional on the class or any
method — pins the 3-segment lifecycle rule.
Tests: 32 new unit tests + 21 regression tests in touched areas, all
53 green:
- ChatOriginTest (6) — value-object invariants + JSON round-trip.
- LocaleAwareToolCallbackToolContextTest (2) — decorator forward.
- DeliveryConfigTest (4) — Jackson round-trip + forward-compat.
- ToolCallbackToolContextForwardArchTest (2) — both ArchUnit guards.
- CronJobRunnerDeliveryGuardTest (3) — channel-cron prefix injection.
- AbstractCronResultDeliveryTest (4) — claim CAS + concurrent CAS.
- ChannelCronResultDeliveryTest (6) — supports / doDeliver / errors.
- ApprovalReplayContinuityTest (5) — Memento round-trip + corrupt
payload fallback + unknown-field tolerance.
Refs: #25, #16
Adds explicit user-intent gating to the provider catalog. Fresh installs
get an empty dropdown by default — built-in cloud + local providers
(OpenAI, Anthropic, Ollama, LM Studio, MLX, llama.cpp, etc.) live in a
new 'Add Provider' drawer until the user opts them in. Existing installs
upgrade conservatively: V55 promotes any provider with evidence of use
(real api_key, OAuth token, recent chat usage, or current default model).
Backend
- V55 migration (H2 + MySQL): adds enabled BOOLEAN DEFAULT FALSE on
mate_model_provider, plus 4 promote-to-true UPDATE rules. Also
CREATE INDEX idx_message_runtime_provider_time so the 30-day usage
lookup doesn't full-scan mate_message on heavy users.
- ModelProviderEntity, ProviderInfoDTO: enabled field.
- ModelProviderService:
* listProviders() now filters WHERE enabled = TRUE — chat path,
ModelSelector, Settings/Models main grid see only opted-in rows.
* listCatalog() new — full catalog (enabled + disabled) for the drawer.
* setEnabled(id, enabled) flips the flag, publishes
ModelConfigChangedEvent (re-probe via the existing listener), and
on disable auto-promotes a replacement default model when the
disabled provider owned the current default. Returns EnableResult
so the frontend can fire a toast.
* createCustomProvider sets enabled=true (user just made the row).
- ProviderInitProbe.listConfiguredProviders also filters enabled=true —
no point probing rows the user can't see.
- ModelConfigController: GET /catalog, POST /{id}/enable, POST /{id}/disable.
- Plugin-registered ChatModels are unaffected — they live in
pluginChatModels (in-memory map), don't go through DB listProviders,
so the enabled filter doesn't strand them.
Frontend
- New types: ProviderInfo.enabled, EnableResult.
- New API: catalog / enableProvider / disableProvider.
- New composable useProviderEnablement: catalog ref, drawerOpen,
togglingId, loadCatalog, openDrawer / closeDrawer, enableProvider,
disableProvider (fires defaultSwitchedToast on auto-switch).
- AddProviderDrawer.vue: lazy-loaded, reuses DoctorDrawer's Teleport +
overlay + slide-in panel pattern. Two groups (cloud / local),
unenabled rows surface to the top of each group, enabled rows show
an 'Enabled' badge instead of a button. Mobile: full-screen sheet
that slides up from below.
- ProviderCard: new 'Disable' button with soft-danger styling on
enabled providers — soft-hide that keeps the config; user can
re-enable from the drawer.
- Settings/Models index.vue:
* Two top CTAs: 'Enable Provider' (drawer) and 'Custom' (existing
custom-create modal) — distinct workflows, both surfaced.
* Empty state with prominent 'Enable Provider' CTA when zero
enabled providers — paired with onMounted auto-open of the
drawer (sessionStorage guard so closing it doesn't bring it
back on the next route visit in the same session).
* Deep-link: ?addProvider=1 query forces the drawer open and
strips itself after, so a back/forward doesn't re-fire the open.
- ModelSelector: when groups.length === 0 and not searching, show
'No providers configured -> Configure' CTA linking to
/settings/models?addProvider=1 — the natural flow when a fresh
user opens chat before configuring anything.
- i18n: 13 new keys per locale (zh-CN + en-US) plus common.close.
Migration safety
- Conservative default policy: only rows with concrete evidence of
use are auto-enabled; everything else stays hidden. Upgrade users
may notice unused built-ins disappearing from their dropdown —
that's the intended cleanup.
- mate_message index added so the 30-day usage rule doesn't full-scan
on large installations; FlywayRepairConfig handles redeploy idempotency.
Tests
- ModelProviderServiceEnableTest covers all 7 enable/disable branches:
flag flip + event publish, no-op on already-{enabled,disabled},
default-switch when disabled provider owned current default,
no-switch when default belongs elsewhere, no-replacement returns
unchanged, getDefaultModel exception path, candidates with no
models are skipped.
- ProviderInitProbeTest: helper provider() now sets enabled=true so
the new probe filter doesn't strand existing fixtures.
- vip.mate.llm.** suite: 125 tests green. vue-tsc 0 errors. Browser
page renders with both new buttons + drawer.
Reshapes the Settings/Models frontend to match the channel-module split
convention (commit 22894ac4 'perf(channels): split Channels.vue...'),
zero behavior change. Paves the way for a follow-up that adds an enabled
column + AddProviderDrawer without bloating useProviders back to monolith.
Frontend split
- useProviders.ts goes from 615-line monolith to a 48-line facade that
composes five single-responsibility slices:
* useProviderList — providers / activeModels / currentProvider,
loaders, status pill, icons
* useProviderForm — create/edit modal + form, save/delete
* useProviderDiscovery — manage-models modal, discovery, connection
and per-model tests
* useProviderOAuth — openai-chatgpt + claude-code OAuth flows
* useProviderPool — manual reprobe (most pool surface inlined to
ProviderInfo.liveness in the prior liveness change)
Cross-composable refs flow via dep-injection arguments — no module-
level state, no circular deps. Each composable stays independently
testable.
- Pure helpers extracted to src/utils:
* safeJson.ts — strict JSON-object parser
* modelProtocol.ts — protocol <-> ChatModel class translation
- Modals (ProviderConfigModal, ManageModelsModal) loaded via
defineAsyncComponent so the route's first paint doesn't drag along
~30KB of form/auth UI.
- el-skeleton placeholder during the initial Promise.all so the page
paints something instead of blank-then-pop.
Layout regression fix
- MainLayout's <keep-alive> slot used :key='workspaceRouteKey' (a
workspace-scoped string), shared between two <component v-if> blocks.
Adding a second keepAlive route would have caused two components to
mount side by side, because Vue saw identical keys and patched in
place across the v-if boundary. Switched the key to
${workspaceRouteKey}:${route.path} so different routes get distinct
vnode identities while workspace switching still busts the cache.
Discovered while implementing the split — the multi-line HTML
comment also had to live OUTSIDE <keep-alive>, since KeepAlive
treats comments as children and rejects 'more than one'.
Embedding section title fix (drive-by)
- EmbeddingModelsSection.vue's scoped style didn't redeclare
.group-title's flex layout, so the icon stacked above the title
text instead of sitting inline. Added the missing flex rules
locally — now matches the local-models / cloud-models group headers.
Verification
- vue-tsc 0 errors.
- Browser end-to-end: 27 cards render correctly, modals open via lazy
load, /channels <-> /settings/models switch four times in a row with
exactly one page title visible at each step (no stacking).
Phase 1 of the model-module refactor: combine pool / cooldown / probe-
completion signals into a single Liveness state surfaced through the
provider DTO, so the dropdown stops listing providers that are provably
unreachable. Zero schema change; one PR backend + frontend.
Backend
- Liveness enum with five mutually-exclusive states: LIVE, COOLDOWN,
REMOVED, UNPROBED, UNCONFIGURED. Computed in ModelProviderService
from AvailableProviderPool / ProviderHealthTracker / ProviderInitProbe
snapshots batched once per listProviders() call.
- ProviderInitProbe.hasBeenProbed exposes a monotonic Set so the UI
can distinguish 'still booting' from 'probed and removed' — without
it the startup window flashes false REMOVED states.
- ProviderInfoDTO gains liveness + unavailableReason +
cooldownRemainingMs + lastProbedAtMs. The legacy 'available' boolean
stays but is now derived from liveness == LIVE so the chat fallback
walker and the dropdown agree about what's usable.
- ProviderInitProbe injected into ModelProviderService via
ObjectProvider to break the startup cycle (probe already depends on
the service).
Frontend
- ProviderInfo type extended with liveness + the three detail fields.
- ModelSelector filters UNCONFIGURED + REMOVED out of the dropdown,
shows COOLDOWN / UNPROBED with a status dot and dimmed rows that the
user can still click to override.
- ProviderCard renders a five-state badge driven by liveness instead
of the old configured + pool-entry combo. Reprobe button now keys
off liveness in {REMOVED, COOLDOWN}.
- useProviders drops loadProviderPool / providerPool — pool data ships
inline on each ProviderInfo, saves a round trip per page load and
keeps a single source of truth.
- i18n: 8 new keys across zh-CN and en-US for liveness labels and the
cooldown countdown tooltips.
Bonus fix (discovered during verification): AgentGraphBuilder.buildOpenAiApi
hard-required a usable API key on every OpenAI-compat provider, ignoring
the per-provider requireApiKey flag. That bug stranded keyless local
runtimes (LM Studio / MLX / llama.cpp) the moment a user actually
launched them; Ollama only worked by accident because its seed row
carries a placeholder string in api_key. keyRequired now honors
requireApiKey, and Spring AI's NoopApiKey is used when no key is needed
so the Authorization header is omitted entirely.
Test
- ModelProviderServiceLivenessTest covers all five Liveness states +
the probe-bean-absent fallback branch.
- vip.mate.llm.** suite (118 tests) green; vue-tsc clean.
- End-to-end browser sanity: 27 raw providers reduce to 6 LIVE groups
in the chat dropdown; LM Studio / MLX / llama.cpp render REMOVED red
badges with reprobe buttons; cloud providers without keys show
UNCONFIGURED.
- Extract create/edit modal into ChannelEditModal.vue (defineAsyncComponent),
shrinking Channels.vue from 1438 to 370 lines and dropping ~30KB from the
initial route chunk.
- Move side-effect logic into composables: useWeixinQrcodePoll (QR + 2s status
poll, auto-cleanup) and useWecomBotAuth (lazy SDK script with module-level
promise dedupe). Pure config-JSON helpers move to utils/channelConfigJson.ts.
- Switch i18n locales from static imports to dynamic import keyed by current
locale; applyLocale becomes async to avoid first-render flicker.
- /channels route opts into keep-alive (meta.keepAlive=true). Channels.vue
pauses status polling in onDeactivated and resumes in onActivated, with an
isActive guard to prevent late-resolving timers from leaking after navigation.
- Initial load goes from serial 3-RTT to Promise.all + 4-card el-skeleton.
When SSE setup fails (e.g. workspace permission denied for shared channel
conversations opened from the web console), the failed turn is never
persisted on the backend. Two issues made the failure invisible to the user:
- The fallback errorInfo dropped data.message, so the inline retry card fell
back to the generic "请求过程中遇到了意外问题" template instead of the
actual reason. Carry rawMessage through, and lower the MessageBubble
display threshold from >8 to >3 chars so short-but-informative messages
(7-char Chinese / "Forbidden") aren't filtered out.
- The status-poll loop in useChat overwrote the local-only failed turn
with the server's "no message" view, erasing the inline retry card.
Skip the merge for turns that exist only locally and are in error state,
so the user can still see the failure and retry.
- Reconcile approval status atomically: DB row, message metadata, in-memory store
- Approve and deny both flip the tool-call card + timeline segment to a terminal
state on the gate message — no more orange spinner stuck after a decision
- Frontend hydrate matches by pendingId and reverse-converges to expired so a
refresh after server-side timeout / consume clears the banner without restart
- Stop sweep, GC timeout, and JVM restart all close the loop with consistent
state
- Remove the dead REST /approve endpoint + matching frontend client export so
there is only one resolve path to maintain
Three small but high-impact fixes that all surfaced together while
verifying the long-form generation flow.
1. ChatConsole onBeforeUnmount no longer kills the backend turn.
Previously, switching tabs / route navigation / any cause that
unmounted the chat view called stopChatGeneration(), which POSTs
/chat/{cid}/stop and aborts the in-flight LLM call. The user
reported a turn dying mid-generation just from switching pages.
Replaced with resetForNewConversation() — front-end SSE disconnect
only, no /stop. Backend keeps running; pollActivity / status probe
reconnects on return. Aligns with the existing comment in
selectConversation: "let A's backend agent run continue running."
2. Agent max_iterations raised 25 → 100 with a hard ceiling.
The previous 25-step ceiling caused LimitExceededNode to fire on
substantive multi-tool tasks (document generation + image conversion
+ retry loops). 100 matches QwenPaw's _MAX_MAX_ITERATIONS upper
bound. New plumbing:
- BaseAgent.MAX_ITERATIONS_HARD_CEILING = 100 public constant
- BaseAgent default field 25 → 100 (Java-side fallback)
- AgentGraphBuilder clamps any per-agent DB override to the
ceiling at runtime; if the row holds 200, runtime sees 100 and
a WARN is logged with the original value.
- V47 migration (h2 + mysql) idempotently bumps the three default
seeded agents (1000000001, 1000000002, 1000000003) only if they
still hold the old defaults (25 / 20). User-customized values
are not touched.
- data-en/zh/-mysql-en/-mysql-zh seed files updated to 100 for
fresh installs.
3. DocxRenderTool tells the LLM not to prepend a host to the URL.
DeepSeek and Claude have both been observed wrapping the
/api/v1/files/generated/{id} relative path returned by renderDocx
into an absolute URL with a hallucinated domain (e.g.
https://ai-tools-system.com/...), breaking the download link in
the rendered chat bubble. The tool's return string now appends an
explicit "must use the relative path verbatim, do not add any
https:// or http:// prefix" instruction, which Claude and
DeepSeek both honor.
- Drag-over highlight (orange border + shadow + arrow icon) on the upload
zone, using dragCounter to prevent flicker over nested children
- Optimistic list items appear immediately on drop/select with UPLOADING
badge and progress bar, before the HTTP request completes
- Wire axios onUploadProgress through api.uploadRaw → store.uploadRawFile
so the progress bar tracks real byte transfer (0–100%), with an
indeterminate shimmer until the first tick
- try/catch around every upload: on failure, placeholder flips to an error
state with ElMessage.error toast and a × dismiss button
- Upload all dropped/selected files concurrently via Promise.all
- i18n keys added (zh-CN + en-US): dropToUpload, uploading, uploadFailed,
status.uploading, progress.uploading
Three bugs fixed:
1. Badge stays "processing" after job completes: pollJobs() never called
fetchRawMaterials() when a job reached terminal status, so
raw.processingStatus stayed processing in the store. Fix: detect
terminal job status in pollJobs, trigger fetchRawMaterials to sync.
2. "Completed" dot pulses instead of solid: dotClass() treated completed
the same as in-progress (target === cur → active). Fix: add
isTerminal computed (includes completed), return done for all
dots at or before the terminal position — no pulse animation.
3. Stage label stays orange at terminal: same cause — active class
applied regardless of terminal state. Fix: use done class for
terminal labels (green instead of orange).
Also: v-if on JobStageBar now shows for terminal status jobs (not just
stage !== queued), so completed/failed stage bars remain visible.
Root cause: processRawMaterial() created a job record at queued stage
but never called jobService.transition() during processing. The job row
stayed at queued forever, so the stage bar never advanced.
Backend (WikiProcessingService):
- Transition job to ROUTING immediately after creation
- Transition to PHASE_A_RUNNING before chunk processing begins
- Transition to COMPLETED/PARTIAL/FAILED at the end based on finalStatus
- Transition to FAILED in the catch block on unhandled exceptions
Backend (WikiProcessingJobService.transition):
- Handle FAILED, PARTIAL terminal stages (set finishedAt + status)
- Handle non-terminal intermediate stages (set status to running)
Frontend (JobStageBar.vue):
- Add stageMapping for backend stages not shown as dots: phase_a_done →
phase_b_running, failed/partial/cancelled → completed position
- Guard stageIndex() against -1 (unknown stages default to all-pending)
- Terminal failure states show red failed dot instead of pulsing active
Three root causes for broken progress:
1. JobStageBar was shown whenever a job record existed (even at queued
stage), hiding the working SSE-driven progress bar. Fix: only show
JobStageBar when job.stage !== queued.
2. SSE connection only opened when hasProcessing was true (status ===
processing), but reprocess sets status to pending first. Fix:
include pending in the hasProcessing check.
3. After reprocess, if processing finished before SSE connected, the
status badge stayed on pending forever. Fix: immediately set local
status to processing after reprocess API call, clear stale job
entries, and add delayed re-fetches (5s/15s) as safety net.
Also: clear rawJobs entries on raw.completed/raw.failed SSE events to
prevent stale JobStageBar from lingering after processing ends.
UI — Failover priority editor
- ProviderConfigRequest + ProviderInfoDTO carry fallbackPriority
- ModelProviderService.updateProviderConfig persists it (null = unchanged);
toProviderInfo exposes the current value to the UI (defaults to 0)
- ProviderConfigModal advanced panel exposes a number input with hint
- ProviderCard shows a "Fallback #N" badge for chain members so the
priority order is visible at a glance without opening the modal
- 5 new i18n keys (zh + en) — verified to resolve at runtime via i18n.global.t
Backend — Per-provider health tracker
- ProviderHealthTracker: ConcurrentHashMap-backed counters; N consecutive
failures (default 3) push the provider into a cooldown window (default
5 min) during which the chain walker skips it. Success resets both
counter and cooldown atomically. Lazy expiry on lookup so dead entries
do not accumulate.
- ProviderHealthProperties exposed under mateclaw.llm.failover.health.*
with sane production defaults
- New FallbackEntry record (providerId + ChatModel) replaces raw
List<ChatModel> in the chain so the walker can correlate cooldown
state to entries; AgentGraphBuilder.buildFallbackChain returns the
new type
- NodeStreamingChatHelper takes the tracker through a new 4-arg
constructor and consults it before each fallback call; records
success/failure on each chain attempt. Legacy 2/3-arg constructors
preserved as @Deprecated wrappers (synthetic providerId means no
health tracking on the legacy path — that path is opt-out anyway)
Tests
- ProviderHealthTrackerTest (9 tests): below/at threshold, success
reset, cooldown expiry (via reflection on the min-clamp setter),
disabled-tracker no-op, null-providerId safety, per-provider
isolation, snapshot output
- NodeStreamingChatHelperFallbackChainTest updated to FallbackEntry
field type — verifies providerId + ChatModel survive the chain
- 168 tests pass (was 159 + 9 new)
Verification
- mvn test green; vue-tsc clean; live UI confirms i18n resolution
- ChannelMessageRouter: include assistantMessageId in message_complete
and done broadcasts so ChatConsole observers can reconcile the
streaming placeholder to the persisted DB row by id instead of
falling back to the FIFO 'claim' heuristic (which occasionally
dropped the assistant bubble on external channel conversations).
Capture the id from ConversationService.saveMessage in both the
sync agentService.chat path and the streaming processWithStreaming
path; switch to HashMap since Map.of rejects null values when save
is skipped (e.g. under approval).
- ChatConsole: add a 'running' indicator on the sidebar so users can
tell which conversations have an in-flight agent run. Pulsing amber
dot on the channel icon (both expanded and collapsed modes) plus a
'生成中…' / 'Generating…' pill in expanded mode.
- ChatConsole: don't cancel the previous conversation's streaming run
when switching conversations — let it keep running in the background
and reconcile when the user comes back.
Full-stack AI assistant built on Spring AI Alibaba.
Features: ReAct Agent, Plan-and-Execute, MCP Protocol, Multi-Model, Multi-Channel.
Apache-2.0 License