Group-chat reply slot fallback:
- The platform blocks proactive sends in group chats; outbound paths
(cron summaries, async-task completions, generated image/music/3D
delivery, TTS audio) silently failed because they fell through to
the proactive-send command. New bounded LRU maps each group chat
to its most recent inbound frame id; the dispatcher prefers that
reply slot and falls through to proactive only for single chats.
- Centralised text and media dispatch through a single helper so the
group rule never has to be re-implemented per outbound path.
Upload size pre-check + auto-downgrade:
- Without client-side limits, oversized uploads streamed for ~1 minute
before the server rejected at the finish step — users saw nothing
arrive in their chat. The new decision layer mirrors the platform's
hard limits and produces three outcomes: rejected with a friendly
reason, downgraded to a generic file delivery with an inline note,
or pass-through unchanged.
- Files over 20MB reject. Images / videos over their 10MB limit
downgrade to file. Voice content that isn't AMR or exceeds 2MB
downgrades. AMR voice within 2MB stays native.
appmsg inbound parsing:
- Forwarded complex messages (document transfers, article links,
miniprogram cards) used to fall into the inbound switch's default
branch and silently drop. The new branch flattens four sub-types
into a text marker the agent reads plus any media that needs to
reach downstream tools — document forwards reuse the same magic-
byte sniff and per-conversation upload layout as native file
inbound, so extension recovery and chat-uploads serving work
identically.
- Article links produce "[链接] title\ndescription\nurl" so the agent
can summarize without round-tripping. Miniprograms surface their
title. Unknown sub-types still emit a generic marker so the agent
is never blind.
WeCom quoted-message context:
- Parse the body.quote field that arrives alongside any inbound message
(text / image / voice / file / mixed sub-types). When a user long-
presses a previous bot bubble and types a follow-up like "解释一下",
the agent now sees both the user's new text and the referenced
content as proper context — replies stay on topic instead of
guessing what was being explained.
- Quoted images / files are downloaded through the same pipeline as
inbound new media (magic-byte sniff, ZIP container peek for
DOCX / XLSX / PPTX recovery, chat-uploads layout) so the vision
sidecar and document tools can actually analyse what was quoted.
- Reading order in the assembled prompt: "[引用消息: ...]\n<user text>"
first, then quoted media parts, then the user's own current-message
media. Mixed quotes flatten into a space-joined summary.
Multimodal sidecar settings preservation:
- The bulk settings PUT used to unconditionally overwrite the vision /
video sidecar model ids — null in a partial payload became "" in
the DB, silently wiping the configured sidecar every time a user
saved an unrelated settings page (System / Music / Image / etc.).
Symptom: "I picked a vision model, saved a different settings tab,
now the bot can't see images anymore."
- Bulk save now guards both keys with non-null checks, matching the
pattern used for music / 3D / image / video / tts / stt blocks.
- A dedicated /settings/sidecar endpoint always writes both keys, so
the sidecar UI can still explicitly clear via null without leaking
the write-on-null semantics into every other settings save.
- Frontend sidecar card switches to the dedicated endpoint; other
settings pages keep their existing partial-payload behaviour.
Inbound (WeCom):
- Save uploaded media under data/chat-uploads/{conversationId}/ with full
fileName/path/fileUrl/storedName/fileSize on the content part. Web mirrors
of an IM conversation now show real thumbnails instead of "未命名".
- Magic-byte sniff (PDF / PNG / JPEG / GIF / Office / ODF / archives /
audio / video) recovers a real extension when the platform omits filename
for forwarded files — no more PDFs labelled "file.bin".
- ZIP container peek distinguishes DOCX / XLSX / PPTX / VSDX / ODT / ODS /
ODP / EPUB / JAR from a plain zip via discriminator paths and the OASIS
mimetype entry.
Outbound (WeCom):
- Chunk upload field name corrected so server-side actually stores the
bytes — file messages used to arrive with correct filename/size but
empty content, breaking every PDF / DOCX / PPTX recipient.
- Scan agent text for served-file URLs in both the text-reply and
content-parts paths; fetch bytes from the in-memory generated-file
cache and dispatch through the native chunk upload + media message
protocol so users receive a tappable file card instead of an
unopenable markdown link. Cache miss surfaces a clear retry hint.
Async tool result forwarding:
- New AsyncTaskMediaDispatcher routes generation completions (image,
video, music, 3D model) to whichever IM channel the conversation is
bound to via ChannelSessionStore + ChannelManager. Web / webchat
conversations are intentionally skipped — their SSE stream already
renders the result.
- Wired into all four generation services so IM users actually receive
generated media as native attachments. Each part now carries an
absolute disk path so adapters read bytes locally instead of round-
tripping through an authenticated served URL.
Slack native file upload:
- SlackChannelAdapter overrides the content-parts dispatch. Image /
audio / video / file / model3d parts ride filesUploadV2 so users see a
file card with preview thumbnail bound to the same thread as the
originating message. Text parts continue through chat.postMessage.
- Resolves bytes from the part's local path, falls back to an HTTP fetch
of fully-qualified URLs.
IM approval hint visibility:
- IM-driven approve / deny / auto-cancel / replay-error hints now go
through saveMessage + tracker broadcast in addition to the channel
adapter, so a Web mirror viewing the same conversationId sees the
resolution. Previously hints reached only the IM channel; the Web
admin console had no record of the outcome.
Adaptive paste-merge debounce:
- WeCom and other IM clients silently split long pasted prompts into
fragments that arrive 0.5-2 seconds apart, missing the existing 500ms
merge window. The agent then saw torn context and emitted multiple
conflicting replies.
- When the merged buffer crosses a content-length threshold, extend the
debounce window so subsequent fragments arrive in time. Default
500ms unchanged for normal short messages.
User-reported field issues + a deeper code audit revealed multiple
overlapping bugs in the prior cron-channel delivery change. This fixes
all six.
#1 — Concurrency race on ToolExecutionExecutor (root cause of 'sometimes
succeeds, sometimes fails' tool calls). The volatile instance fields
currentRequesterId / currentWorkspaceBasePath / currentChatOrigin
were shared by every conversation routed through the same per-agent
executor; one user mid-build-loop while another's execute()
overwrote the field would cross-contaminate the captured values into
PreparedToolCall. Fix: kill the instance fields, thread
origin/requester/workspace as method params straight into
PreparedToolCall snapshot. Comment pins the rule so it cannot regress.
#2 — CHAT_ORIGIN missing from KeyStrategyFactory (latent timebomb,
masked by spring-ai-alibaba-graph-core's non-filtering builder path).
Without an addStrategy registration, multi-node state merges in long
ReAct / Plan-Execute loops drop the key, ActionNode reads
ChatOrigin.EMPTY, and the cron persists with channel_id=NULL. Also
caught 4 more keys that were latently unregistered:
WORKSPACE_BASE_PATH, STOP_REQUESTED, RETURN_DIRECT_TRIGGERED,
DIRECT_TOOL_OUTPUTS. All five now registered in both ReAct and
Plan-Execute factories.
#3 — CronJobs UI didn't surface channel binding. CronJobDTO carried
channelId / deliveryConfig but the list page never rendered them.
Added: (a) 'channel' column on list page, (b) channel + targetId
rows in the detail modal, (c) backend batch-loads channel names via
ChannelMapper.selectBatchIds so the column shows the human-readable
name, (d) i18n keys (zh + en), (e) channelName field on TS CronJob
type.
#4a — DingTalk targetId expiry. ChannelChatOriginFactory.resolveTargetId
used to prefer ChannelMessage.replyToken which for DingTalk encodes
a sessionWebhook URL that expires ~90 minutes after the inbound
message. Cron persisted with that webhook then dies with 401/403 and
marks NOT_DELIVERED forever. Fix: prefer the stable chatId, fall
back to senderId — both work indefinitely via DingTalk's Robot API.
#4b — Scheduler pool exhaustion under long LLM. CronJobService's
ThreadPoolTaskScheduler ran with poolSize=4 AND the LLM call lived
on the scheduler thread. Four concurrent crons saturated the pool
and the 5th silently missed its tick. Fix: keep scheduler tiny (it
just fires triggers) and offload runAgent to a dedicated
virtual-thread executor (cron-execute-* threads). LLM workload is
I/O-bound — virtual threads scale to thousands at trivial cost.
#5 — Minor latent bugs:
- AbstractCronResultDelivery.claimRun used .in(... 'NONE','PENDING',null),
but SQL IN never matches NULL. Rewrote as IS NULL OR IN
(NONE,PENDING) so legacy pre-V57 rows can still claim.
- CronDeliveryListener.onCompletedRaw was an empty @EventListener
with a wrong-headed comment about test fallbackExecution. Removed.
- CronJobTool.resolveAgentId silently returned 1L when origin
lacked an agentId — would silently bind to whatever agent #1
happens to be. Replaced with explicit error so wiring bugs surface
immediately instead of producing scheduled-but-never-runs crons.
State-key registration guard. New StateKeyRegistrationCoverageTest
scans MateClawStateKeys via reflection and parses
AgentGraphBuilder.java to extract every
.addStrategy(MateClawStateKeys.X, ...). Asserts every non-_NODE
constant appears in at least one factory. Caught the 4 unregistered
keys above on first run; will catch any future 'forgot to register'
regression.
Tests: 33 unit/arch tests + 27 regression in touched areas — all green.
Vue typecheck clean.
Refs: #25, #16
Replaces the prior ThreadLocal context plumbing with explicit Spring AI
ToolContext threading carried by an immutable ChatOrigin value object,
so a cron created from inside WeChat (or any IM channel) delivers its
results back to the originating channel.
Architecture
- ChatOrigin / ChannelTarget value objects + per-entry-point factories
(ChannelChatOriginFactory in vip.mate.channel, CronChatOriginFactory
in vip.mate.cron — symmetric, no cyclic deps).
- LocaleAwareToolCallback now forwards call(String, ToolContext) and
getToolMetadata so the decorator chain cannot silently drop the origin.
- AgentService 6-method overhaul + ChatOriginHolder bridge into
StateGraph buildInitialState which writes CHAT_ORIGIN; ActionNode +
StepExecutionNode forward it to ToolExecutionExecutor.
- ToolExecutionExecutor builds ToolContext per call; 8/8 tools migrated
(CronJobTool, WorkspacePathGuard, Video/Image/Browser/ReadFile/Music,
DelegateAgentTool with parent-origin inheritance).
- CronJobRunner + CronJobLifecycleService 3-segment REQUIRES_NEW model
(T1 startRun / no-tx runAgent / T2 finishRunAndPublish); ArchUnit
pins CronJobRunner as @Transactional-free.
- CronResultDelivery Strategy + AbstractCronResultDelivery Template
with SQL CAS idempotency on mate_cron_job_run.delivery_status —
replaces the prior process-local Caffeine TTL, cluster-safe.
- CronJobCompletedEvent + @Async @TransactionalEventListener(AFTER_COMMIT);
cronDeliveryExecutor (core=2, max=4, queue=1000, AbortPolicy + audit).
- CronRunStaleCleanup @Scheduled(5min) sweeps PENDING-15min and
status='running'-30min in one query each.
- CronJobRunner.wrapWithDeliveryGuard prepends a system note for
channel-bound crons to suppress hallucinated 'install CLI to send
WeChat' suggestions.
- ApprovalWorkflowService Memento: persist ChatOrigin snapshot on
create, restore on replay so cross-restart approvals keep channel
binding; ChannelMessageRouter + ChatController web-replay both prefer
the Memento and fall back to fresh-build.
- ChannelManager.sendToChannel 4-arg DeliveryOptions overload;
ChannelAdapter#proactiveSend default 4-arg pass-through; Slack
overrides for thread_ts and Telegram overrides for message_thread_id.
- CronJobs UI: read-only 'last delivery' badge driven by
CronJobMapper.selectListWithDeliveryStatus subquery.
Schema migrations V57/V58/V59 (V56 was already taken by an unrelated
provider migration — Flyway processes versions in order regardless of
gaps):
- V57: mate_cron_job_run delivery_status / target / error + composite
index (delivery_status, started_at) covering the cleanup sweep.
- V58: mate_cron_job channel_id (indexed) + delivery_config TEXT (JSON
via MyBatis Plus JacksonTypeHandler).
- V59: mate_tool_approval chat_origin TEXT (Memento).
All idempotent in both H2 (IF NOT EXISTS) and MySQL (INFORMATION_SCHEMA
guard + PREPARE).
ArchUnit guards (test scope, archunit-junit5 1.3.0):
- every concrete vip.mate.* ToolCallback must override
call(String, ToolContext) — pins the decorator-forward fix.
- CronJobRunner must NOT carry @Transactional on the class or any
method — pins the 3-segment lifecycle rule.
Tests: 32 new unit tests + 21 regression tests in touched areas, all
53 green:
- ChatOriginTest (6) — value-object invariants + JSON round-trip.
- LocaleAwareToolCallbackToolContextTest (2) — decorator forward.
- DeliveryConfigTest (4) — Jackson round-trip + forward-compat.
- ToolCallbackToolContextForwardArchTest (2) — both ArchUnit guards.
- CronJobRunnerDeliveryGuardTest (3) — channel-cron prefix injection.
- AbstractCronResultDeliveryTest (4) — claim CAS + concurrent CAS.
- ChannelCronResultDeliveryTest (6) — supports / doDeliver / errors.
- ApprovalReplayContinuityTest (5) — Memento round-trip + corrupt
payload fallback + unknown-field tolerance.
Refs: #25, #16
Three knots untangled so an image sent from DingTalk lands in both the
LLM's multimodal prompt and the chat history bubble:
- Prefer MessageContent.downloadCode (universal, used by the new
api.dingtalk.com messageFiles/download) over pictureDownloadCode
(legacy oapi field). Sending the legacy code to the new API got
HTTP 500 unknownError, which was the original 'image not recognized'.
- After fetching bytes, persist to ~/.mateclaw/media/dingtalk/ so vision
can read via FileSystemResource, AND stuff the same bytes into
GeneratedFileCache so the UI gets an /api/v1/files/generated/{id} URL
to render. Without the URL the message bubble showed an empty card.
- Carry filename / contentType / size on the MessageContentPart so the
chat history doesn't fall back to the 'unknown' caption.
Same treatment applied to the richText branch (inline images from the
PC client) and threaded through the Stream SDK path.
Bundles in the prerequisite ChannelManager wiring of GeneratedFileCache
into DingTalkChannelAdapter and the new DingTalkMediaUploader used by
the outbound attachment flow that this work depends on.
Known limit: GeneratedFileCache TTL is 10 min — fresh refreshes work,
but viewing the image after a JVM restart needs a stable on-disk
serving endpoint, which is intentionally out of scope here.
The stream SDK delivers voice messages as ChatbotMessage with msgtype=audio
and the server-side ASR result already filled into MessageContent.recognition
(same shape as WeCom's voice.content). The adapter's handleStreamMessage
only read msg.getText(), which is null for audio events, so the message
landed in handleWebhook with no msgtype, fell through to the default text
branch, found null content, and got dropped at 'Empty message content,
ignoring'. From the user's side: send a voice, nothing happens, no log of
the attempt.
Two surgical edits:
- handleStreamMessage now checks getContent().getRecognition() first; if
present and non-blank, builds payload {msgtype: audio, audio: {recognition}}
before falling back to the existing text path. The earlier comment about
richText being handled inside handleWebhook was wrong — picture and
richText also need their fields propagated through the payload Map; left
a TODO for them.
- handleWebhook gains an explicit case 'audio' branch that pulls text out
of audio.recognition and pushes it onto contentParts.
- ChannelMessage.inputMode now reflects 'voice' when msgtype=audio,
mirroring feishu's behavior so downstream code (memory-extraction
filters, voice-themed system prompts) can tell text vs voice turns apart.
No STT call required — DingTalk transcribes server-side and ships text in
the webhook, so this is a 0-network, 0-config fix.
Mirrors the feishu one-click flow: scan a QR with the DingTalk app,
approve, and the bot's client_id / client_secret get auto-filled instead
of forcing the user through the open-dev console. Saves about seven
manual steps per channel setup.
Backend
- Bump dingtalk-stream from 1.3.5 to 1.3.12. Diff against the classes we
depend on (OpenDingTalkStreamClient, ChatbotMessage, MessageContent,
GenericEventListener) is empty — pure point-release bumps, no API churn.
- New DingTalkAppRegistrationService: synchronously runs init + begin
against /app/registration/{init,begin} on oapi.dingtalk.com to obtain
the device_code and verification URL, then spawns a daemon worker that
polls /app/registration/poll every 5s until SUCCESS / FAIL / EXPIRED is
returned. Sessions evict after 7 minutes, worker has a 6-minute hard
runtime cap, transient HTTP errors do not terminate the loop. Same
shape as the feishu service, but written from scratch because the
dingtalk-stream SDK doesn't wrap this OAuth device flow.
- Two new endpoints under /api/v1/channels/webhook:
POST /dingtalk/register/begin returns session_id;
GET /dingtalk/register/status returns status + qrcode_img (data URI
PNG, ZXing-encoded from the verification URL, matching the feishu and
weixin flows). Status surface: waiting / confirmed / expired / denied.
Frontend
- channelApi.dingtalkRegisterBegin / dingtalkRegisterStatus.
- New useDingTalkAppRegister composable, structurally identical to
useFeishuAppRegister minus the domain argument. Stops polling on
terminal status, fires onConfirmed with {clientId, clientSecret}.
- ChannelEditModal: dingtalk-register-card rendered when channelType is
dingtalk, scoped DingTalk blue (#1f79ff) to differentiate from feishu's
indigo. onConfirmed writes channelConfig.client_id / client_secret so
the existing form fields update reactively.
- i18n: channels.dingtalkRegister.* keys for title / hint / button states
/ scan / confirmed / expired / denied / startFailed.
Saves the user the entire 'go to the open platform -> create an enterprise
app -> copy App ID and Secret' detour. Click a button in the channel form,
scan the QR code, confirm authorization, credentials are auto-filled.
Backend
- Bump com.larksuite.oapi:oapi-sdk from 2.5.3 to 2.6.1, which adds the
scene/registration package wrapping the device-code flow.
- New FeishuAppRegistrationService: each begin() creates a sessionId,
spawns a worker thread, runs the SDK's blocking RegisterApp.register
with onQRCode and onStatusChange wired into a per-session state machine
(PENDING -> WAITING -> CONFIRMED / EXPIRED / DENIED / ERROR). The
session caches the QR data URI so ZXing only encodes once per attempt.
Sessions evict after 5 minutes so closed browsers don't leak the map.
- Two new webhook endpoints under /api/v1/channels/webhook/feishu:
POST /register/begin returns session_id, GET /register/status returns
status + qrcode_img (data URI base64 PNG, ZXing-encoded from the SDK's
verification URL — the raw URL would render as a broken image, so the
encoding step matches the WeCom flow).
- SDK detail caught the hard way: don't pass .domain() or .larkDomain().
The SDK defaults are accounts.feishu.cn / accounts.larksuite.com (the
registration endpoints). open.feishu.cn is the open-API endpoint, a
completely different service. Passing the wrong one makes the SDK parse
HTML as JSON and emit invalid_response.
Frontend
- channelApi: feishuRegisterBegin / feishuRegisterStatus.
- New useFeishuAppRegister composable: state machine that begins the
session, polls status every 2s, prefers qrcode_img over qrcode_url for
the <img> src, stops on terminal status, fires onConfirmed with
{appId, appSecret}.
- ChannelEditModal: a new feishu-register-card above the wecom one. The
composable's onConfirmed writes channelConfig.app_id / app_secret, so
the existing form fields update reactively.
- i18n: channels.feishuRegister.* keys for title / hint / button states /
scan / confirmed / expired / denied / error.
Backend (FeishuChannelAdapter):
- Default connection_mode flips webhook -> websocket on doStart and doReconnect.
- Stale event filter: drop events whose message.create_time is older than
stale_event_threshold_seconds (default 30s) so SDK reconnect replays do not
re-trigger the agent.
- Silent disconnect watchdog runs every 60s; if no events arrive for
silent_disconnect_threshold_seconds (default 1800s) after the first event,
call onDisconnected to force a reconnect cycle. Setting the threshold to 0
disables the watchdog. The watchdog is scheduled before wsClient.start() on
the bring-up path because that call blocks indefinitely.
- Quoted message context: when a reply has parent_id set, fetch the parent
via GET /open-apis/im/v1/messages/{id}, summarize per msg_type (text / post
first paragraph / [Image]/[File]/[Audio]/[Video] placeholders, capped at
200 chars), and prepend [Quoted: ...] to both content text and the first
content part. LRU-cached (200) per message_id.
- AbstractChannelAdapter gains getConfigLong helper for numeric config keys.
Frontend:
- types/index.ts feishu fields: default connection_mode is websocket; the
recommended option moves to the top; verification_token and encrypt_key
get showIf so they only render in webhook mode; new enable_quoted_context
switch (default on) exposes the quoted-message feature.
- ChannelEditModal builds a feishu-specific WEBHOOK_GUIDES path that picks
webhookStep vs websocketStep based on connection_mode, so users only see
steps for the mode they're using.
- i18n: split feishu.step3/step4 into webhookStep/websocketStep, rename
step5 to permissionStep. Channel type labels in zh-CN drop bilingual
prefix (e.g. 'Feishu / Lark (飞书)' -> '飞书').
Migrations:
- V52 was a no-op the first time it ran (matched compact JSON only) and
Flyway refused to re-run after the SQL was fixed. V52 is documented as a
no-op; V53 carries the actual UPDATE with REPLACE covering both compact
and pretty-printed JSON, and an idempotent WHERE for rows already on
websocket. h2 and mysql variants stay in lockstep.
- Reconcile approval status atomically: DB row, message metadata, in-memory store
- Approve and deny both flip the tool-call card + timeline segment to a terminal
state on the gate message — no more orange spinner stuck after a decision
- Frontend hydrate matches by pendingId and reverse-converges to expired so a
refresh after server-side timeout / consume clears the banner without restart
- Stop sweep, GC timeout, and JVM restart all close the loop with consistent
state
- Remove the dead REST /approve endpoint + matching frontend client export so
there is only one resolve path to maintain
Foundation for the ghost-approval root-cause fix.
Adds ResolveOutcome / MetadataDecision; rewrites ApprovalWorkflowService so
every resolve / consume / timeout / supersede transitions through one
two-phase contract: snapshot → DB UPDATE conditional on status=PENDING →
metadata reconciliation → afterCommit memory mutation. ChatController,
ChannelMessageRouter, and ApprovalController all switch to the workflow;
ApprovalService.resolve / resolveAndConsume / consumeApproved /
cancelStalePending / denyAllByConversation are physically removed so
DB-bypass is no longer reachable at compile time.
Specific fixes:
- recoverFromDb preserves DB pendingId + createdAt (was generating fresh
random ids, breaking every later DB sync)
- effectiveExpireAt = expireAt ?? createdAt + PENDING_TTL: legacy rows
with NULL expireAt no longer resurrect as live PENDING after restart
- markPendingApprovalsResolved flips pendingApproval.status + currentPhase
+ MessageEntity.status atomically (was only flipping the first field;
message.status uses existing completed/stopped, not approved/denied,
to stay within the frontend Message.status union)
- GC scheduler moves to ApprovalWorkflowService; timeouts and overflow
evictions now sync DB + metadata + memory through markTimeout
- DB UPDATE rows=0 returns alreadyResolved (concurrent-resolve safe);
exception propagates so @Transactional rolls back; memory stays untouched
- expireRecoveredRow gates metadata write on DB success (was writing
metadata even when DB update failed, producing the worst-case ghost)
- Mockito JDK 21 agent attach fixed via maven-dependency-plugin properties
+ surefire argLine (no more flaky self-attach across machines)
Tests: 34 new across 4 classes (recovery, resolve, GC, metadata sync).
Full suite: 788 / 788.
The two tools-sync scripts ran on every startup and used H2 MERGE INTO
... KEY(id), which overwrites every column on existing rows. That
silently reverted UI-toggled `enabled` and was the proximate cause of
a recent WriteFileTool/EditFileTool outage.
They were also a strict subset of the fresh-install seed (data-zh.sql /
data-en.sql register all 19 builtins; the sync scripts only 16) and out
of date. Per-tool Flyway migrations (V3, V31) are already the canonical
'register a new builtin' path, so the sync layer was duplicated and
error-prone.
Delete both files and the runToolSyncScript() loader. Tool descriptions
shown to the LLM come from @Tool annotations in code, not the DB row,
so removing per-startup metadata refresh has no functional impact.