Go to file
matevip b4697f2806 fix(cron): post-deploy bug bundle — flakiness, scheduler, channel UI
User-reported field issues + a deeper code audit revealed multiple
overlapping bugs in the prior cron-channel delivery change. This fixes
all six.

#1 — Concurrency race on ToolExecutionExecutor (root cause of 'sometimes
   succeeds, sometimes fails' tool calls). The volatile instance fields
   currentRequesterId / currentWorkspaceBasePath / currentChatOrigin
   were shared by every conversation routed through the same per-agent
   executor; one user mid-build-loop while another's execute()
   overwrote the field would cross-contaminate the captured values into
   PreparedToolCall. Fix: kill the instance fields, thread
   origin/requester/workspace as method params straight into
   PreparedToolCall snapshot. Comment pins the rule so it cannot regress.

#2 — CHAT_ORIGIN missing from KeyStrategyFactory (latent timebomb,
   masked by spring-ai-alibaba-graph-core's non-filtering builder path).
   Without an addStrategy registration, multi-node state merges in long
   ReAct / Plan-Execute loops drop the key, ActionNode reads
   ChatOrigin.EMPTY, and the cron persists with channel_id=NULL. Also
   caught 4 more keys that were latently unregistered:
   WORKSPACE_BASE_PATH, STOP_REQUESTED, RETURN_DIRECT_TRIGGERED,
   DIRECT_TOOL_OUTPUTS. All five now registered in both ReAct and
   Plan-Execute factories.

#3 — CronJobs UI didn't surface channel binding. CronJobDTO carried
   channelId / deliveryConfig but the list page never rendered them.
   Added: (a) 'channel' column on list page, (b) channel + targetId
   rows in the detail modal, (c) backend batch-loads channel names via
   ChannelMapper.selectBatchIds so the column shows the human-readable
   name, (d) i18n keys (zh + en), (e) channelName field on TS CronJob
   type.

#4a — DingTalk targetId expiry. ChannelChatOriginFactory.resolveTargetId
   used to prefer ChannelMessage.replyToken which for DingTalk encodes
   a sessionWebhook URL that expires ~90 minutes after the inbound
   message. Cron persisted with that webhook then dies with 401/403 and
   marks NOT_DELIVERED forever. Fix: prefer the stable chatId, fall
   back to senderId — both work indefinitely via DingTalk's Robot API.

#4b — Scheduler pool exhaustion under long LLM. CronJobService's
   ThreadPoolTaskScheduler ran with poolSize=4 AND the LLM call lived
   on the scheduler thread. Four concurrent crons saturated the pool
   and the 5th silently missed its tick. Fix: keep scheduler tiny (it
   just fires triggers) and offload runAgent to a dedicated
   virtual-thread executor (cron-execute-* threads). LLM workload is
   I/O-bound — virtual threads scale to thousands at trivial cost.

#5 — Minor latent bugs:
   - AbstractCronResultDelivery.claimRun used .in(... 'NONE','PENDING',null),
     but SQL IN never matches NULL. Rewrote as IS NULL OR IN
     (NONE,PENDING) so legacy pre-V57 rows can still claim.
   - CronDeliveryListener.onCompletedRaw was an empty @EventListener
     with a wrong-headed comment about test fallbackExecution. Removed.
   - CronJobTool.resolveAgentId silently returned 1L when origin
     lacked an agentId — would silently bind to whatever agent #1
     happens to be. Replaced with explicit error so wiring bugs surface
     immediately instead of producing scheduled-but-never-runs crons.

State-key registration guard. New StateKeyRegistrationCoverageTest
scans MateClawStateKeys via reflection and parses
AgentGraphBuilder.java to extract every
.addStrategy(MateClawStateKeys.X, ...). Asserts every non-_NODE
constant appears in at least one factory. Caught the 4 unregistered
keys above on first run; will catch any future 'forgot to register'
regression.

Tests: 33 unit/arch tests + 27 regression in touched areas — all green.
Vue typecheck clean.

Refs: #25, #16
2026-04-28 21:45:07 +08:00
assets docs: add preview screenshot to README 2026-04-10 18:35:13 +08:00
docker/searxng fix(docker): bake searxng settings.yml into custom image 2026-04-24 23:25:03 +08:00
mateclaw-plugin-api feat(plugin): Plugin SDK + UI layout improvements 2026-04-13 18:38:03 +08:00
mateclaw-plugin-sample feat(plugin): Plugin SDK + UI layout improvements 2026-04-13 18:38:03 +08:00
mateclaw-server fix(cron): post-deploy bug bundle — flakiness, scheduler, channel UI 2026-04-28 21:45:07 +08:00
mateclaw-ui fix(cron): post-deploy bug bundle — flakiness, scheduler, channel UI 2026-04-28 21:45:07 +08:00
mateclaw-webchat release: v1.1.0 2026-04-17 19:41:03 +08:00
.dockerignore chore: bump version to 1.1.137-SNAPSHOT 2026-04-18 21:58:54 +08:00
.env.example build(docker): pass MAVEN_FLAGS build-arg to support aliyun-first profile in CN builds 2026-04-25 10:32:23 +08:00
.gitignore fix(approval): unify tool-approval state machine across DB / message metadata / memory 2026-04-27 22:25:27 +08:00
docker-compose.yml build(docker): pass MAVEN_FLAGS build-arg to support aliyun-first profile in CN builds 2026-04-25 10:32:23 +08:00
LICENSE chore: add Apache-2.0 license 2026-04-04 23:29:51 +08:00
README_zh.md docs: update README title and tagline 2026-04-21 09:27:42 +08:00
README.md docs: update README title and tagline 2026-04-21 09:27:42 +08:00

MateClaw Logo

MateClaw

Your second brain

GitHub Repo Documentation Live Demo Website Java Version Spring Boot Vue Last Commit License

[Website] [Live Demo] [Documentation] [中文]

MateClaw Preview


Most AI tools die when their vendor has a bad day. Most forget you the moment the tab closes. Most give you a chatbox and call it a product.

MateClaw is the whole widget. One deployment. Reasoning, knowledge, memory, tools, channels — built together, not bolted on. And when your primary model goes down, the next one picks up mid-sentence.


Three things that make it different

1 · Your AI doesn't die when a model does

Primary key expired. Vendor returns 401. Network blip. Quota drained.

Other tools hand you a red error card. MateClaw routes to the next healthy provider — DashScope, OpenAI, Anthropic, Gemini, DeepSeek, Kimi, Ollama, LM Studio, MLX, 14+ in total — and the user sees the reply finish. A provider health tracker parks bad vendors in a cooldown window so they don't waste seconds on every turn.

You don't write a retry script. You drag providers into priority order in Settings → Models and watch the health dashboard fill with green dots as requests route around failures in real time.

Upload a PDF, a batch of markdown, a scraped page — raw material in.

MateClaw's LLM Wiki digests it into structured pages, builds [[links]] between them, and remembers where every sentence came from. Click a citation, see the exact source chunk. Ask a question, the page you get is stitched from the right chunks — with references you can verify.

This is the difference between a warehouse and a library.

3 · One product, five surfaces

Surface What it is
Web Console Full admin — agents, models, tools, skills, knowledge, security, cron
Desktop Electron app with a bundled JRE 21. Double-click, run. No Java install
Webchat Widget One <script> tag embed. Drop it on any site
IM Channels DingTalk · Feishu · WeChat Work · Telegram · Discord · QQ
Plugin SDK Java module for third-party capability packs

Same brain. Same memory. Same tools. Different doors.

$0 · No tokens metered. No seats billed. Your server. Your data. Your keys.


What's in the box

Agent runtime

ReAct for iterative reasoning. Plan-and-Execute for complex multi-step work. Dynamic context pruning, smart truncation, stale-stream cleanup — the boring stuff that makes long conversations actually work.

Knowledge & memory

  • LLM Wiki — raw materials digest into linked pages with citations
  • Workspace memoryAGENTS.md, SOUL.md, PROFILE.md, MEMORY.md, daily notes
  • Memory lifecycle — post-conversation extraction, scheduled consolidation, dreaming workflows

Tools, skills, MCP

Built-in tools for web search, files, memory, date/time. MCP over stdio / SSE / Streamable HTTP. SKILL.md packages from the ClawHub marketplace. A Tool Guard layer with RBAC, approval flows, and path protection — capability needs boundaries.

Multimodal creation

Text-to-speech · Speech-to-text · Image · Music · Video. First-class, not add-ons.

Enterprise-ready

RBAC + JWT. Full audit trail. Flyway-managed schema that auto-heals on upgrade. One JAR to ship. MySQL in production, H2 for dev — nothing to change in your code.


AI is becoming infrastructure

On March 2, 2026, Claude went dark for 4 hours across API, web, and mobile. Three weeks later, another 5 hours. Every company that bet their AI strategy on a single vendor spent those outages staring at red error cards.

This is the same shift databases went through around 2010 and cloud went through around 2018: the winning layer stops being tied to one supplier. 57% of companies now run AI agents in production. None of them want one vendor's bad day to become their bad day.

MateClaw is that layer — built the Spring Boot way.


Why MateClaw

MateClaw OpenClaw Hermes Agent Claude Code Cursor
Multi-vendor failover Chain + health tracker + cooldown Swap providers via config Orchestration w/ retry Anthropic only One model
Knowledge digestion LLM Wiki + page-level citations Canvas + memory Skills Hub + memory Code index
Multi-user admin RBAC + approval flow + audit Config-file first Single-user CLI Enterprise tier Teams plan
Surfaces Web admin + Desktop + Widget + SDK + 6 IM 25+ chat channels 15+ channels (CLI-led) 3 IM preview IDE only
Stack Java (Spring Boot) TypeScript Python TypeScript Electron/TS
License / Price Apache 2.0 · Free MIT · Free MIT · Free Proprietary · $20200/mo Proprietary · $0200/mo

OpenClaw and Hermes Agent are excellent personal AI platforms — pick either if you're running one user on one laptop, building your own agent from CLI, and treating everything as config files to hand-tune. Both have bigger communities than MateClaw today.

MateClaw is the version built for teams. RBAC per agent, per model, per tool. An approval flow that pauses risky actions for review. Full audit trail. A web admin dashboard where one operator manages 50 agents across 14 vendors. Spring Boot inside — drop-in for any Java shop already running production services.

Same "whole widget" philosophy. Different center of gravity.


Quick start

# Backend
cd mateclaw-server
mvn spring-boot:run           # http://localhost:18088

# Frontend
cd mateclaw-ui
pnpm install && pnpm dev      # http://localhost:5173

Login: admin / admin123

Docker

cp .env.example .env
docker compose up -d          # http://localhost:18080

Desktop

Download from GitHub Releases. Bundles JRE 21. No Java install needed.


Architecture

Business Architecture

Technical architecture

Technical Architecture


Project structure

mateclaw/
├── mateclaw-server/        Spring Boot 3.5 backend (Spring AI Alibaba, StateGraph runtime)
├── mateclaw-ui/            Vue 3 + TypeScript admin SPA (built into the server JAR)
├── mateclaw-webchat/       Embeddable chat widget (UMD / ES bundles)
├── mateclaw-plugin-api/    Java SDK for third-party capability plugins
├── mateclaw-plugin-sample/ Reference plugin implementation
├── docker-compose.yml
└── .env.example

Desktop binaries ship via GitHub Releases with a bundled JRE 21 — no Java install needed.

Tech stack

Layer Technology
Backend Spring Boot 3.5 · Spring AI Alibaba 1.1 · MyBatis Plus · Flyway
Agent StateGraph runtime · ReAct + Plan-Execute
Database H2 (dev) · MySQL 8.0+ (prod)
Auth Spring Security + JWT
Frontend Vue 3 · TypeScript · Vite · Element Plus · TailwindCSS 4
Desktop Electron · electron-updater · JRE 21 (bundled)
Widget Vite library mode · UMD + ES bundles

Documentation

Full docs at claw.mate.vip/docs — setup, architecture, each subsystem, API reference.

Roadmap

Sharper multi-agent collaboration · Smarter model routing · Deeper multimodal understanding · Longer-lived memory · A richer ClawHub.

Contributing

git clone https://github.com/matevip/mateclaw.git
cd mateclaw
cd mateclaw-server && mvn clean compile
cd ../mateclaw-ui && pnpm install && pnpm dev

Why the name

Mate is companion. Claw is capability.

Something that stays with you — and grabs work and moves it.

License

Apache License 2.0. No asterisks.