docs: add user guide for the persistent-goal feature

This commit is contained in:
matevip 2026-05-21 16:25:50 +08:00
parent 2b6a4c64c9
commit e8ec612ca8
2 changed files with 507 additions and 0 deletions

View File

@ -0,0 +1,254 @@
---
title: Persistent Goals — lock in across turns, let the worker follow up
description: MateClaw's Goal system lets a digital worker lock a multi-turn task as a goal, self-evaluate progress, and optionally drive itself forward until done or out of budget.
head:
- - meta
- name: keywords
content: Goal,Agent,multi-turn,auto-evaluation,auto-followup,persistent,MateClaw
---
# Persistent Goals
> **You used to repeat the context every turn. Now you set a goal once, the worker follows.**
You say "deploy this blog to fly.io" in one turn, the worker answers, and stops. Next turn you have to remember to ask "is DNS set? cert signed? tests run?" — you're keeping the goal in your head, not the worker.
Goals flip that. **You say it once, the worker locks the goal and self-checks every turn: what's still missing? Should I take the next step myself?**
It is not a new tab or a new feature. It is a **state** the worker has. A ring appears around the assistant avatar. How filled the ring is, is how close you are to done. When done, the ring goes away.
---
## What it looks like
Not a banner. Not a dialog. Not a separate page.
A **ring around the assistant avatar**.
| State | Visual | Meaning |
|---|---|---|
| No goal | Plain avatar | This conversation has no goal — same as before |
| Active | Avatar + orange ring | Goal in flight, ring fills to progress |
| Evaluating | Avatar + sand breathing halo | Backend is judging this turn's answer |
| Completed | Avatar + green ring (briefly) | Goal reached; ring fades, conversation continues |
| Exhausted | Avatar + red-orange ring | Budget used up — your call to extend or let go |
**Hover the avatar** to see the full tooltip — title + what's still missing. Don't hover, don't get bothered. That's the design.
---
## Three ways to set a goal
In increasing order of how much you have to spell out:
### Way 1 — Let the worker decide
State the multi-turn nature of the task plus an explicit setGoal request:
> I want to do a complete project: translate the README to English, open a PR, address review feedback, merge. This spans many turns. **Please use setGoal to lock it in**, self-evaluate each turn, turnBudget=8, autoFollowup on.
The worker picks up the two signals ("long task" + "setGoal requested") and creates the goal, auto-summarizing the title from context. You see a ring next to its avatar — goal is locked.
### Way 2 — Direct tool command
Tell the worker exactly which tool to call with which params:
> Please call setGoal immediately, title="Deploy blog to fly.io", turnBudget=10, autoFollowup=true. Do not ask any clarifying questions.
The "do not ask clarifying questions" clause matters — otherwise the worker's instinct is to ask "where's the code? what domain?" first.
### Way 3 — Programmatic via the REST API
For automation and external scripts, the endpoint is direct:
```
POST /api/v1/goals
{
"conversationId": "conv-xxx",
"agentId": "1000000001",
"workspaceId": 1,
"title": "Deploy blog to fly.io",
"description": "...",
"exitCriteria": "DNS + SSL + healthcheck + tests pass",
"turnBudget": 10,
"llmCallBudget": 200,
"autoFollowupEnabled": false
}
```
Full surface in the [API reference](./api).
---
## What a goal carries
Four required:
| Field | Meaning |
|---|---|
| **title** | Short label, shown on avatar hover |
| **description** | Full statement of what you want |
| **exitCriteria** | LLM-readable bar the evaluator scores against (e.g. "tests pass + deployed") |
| **budgets (turnBudget + llmCallBudget)** | Failsafes against runaway iteration |
Optional:
- **autoFollowupEnabled** — when on, the worker may continue itself if it judges the goal incomplete, without waiting for your next message
- **followupCooldownSeconds** — minimum delay between two consecutive auto-followups
---
## How it runs
After every turn, a backend evaluator node runs:
1. Reads the worker's final answer + last few messages of context
2. Calls a lightweight evaluator model (point this at a cheap one) asking: completion 01? what's the gap? continue or done?
3. Writes the result into the `mate_agent_goal_event` timeline
4. Decides next step: complete / exhaust budget / continue / auto-followup
**Key invariant**: evaluation runs *after* the final answer has streamed to your screen — it never blocks you seeing the reply. You see the answer appear → the ring updates a moment later.
### Auto-followup
When `autoFollowupEnabled=true` and this turn's evaluator decision is "continue", the backend:
1. Writes a `followup_injected` event to the timeline
2. APPENDs a user message to the conversation: *"Continue working on the goal. Still missing: {gap}. Take the next concrete step."*
3. Re-enters the reasoning loop — the next assistant reply lands right after the first
Feels like: the worker answers a segment → pauses a beat → **keeps going** — like a person who finished one step, thought for a second, and continued.
---
## Four built-in tools (worker-callable)
These four ship as agent-wide system tools — no binding setup needed:
| Tool | Purpose | Prompt example |
|---|---|---|
| **setGoal** | Create a goal | "Use setGoal to lock in this task, title=..." |
| **addGoalCriterion** | Append a sub-criterion to the active goal | "Add: must support IPv6" |
| **completeGoal** | Explicitly mark done | "All items done — call completeGoal" |
| **getGoalStatus** | Inspect current state | "How are we doing?" |
On completion (`completeGoal` or evaluator score ≥ 0.95), the worker forwards a summary to its [long-term memory](./memory) so future conversations can recall it.
---
## Sub-agents cannot mutate the parent's goal
In [multi-agent collaboration](./agents) a parent worker can delegate to a child worker. Children **don't see** the four goal tools — the goal is the parent conversation's state, the child is a stateless executor.
> This is intentional. Children do work for the parent, but the goal stays owned by the parent.
---
## When the budget runs out
```
turnsUsed >= turnBudget OR (agentLlmCallsUsed + evalLlmCallsUsed) >= llmCallBudget
```
Either one hit → goal status flips to **exhausted**, no more evaluations, no more follow-ups, ring turns red-orange. The last turn's assistant reply still goes through.
Your options:
- **Raise the budget and resume**`PATCH /api/v1/goals/{id}` to widen budgets then resume (no UI button in v1 — use the API or abandon and re-create)
- **Let it go** — call abandon; the conversation slot is freed for a new goal
---
## State machine
```
create
active
↓ ↑
paused
active ──evaluator score≥0.95 / completeGoal──→ completed (terminal)
active ──turns_used / llm_calls exhausted ────→ exhausted (terminal)
active ──user abandon ────────────────────────→ abandoned (terminal)
```
Terminal states (completed / exhausted / abandoned) cannot revive. To keep going, create a fresh goal — intentional simplicity, avoids messy "resurrect with what budget" semantics.
**One active goal per conversation**: at most one active row at any time. Terminal rows stay in history, don't count against the slot. Enforced at the DB layer with a generated column + unique index (H2 / MySQL), plus service-level precheck and audit — defense in depth.
---
## What this system does not do
A few deliberate non-features:
- **No nested goals / goal trees** — one goal per conversation, no OKR stack
- **No "goal templates"** — every goal is hand-written
- **No cross-conversation goal migration** — use a [workflow](./workflow) for that
- **No completion score in the UI**`completionScore` is an internal engineering protocol, not user vocabulary. The UI speaks via a ring; hover reveals the natural-language gap the evaluator wrote. The numeric score stays in logs and the API for debugging
---
## Full event timeline (drawer view)
Each goal has an append-only event log, newest first:
| Event | Trigger |
|---|---|
| `created` | setGoal tool or REST POST |
| `evaluated` | every turn after evaluator runs |
| `followup_injected` | autoFollowup fired and injected a prompt |
| `completed` | evaluator concluded done, or completeGoal tool |
| `exhausted` | budget hit |
| `paused` / `resumed` / `abandoned` | user actions |
| `criterion_added` | addGoalCriterion tool |
Pull via `GET /api/v1/goals/{id}/events`. See [API reference](./api).
---
## Configuration
`application.yml`:
```yaml
mateclaw:
goal:
# Master switch; when off, the graph node passes through for every call.
enabled: true
# Default turn budget when the user doesn't override.
default-turn-budget: 20
# Default combined (agent + evaluator) LLM call budget.
default-llm-call-budget: 200
# Minimum seconds between two consecutive auto-followups.
auto-followup-cooldown-seconds: 0
# Model used by the evaluator. Empty = same model as the chat agent.
# Recommended: a cheap model like qwen-turbo / glm-4-flash.
evaluator-model: ""
# Max recent messages included in the evaluator prompt.
evaluator-context-messages: 8
```
---
## Database
Two tables, all `mate_`-prefixed:
| Table | Purpose |
|---|---|
| `mate_agent_goal` | Goal itself; status / budgets / dual LLM counters / auto-followup config |
| `mate_agent_goal_event` | Append-only event log; powers the timeline view |
Flyway migration `V120__agent_goal.sql` (H2 + MySQL dialects).
---
## One-liner
**A goal isn't a new feature on the worker. It's a state change.**
Before, the worker forgot the moment it answered. Goals make a worker remember one thing across many turns — what it's working on, what's still missing, when it counts as done. You say it once. The ring next to the avatar tracks the rest.

View File

@ -0,0 +1,253 @@
---
title: 持久化目标 — 跨多轮锁定,让员工自己跟进
description: MateClaw 的 Goal 系统让数字员工把跨多轮的任务锁成一个目标,自己评估进度、自己续命,直到完成或耗尽预算。
head:
- - meta
- name: keywords
content: Goal,目标管理,Agent,多轮对话,自动评估,auto-followup,持久化,MateClaw
---
# 持久化目标
> **以前你每轮都要把上下文重复一遍。现在你定一个目标,员工自己跟。**
一次对话里你说"帮我把这个博客部署到 fly.io",员工答完一轮就停了。下一轮你要再问"DNS 配好没?证书呢?测试跑了吗?"——你在替它记目标。
Goal 把这件事翻过来。**你说一次,员工锁住目标,自己每轮自检:还差什么?要不要自己再做一步?**
它不是聊天里的一个新功能。它是员工的一种**状态**。员工头像周围多了一圈光,光填多少就是离完成多远。完成了,光消失。
---
## 它在视觉上长什么样
不是一个 banner。不是一个 dialog。不是一个独立的标签页。
**assistant 头像周围的一圈环**
| 状态 | 视觉 | 含义 |
|------|------|------|
| 无目标 | 头像就是头像 | 这条对话没绑目标,跟过去一样 |
| 进行中 | 头像 + 橙色环 | 有目标在跟,光填到进度处 |
| 评估中 | 头像 + 沙金呼吸光晕 | 后台正在判断这轮答案 |
| 已完成 | 头像 + 绿色环(短暂出现) | 目标达成,环随后消失,对话继续 |
| 预算耗尽 | 头像 + 红橙色环 | 用完 budget需要你决定加预算还是放手 |
**hover 头像**才显示完整 tooltip — 标题 + 还差什么。不 hover 就不打扰你。这是设计意图。
---
## 怎么定一个目标
三种方式,按门槛从低到高:
### 方式 1 — 让员工自己定
你只要在第一次描述任务时让员工知道这是个长任务:
> 我要做一个完整的项目:把 README 翻译成英文、提 PR、走 review、合并。这跨多轮**请你用 setGoal 锁定**每轮自我评估turnBudget=8autoFollowup 开启。
员工识别到"长任务"+"明确要求 setGoal"两条信号会自动调用工具创建目标title 从对话上下文自动归纳。你只需点开它的回答,看见头像旁边多了一圈光,就知道目标已锁。
### 方式 2 — 直接命令工具
不想让员工判断,你直接告诉它调哪个工具、传什么参数:
> 请立刻调用 setGoal 工具title="部署博客到 fly.io"turnBudget=10autoFollowup=true。不要问任何前置确认。
"不要问前置确认"这一句很重要 — 否则员工会先问"代码在哪?域名是什么?" 它的本能就是先澄清。
### 方式 3 — 通过 API 程序化创建
对自动化、外部脚本REST 端点直接可用:
```
POST /api/v1/goals
{
"conversationId": "conv-xxx",
"agentId": "1000000001",
"workspaceId": 1,
"title": "部署博客到 fly.io",
"description": "...",
"exitCriteria": "DNS+SSL+健康检查+测试通过",
"turnBudget": 10,
"llmCallBudget": 200,
"autoFollowupEnabled": false
}
```
完整接口列表见 [API 参考](./api)。
---
## 一个目标里有什么
最少四样:
| 字段 | 含义 |
|---|---|
| **标题 (title)** | 短句,光环 hover 时显示 |
| **描述 (description)** | 完整诉求 |
| **退出判据 (exitCriteria)** | LLM 可读的判据evaluator 按这个打分(比如 "DNS 配好+测试通过" |
| **预算 (turnBudget + llmCallBudget)** | 防失控上限 |
可选:
- **自动延续 (autoFollowupEnabled)**:开了之后,员工答完一轮如果觉得"还没完成",会自己接着做下一步,不等你催
- **冷却 (followupCooldownSeconds)**:两次自动延续之间至少隔多久
---
## 它在后台是怎么运转的
每次员工回答完一轮,后台会跑一个评估节点。这个节点:
1. 取员工这一轮的最终回答 + 最近几条消息上下文
2. 调一个轻量 evaluator建议指向便宜的小模型完成度多少0~1还差什么该继续还是已完成
3. 把答案写到 `mate_agent_goal_event` 时间线表里
4. 决定下一步:完成 / 预算耗尽 / 继续 / 自动延续
**关键不变量**:评估发生在 final answer 已经串给你看完之后 — **不阻塞用户看回答**。你看到回答出现 → 短暂后头像旁边的光环进度变化。
### 自动延续是怎么发生的
如果 `autoFollowupEnabled=true` 且这一轮 evaluator 判 "continue",后台会:
1. 写一条 `followup_injected` 事件到时间线
2. 给对话末尾 APPEND 一条用户消息:"Continue working on the goal. Still missing: {gap}. Take the next concrete step."
3. 让员工再跑一轮 reasoning**这一轮的回答就直接接在第一轮后面**
你的体感是:员工答完一段 → 停半拍 → **继续往下做** — 就像一个人做完一步停了一下想了想然后继续。
---
## 4 个内置工具(员工可用)
员工的工具集里默认包含这 4 个(无需手动绑定,是 agent-wide 系统级工具):
| 工具 | 用途 | 触发提示词示例 |
|---|---|---|
| **setGoal** | 创建目标 | "请用 setGoal 锁定本次目标title=..." |
| **addGoalCriterion** | 追加子准则到已有目标 | "再加一条准则:必须支持 IPv6" |
| **completeGoal** | 显式标记完成 | "所有事项已做完,请 completeGoal" |
| **getGoalStatus** | 查询当前 goal 状态 | "我们现在进展到哪了?" |
完成时 (`completeGoal` 或 evaluator 判 score≥0.95),员工会把这个目标的总结同步到[长期记忆](./memory),后续对话能查得回来。
---
## 子员工不能改父员工的目标
[多员工协作](./agents)里 parent 员工可以委派 child 员工干活。Child **看不到**这 4 个 goal 工具 — 目标是 parent 会话的状态child 是无状态的执行体。
> 这一条是设计意图,不是 bug。child 帮 parent 做事,但目标的"所有权"留在 parent 那。
---
## 预算耗尽时
```
turnsUsed >= turnBudget 或 (agentLlmCallsUsed + evalLlmCallsUsed) >= llmCallBudget
```
任一条命中 → 目标状态翻为 **exhausted**,不再触发评估、不再注入 follow-up光环变橙红色。员工的最后一轮回答会正常发送给你。
你的选择:
- **加预算 + 恢复** — 通过 `PATCH /api/v1/goals/{id}` 改 budget 后 resumev1 暂未给 UI 提供按钮,可以走 API 或先 abandon 重新创建)
- **放手** — 调 abandonconversation 上释放槽位,可以重设新目标
---
## 状态机
```
create
active
↓ ↑
paused
active ──evaluator score≥0.95 / completeGoal──→ completed (终态)
active ──turns_used/llm_calls 用完 ─────────→ exhausted (终态)
active ──user abandon ─────────────────────→ abandoned (终态)
```
终态 (completed / exhausted / abandoned) 不能复活。要继续就开新 goal — 这是有意保留的简单约束,避免 "重启" 带来的预算账目混乱。
**一会话一目标**:每个 conversation 同一时刻最多一个 active goal。终态 goal 留在历史里不占名额。底层用 H2 / MySQL 的生成列 + 唯一索引保证并发安全service 层 + DB 层双重防御。
---
## 这套系统不做什么
按设计原则保留了几个"不做"
- **不做嵌套目标 / 目标树** — 一个 conversation 一个目标,不堆 OKR
- **不做"目标模板"** — 每个目标是手写的,不是从库里挑的
- **不做跨 conversation 迁移目标** — 想要那效果,请用[工作流](./workflow)
- **不暴露评估分数给用户** — 那个 `completionScore` 是工程内部协议不是用户语言。UI 用一圈光说话hover 显示 evaluator 写的 gap 文本(自然语言)。后端日志和 API 里仍可见数值,方便调试
---
## 完整事件时间线drawer 抽屉视图)
每个目标都有一份只增不删的事件时间线,按时间倒序展示:
| 事件 | 触发 |
|---|---|
| `created` | setGoal 工具或 REST POST |
| `evaluated` | 每轮答完evaluator 跑完一次 |
| `followup_injected` | autoFollowup 触发,注入了 prompt |
| `completed` | evaluator 判完成或 completeGoal 工具 |
| `exhausted` | budget 用尽 |
| `paused` / `resumed` / `abandoned` | 用户手动操作 |
| `criterion_added` | addGoalCriterion 工具 |
通过 `GET /api/v1/goals/{id}/events` 拉取(详见 [API 参考](./api))。
---
## 配置项
`application.yml`
```yaml
mateclaw:
goal:
# 主开关;关闭后图节点对所有调用 pass-through
enabled: true
# 默认 turn 预算
default-turn-budget: 20
# 默认 LLM 调用预算agent + evaluator 之和)
default-llm-call-budget: 200
# 自动延续之间至少隔多久(秒)
auto-followup-cooldown-seconds: 0
# 评估器使用的模型;空字符串 = 沿用对话当前模型便宜的小模型推荐qwen-turbo / glm-4-flash
evaluator-model: ""
# 评估 prompt 携带的历史消息条数上限
evaluator-context-messages: 8
```
---
## 数据库
两张表,都用 `mate_` 前缀:
| 表 | 用途 |
|---|---|
| `mate_agent_goal` | 目标本体;含 status / budget / 双 LLM 计数器 / 自动延续配置 |
| `mate_agent_goal_event` | 目标的事件追加日志drawer 时间线读它 |
迁移由 Flyway 跑 `V120__agent_goal.sql`H2 + MySQL 双方言)。
---
## 一句话总结
**Goal 不是给员工加一个功能。是改它的状态。**
以前的员工"答完就忘"。Goal 让员工跨多轮记住一件事 — 它在干什么、还差什么、什么时候算完。你只用说一次。剩下的,让头像旁边那圈光替你跟。