编程

Plan

根据风险信号决定先规划还是直接执行,并按风险等级匹配规划深度,包含步骤、估算与回滚。

它能做什么

它先扫描任务信号(不可逆、依赖、时长等),决定究竟该规划还是直接执行,然后按最高信号匹配 L0–L4 的深度。每一步都要带可观测的完成判据,并把最可能推翻方案的风险假设尽量前置验证;不可逆动作的回滚步骤要先于该动作写好并测试通过。估算用区间表示,高低比超过 3 倍就改用 spike 而不是硬估;执行中连续两次偏离就会触发重新规划,完成结果再反馈回按类型累计的深度默认值。

什么时候用它

  • 迁移、部署、删除或对外发送动作前的方案设计
  • 恢复或交接跨天 / 跨会话的任务
  • 一次性尝试刚失败,或执行已偏离原方案时
  • 用户要求做 scoping、拆解或时间估算

技能文档

User preferences, the outcome log, and active plans live in ~/Clawic/data/plan/ (see setup.md on first use). If you have data at an old location (~/plan/ or ~/clawic/plan/), move it to ~/Clawic/data/plan/.

When To Use

  • A task has multiple steps, dependencies between them, or any irreversible action (deploy, migration, delete, send)
  • Success criteria are unclear, or the estimate would be a guess
  • The user asks for a plan, breakdown, scoping, or estimate before work starts
  • Execution has drifted from an existing plan, or a one-shot attempt just failed
  • Resuming or handing off work that spans sessions or days
  • Not for personal productivity systems, time blocking, or calendars — that is productivity
  • Not for managing a portfolio of ongoing projects — that is projects

Modes: act-as (plan your own execution — the default) and advise (draft a plan the user will execute); both use the same depth rules. Always make the risk decision first, even when the answer is "no plan needed".

Quick Reference

SituationDo this
Done before successfully, fully reversibleExecute directly (L0)
Single deliverable, ≤30 min (quick_task_minutes), reversibleThink through steps, no doc (L1)
Multi-step, >30 min, or any irreversible stepBullet plan, share with human (L2)
Dependencies across components, or spans >1 dayMilestone plan with checkpoints (L3) → long-horizon.md
High stakes AND novel task typeFull plan, human validation before step 1 (L4) → approval.md
Estimate high/low ratio >3xSpike first → estimation.md, strategies.md
A step has no attachable done-checkConvert to decision doc or spike → decomposition.md
Execution drifting from the plan2-consecutive-deviation trigger → replanning.md
Step blocked, or scope added mid-flightBlock and bolt-on protocols → replanning.md
Human approves instantly, or "just do it"approval.md
Resuming yesterday's (or last week's) planResume protocol → long-horizon.md
Anything else / unsurePlan at L2 — the cost is asymmetric (Core Rule 8)

Depth on demand: decomposition.md steps, altitude, done-checks · estimation.md ranges, spikes, calibration · risk.md irreversibility, rollback, blast radius · strategies.md sequential, parallel, iterative, spike, checkpoint · replanning.md deviations, blocks, abandonment · approval.md validation, checkpoints, scope changes · long-horizon.md multi-day, resume, handoff · outcomes.md logging and learning · setup.md first use.

Core Rules

  1. A plan's value is the risk decision, not the step list. Force it early: what single assumption could invalidate the whole approach, and can it be tested cheaply first? A to-do list with no risk ordering adds ceremony, not safety (risk.md).
  2. Depth = the highest level any single signal triggers, never an average. A 20-minute task with one irreversible step is L2, not L1 — the short duration does not dilute the irreversibility.
  3. Irreversibility dominates. One irreversible step anywhere forces at least L2 — you cannot iterate your way out of a deleted database or a sent email. Classification and blast radius in risk.md.
  4. Every step carries an observable done-check. A step you cannot attach a check to is not a step, it is a hope: "investigate X" becomes "decision doc: X vs Y, with the pick". Check catalog in decomposition.md.
  5. Order by information, not convenience. The step most likely to invalidate the plan goes as early as dependencies allow. Migration example: write and test the rollback script as step 1, not last — if rollback is impossible, you want to know before touching data.
  6. Estimates are ranges; ratio >3x means spike first. If the high/low ratio exceeds 3x, that is not an estimate, it is an unknown. 2-8h (4x) → spike; 3-6h (2x) → plan. Building and calibrating ranges in estimation.md.
  7. Detail decays past the first unknown. Steps written beyond it are speculation you will rewrite. Plan in steps to the first checkpoint; milestones beyond (long-horizon.md).
  8. When uncertain, plan — the cost is asymmetric. A bullet plan costs minutes; a failed one-shot costs the redo plus the cleanup plus the trust.

The Planning Decision

Before executing, scan for signals:

SignalOne-shot OKPlan needed
Task done before successfully
Clear single deliverable
Reversible if wrong
Multiple components
Dependencies between steps
Any irreversible step
Touches production data or external users
Ambiguous success criteria
Estimated > quick_task_minutes of work

Reversible vs recoverable — the distinction that decides the column: reversible means undo restores the prior state (git revert); recoverable means a good state is reachable at a cost (restore last night's backup, lose a day of writes). Recoverable-at-cost counts as "plan needed", not "reversible" — full taxonomy and recovery-window decay in risk.md.

Plan Depth Levels

LevelTriggerFormat
L0Done before successfully, fully reversibleExecute directly
L1Single deliverable, ≤ quick_task_minutes, reversibleThink through steps; no doc
L2Multi-step, or > quick_task_minutes, or any irreversible stepBullet plan shared with the human
L3Dependencies between components, or spans >1 dayMilestone plan with checkpoints; persisted (long-horizon.md)
L4High stakes AND novel (this task type never done)Full plan; human validation before step 1 (approval.md)

Depth = highest single signal (Core Rule 2). Per-type learned defaults override this table once the outcome log has evidence (outcomes.md, Current Defaults).

Plan Format (L2-L4)

📋 Plan: [goal]

Why planned: [the signal that triggered planning — one line]

Steps:
1. [step] — [observable output that proves it is done]
2. [step] — [observable output]
3. [step] — [observable output]

Riskiest assumption: [what invalidates the approach] — tested in step [N]
Irreversible steps: [numbers] — rollback: [how, tested in step M] (or "none" + mitigation)

Estimate: [low–high range] — high end fires if [driver]
Validation: [none | human approves before step N]

[The specific question, not "Ready to start?" — approval.md]

Rules that make the format work:

  • 3-7 steps. More than 7 means wrong altitude: group into milestones and plan only the first milestone in step detail (decomposition.md).
  • Every step passes the done-check test (Core Rule 4); activity verbs get converted before the plan ships.
  • The rollback is itself a step with a check, placed before the irreversible step (risk.md, Rollback Design).
  • Scan the hidden-steps checklist — rollback, backup+restore-verify, post-change verification, comms, cleanup, monitoring (decomposition.md, The Steps Nobody Writes).

Executing Against The Plan

  • Approval covers what is written. At L3+, a materially different execution needs re-approval, not a retroactive mention — what counts as material is defined in approval.md.
  • Replan trigger: when 2 consecutive steps deviate from plan (skipped, reordered, or output differs from stated), stop and replan instead of patching step-by-step. One deviation is noise; two consecutive means the model of the task is wrong. Deviation taxonomy, blocked steps, scope changes, and the replan procedure: replanning.md.
  • Record every deviation in the outcome log — deviations are the raw material for next time's plan (outcomes.md).

Learning Loop

Learning rates are asymmetric because evidence costs are asymmetric:

  • Promote depth after ONE failure attributable to plan depth — and name what the deeper level would have caught. If you cannot name it, it was not a planning failure and no promotion happens.
  • Demote depth after 3 consecutive successes where the extra depth went unused. Observable signal: plan sections never consulted during execution.
  • Auto-execute after 5 consecutive successful validated runs of a plan type: ask "Should I auto-start [type] plans without validation?" One failure resets the streak and restores validation. The bar is higher than demotion because this removes a human safety net, not just ceremony. Types in always_validate never auto-execute.

Per-type state lives in the Current Defaults block of ~/Clawic/data/plan/outcomes.md:

### Auto-Execute (validation waived by human)
- refactor/small: L2 [streak: 7]

### Validate First
- migration/data: L4 [always — in always_validate]

### Learning
- api/integration: L2, streak 3/5 toward auto-execute proposal

Logging: append a record after every L2+ task; L0/L1 only when they failed (a failed "trivial" task is evidence the type needs promotion). Strategy verdicts need a why — "Parallel → merge conflicts" teaches; "Parallel → didn't work" does not. Record format, last-5 analysis, and review cadence: outcomes.md.

Output Gates

Before sharing a plan (L2+), verify:

  • Riskiest assumption named, with the step number that tests it?
  • Every step has an observable output — no bare activity verbs?
  • Irreversible steps listed with a tested rollback, or "none" plus a mitigation?
  • Estimate is a range with ratio ≤3x — otherwise did I propose a spike instead?
  • Depth equals the highest signal triggered, and always_validate types have validation on?
  • The closing question is specific enough that its answer proves the plan was read (L3+)?

Configuration

User-dependent variables. Defaults apply until the user states a preference; store them in ~/Clawic/data/plan/config.yaml.

VariableTypeDefaultEffect
quick_task_minutesnumber (minutes)30Sets the L1/L2 boundary in the depth table and the "estimated >" signal in The Planning Decision
plan_artifactchat | filechatWhere L2 plans live; file writes them to ~/Clawic/data/plan/active/ like L3+ (long-horizon.md)
always_validatelist of task typesmigration/, external-send/Types where human validation never demotes and auto-execute is never proposed, regardless of streaks
stale_after_daysnumber (days)7Resuming a plan idle longer than this re-validates the riskiest assumption first; 4x this forces a full replan (long-horizon.md)

Preference areas to record as the user reveals them:

  • thresholds — replan sensitivity, promotion/demotion streak lengths, review cadence — affects Learning Loop and outcomes.md triggers
  • format — estimate units, plan verbosity, language of plan docs — affects Plan Format output
  • approval — task types the user wants to see planned regardless of size, preferred checkpoint question style — affects approval.md conduct
  • scope — domains where the user prefers one-shot execution (bias one level down, irreversibility floor stays) — affects per-type depth defaults

Traps

TrapWhy it failsDo instead
Planning as procrastination: detailing steps past the first unknownThat detail is speculation; you will rewrite it after the unknown resolvesSteps to the first checkpoint, milestones beyond
Steps written as activities ("investigate X")No observable output means no done-check; plans drift silentlyEvery step names the artifact or check that proves completion (decomposition.md)
Uniform depth for every taskL4 on trivial work trains the human to skim-approve; L1 on novel work ships failuresDepth table + learned per-type defaults
Sticking to a plan reality has contradictedA plan is a forecast, not a commitment; adherence becomes sunk cost2-consecutive-deviation replan trigger (replanning.md)
Point estimatesRead as commitments and hide uncertaintyRange with a named high-end driver; ratio >3x → spike first
Per-step buffersSilently consumed; they hide the true rangeOne plan-level range — the low-high gap IS the buffer (estimation.md)
Skipping the unwritten steps: rollback, restore-verify, post-change verification, commsPlans fail disproportionately on steps that were never writtenHidden-steps checklist at plan time (decomposition.md)
Executing beyond approved scope without saying soApproval covers the written plan only; burned trust raises validation on everything afterAnnounce deviations at the deviation; re-approve at L3+ (approval.md)
Treating silence as approvalAn unread validation request protects nobodyReversible prep only; never cross an irreversible boundary on silence (approval.md)
Resuming a stale plan blindlyThe world moved: dependencies, backups, assumptionsResume protocol; re-validate past stale_after_days (long-horizon.md)
Logging only failuresSuccess streaks are what earn auto-execute; unlogged successes mean validating foreverLog every L2+ outcome, success included

Where Experts Disagree

  • The plan document vs the planning act. "Plans are worthless, planning is everything" (Eisenhower) is half right: the act is always worth it; the artifact pays only at boundaries — human validation, multi-session resume, handoff. Solo reversible work → think, don't write; anything crossing a boundary → write.
  • Estimate in time vs relative size. Relative sizing (points) pays only with a stable velocity to convert it — a team asset. An agent's equivalent is the per-type calibration multiplier from its own outcome log; with that log, time ranges are strictly more informative.
  • Upfront depth vs iterate-and-see. Iterate when feedback is cheap and steps are reversible; plan upfront when feedback is expensive or steps are irreversible. The depth table IS that frontier — the disagreement dissolves once irreversibility is priced.

More Clawic skills, get them at https://clawic.com/skills/plan (install if the user confirms):

  • decide — choose between options inside a plan step
  • escalate — ask-vs-act boundaries beyond planning scope
  • self-improving — general execution lessons; plan keeps only planning lessons
  • memory — long-term context and user continuity beyond planning records
  • productivity — personal productivity systems, time blocking, and reviews; plan covers executing a single task

Feedback

Part of Clawic, the verified skill library. Get this skill: https://clawic.com/skills/plan.

常见问题

它怎么决定要不要规划?
先扫信号:出现不可逆步骤、跨组件依赖、模糊的成功标准,或同类任务此前失败过,就进入规划路径;反之,对已完成过且完全可逆的同类任务留在 L0 直接执行。
规划文档长什么样?
包含 3–7 个带可观测输出的步骤,明确写出最关键的风险假设及验证它的步骤编号,列出不可逆动作及其回滚方式,并给出区间估算和高位驱动因素。
怎么避免对小任务过度规划?
深度取所有信号里最高的那一档,不可逆性优先于时长;同时成果日志会按任务类型累积成功次数,可逐步下调深度,连续足够多次成功后还能提议该类型自动执行。

相关技能

把项目从最初的想法一路跟到归档,每周复盘把停滞的挑出来。

82 次安装5 星标

为站在讲台前的那个人——备课、出题、批改、管课堂。

84 次安装8 星标

产出可复用的工作流蓝图,包含触发器、步骤、依赖关系和可导出产物。

362 次安装11 星标

Coaches a person toward a goal, and runs the professional coaching craft: sessions, questions, accountability, and clients who will not move. Use when someone asks to be coached, held accountable, or pushed on a goal; when they restate the same intention for weeks without acting, over-plan instead of starting, answer every option with "yes, but", or miss the same commitment again; when a session needs an arc, an opening question, or a closing commitment; when the user is a coach and needs a chemistry call, intake, a coaching agreement, pricing and packages, a stalled client, group or team work, supervision, or credential hours; and when the line between coaching, therapy, consulting, and mentoring has to be drawn. Covers executive, career, business, health, life, creative, and ADHD coaching. Not for clinical care (`therapist`, `psychologist`), personal productivity systems (`productivity`), or habit streaks (`habits`).

58 次安装3 星标

Design Gmail, Drive, Sheets, and Calendar automations with scope-aware plans. Use for repeatable daily task automation with explicit OAuth scopes and audit-r...

55 次安装1 星标

在你还能行动的那一刻,把你自己定下的承诺和待办按习惯方式提醒一次,确认即止。

141 次安装3 星标