The problem: "done" is a claim, not a fact
Give several agents a plan and they report progress in prose. Prose is cheap. An agent can say it finished while the tests fail, quietly edit files outside its task, or stall with a claim nobody releases. Asking another model to check the first one just moves the problem.
pytest tests/auth exit 0 in the task's worktree, and did only
src/auth/** change?ALE's rule: anything code can decide, code decides. Models plan and do the work. Deciding whether that work is accepted is left to code: exit codes, file paths and timestamps.
How it works
A plan becomes labels, labels become a run, and every step of the run is an event in an append-only log.
Bake
A literal parser turns Files:, Run: and Depends on Task N lines into labels. It never guesses; gaps are listed for the planner to fill.
Dispatch
Ready tasks get a git worktree and branch each, and the executor your roster picks for the label's role and tier.
Verify
The only path to accepted. It runs the label's acceptance commands itself and checks changed paths against the label.
Integrate
Executors never commit. ale integrate commits only the allowed changes and merges the task branch.
Anatomy of a label
A label is a small JSON document. It is the contract between the orchestrator, the executor and
the verifier. This is an abridged copy of examples/run/labels/T01.json.
{
"task_id": "T01",
"title": "Add token refresh endpoint",
"labels": {
"role": "backend", "model_tier": "standard",
"lane": "pane", "risk": "low",
"effort": "M", "locality": "any"
},
"context": {
"spec_path": "specs/T01.md",
"pointers": ["src/auth/session.py"],
"allowed_paths": ["src/auth/**", "tests/auth/**"],
"depends_on": []
},
"acceptance": [
{"id": "A1", "cmd": "true", "expect": "exit0"},
{"id": "A2", "cmd": "test -d .", "expect": "exit0"}
],
"provenance": {
"lane_reason": "Runs about 40 minutes
unattended, so pane."
}
}
- labels
- A closed vocabulary from your roster.
role+model_tierpick the executor and model; changing models is a one-line roster edit. - lane
inline,workfloworpane. Always the planner's call, with a writtenlane_reason. No classifier chooses it.- allowed_paths
- The files the task may change. Hooks deny edits outside them;
ale verify --basechecks the final diff. - depends_on
- Tasks that must be accepted first.
ale readyonly offers tasks whose dependencies are done. - acceptance
- 2 to 5 shell commands.
expectis an exit code and nothing else:exit0orexit:N.
The board and the watchdog
The board is one file per run, events.jsonl, append only. Status is computed by replaying it,
so there is no state to corrupt and every decision has a line you can point at. Executors may only
write their own events (claim, heartbeat, submit, input-required, note) on tasks they own. Only the
system writes accepted, rejected and lease_expired.
A claim is a lease: it lives while heartbeats arrive. ale watchdog has no model in it. It reads the log and reports:
| Breach | Meaning |
|---|---|
lease_expired | No heartbeat within heartbeat_timeout_s. The claim is released for someone else. |
stuck | Claimed, but no progress past stuck_after_s. |
overrun | Running past max_duration_s. |
unverified | Submitted, but nobody ran ale verify in time. |
| attempts | More attempts than max_attempts. |
Every command reports through its exit code, so scripts and hooks can branch on it:
| Exit | Meaning |
|---|---|
0 | ok |
1 | check failed (for example, acceptance) |
2 | usage error |
3 | claim lost: someone else owns the task |
4 | lease lost: stop writing immediately |
5 | needs human sign-off |
6 | watchdog found breaches |
To watch a run live: ale status, ale timeline, ale meta for token usage, or the web board with ale board --open.
Quick start: watch it refuse a false "done"
No agent needed. You play both roles in a throwaway repo. Install the CLI first (below), then:
- Write a two-task plan and bake it.
mkdir /tmp/ale-demo && cd /tmp/ale-demo git init -q && printf '.ale/\n' > .gitignore && git add .gitignore && git commit -qm init ale setup # write PLAN.md: copy it from the README's quick start ale plan bake PLAN.md --write # exit 1: lists lane / lane_reason gaps
- Fill the planner-only fields and commit.
sed -i.bak 's/"lane":null/"lane":"inline"/; s/"lane_reason": null/"lane_reason": "Small and watched, so inline."/' PLAN.md && rm PLAN.md.bak ale plan bake PLAN.md && git add -A && git commit -qm plan
- Start the run and give T1 a worktree.
ale init-run --plan PLAN.md --set-current ale dispatch --no-exec > .ale/dispatch.jsonl AGENT=$(python3 -c "import json; print(json.loads(open('.ale/dispatch.jsonl').readline())['agent_id'])") WT=.ale/runs/PLAN/wt/T1 - Claim done without doing anything.
ale claim --task T1 --agent "$AGENT" ale submit --task T1 --agent "$AGENT" --summary "done" ale verify --task T1 --cwd "$WT" # acceptance failed: A1, A2The task is now rejected. The agent's word counted for nothing. - Do the work and verify again.
mkdir -p "$WT/src" && printf 'def greet(name):\n return "Hello, %%s!" %% name\n' > "$WT/src/greet.py" ale reopen --task T1 --reason "greet.py written" ale verify --task T1 --cwd "$WT" # exit 0 ale integrate --task T1 # commits allowed changes, merges the branch
Now it is accepted, and T2, which depended on it, is ready.
With real agents, ale run PLAN.md (or /label-layer PLAN.md in Claude Code) does steps 3 to 5
for every task, opens focused fix tasks for rejections, and stops for you when it needs a decision.
Install
Needs Python 3.9+, git, and macOS or Linux. The package bundles the agent catalog and launchers.
python3 -m venv ~/.venvs/ale && . ~/.venvs/ale/bin/activate
pip install git+https://github.com/yodem/agentic-label-engineering.git
ale --help
# or: uv tool install git+https://github.com/yodem/agentic-label-engineering.git
/plugin marketplace add yodem/agentic-label-engineering /plugin install ale@agentic-label-engineering
Adds the executor hooks (path guard, auto heartbeat, stop gate), the /label-layer and /ale:board skills, and the /ale-board Mod, which also needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. Install the CLI too.
pi -e adapters/pi/ale.ts # one run cp adapters/pi/ale.ts ~/.pi/agent/extensions/ # every run
The extension is active only when ALE_TASK is set. See adapters/pi.
Copy adapters/codex/hooks.json into your Codex hook configuration and set ALE_PLUGIN_ROOT to your clone. The codex-exec executor wraps codex exec --json with bin/ale-exec for heartbeats, the exit gate and usage.
What each harness enforces
| Rule | Claude Code | Codex | Pi | ale-exec wrapper |
|---|---|---|---|---|
| Path guard | enforced | enforced | enforced | verify only |
| Auto heartbeat | enforced | no | enforced | enforced |
| Stop / submit gate | enforced | no | advisory | enforced |
| Usage capture | enforced | advisory | enforced | enforced |
| Session context | enforced | no | enforced | no |
Trust boundary
- Agent ids are self-asserted. The log stops accidents and honest mistakes, not a malicious local process forging events.
- Shells can write anywhere. Hooks guard edit tools.
ale verify --basechecks the final diff and is the backstop. - Local filesystems only. The append guarantee does not hold on network mounts.
- The judge is optional and off. ALE can collect shadow votes from an external judge command. They never change a decision, and nothing leaves your machine unless you turn it on. See judge.md.
Where to go next
Label layer
Plan format, every label field, worktrees, fix tasks and the run loop.
Executor protocol
The ten rules an executing agent follows.
Agents
The role taxonomy and how a label resolves to an agent definition.
Hooks
What is enforced where, and what hooks cannot do.
Judge
The optional shadow judge and its JSON contract.
Harness facts
Hook behaviour verified against Claude Code, Pi and Codex.