Agentic Label Engineering · ALE

No agent gets to mark its own work done.

ALE coordinates coding agents working from one plan. Every task carries a typed label that says who does it, which files it may touch, and which commands prove it is finished. Code decides the rest: claims, leases, acceptance and alarms.

The problem: "done" is a claim, not a fact

Give several agents a plan and they report progress in prose. Prose is cheap. An agent can say it finished while the tests fail, quietly edit files outside its task, or stall with a claim nobody releases. Asking another model to check the first one just moves the problem.

What the agent says"Implemented the refresh endpoint, all good."
What ALE checksDid pytest tests/auth exit 0 in the task's worktree, and did only src/auth/** change?

ALE's rule: anything code can decide, code decides. Models plan and do the work. Deciding whether that work is accepted is left to code: exit codes, file paths and timestamps.

How it works

A plan becomes labels, labels become a run, and every step of the run is an event in an append-only log.

Bake

A literal parser turns Files:, Run: and Depends on Task N lines into labels. It never guesses; gaps are listed for the planner to fill.

Dispatch

Ready tasks get a git worktree and branch each, and the executor your roster picks for the label's role and tier.

Verify

The only path to accepted. It runs the label's acceptance commands itself and checks changed paths against the label.

Integrate

Executors never commit. ale integrate commits only the allowed changes and merges the task branch.

Anatomy of a label

A label is a small JSON document. It is the contract between the orchestrator, the executor and the verifier. This is an abridged copy of examples/run/labels/T01.json.

{
  "task_id": "T01",
  "title": "Add token refresh endpoint",
  "labels": {
    "role": "backend", "model_tier": "standard",
    "lane": "pane", "risk": "low",
    "effort": "M", "locality": "any"
  },
  "context": {
    "spec_path": "specs/T01.md",
    "pointers": ["src/auth/session.py"],
    "allowed_paths": ["src/auth/**", "tests/auth/**"],
    "depends_on": []
  },
  "acceptance": [
    {"id": "A1", "cmd": "true", "expect": "exit0"},
    {"id": "A2", "cmd": "test -d .", "expect": "exit0"}
  ],
  "provenance": {
    "lane_reason": "Runs about 40 minutes
                    unattended, so pane."
  }
}
labels
A closed vocabulary from your roster. role + model_tier pick the executor and model; changing models is a one-line roster edit.
lane
inline, workflow or pane. Always the planner's call, with a written lane_reason. No classifier chooses it.
allowed_paths
The files the task may change. Hooks deny edits outside them; ale verify --base checks the final diff.
depends_on
Tasks that must be accepted first. ale ready only offers tasks whose dependencies are done.
acceptance
2 to 5 shell commands. expect is an exit code and nothing else: exit0 or exit:N.

The board and the watchdog

The board is one file per run, events.jsonl, append only. Status is computed by replaying it, so there is no state to corrupt and every decision has a line you can point at. Executors may only write their own events (claim, heartbeat, submit, input-required, note) on tasks they own. Only the system writes accepted, rejected and lease_expired.

A claim is a lease: it lives while heartbeats arrive. ale watchdog has no model in it. It reads the log and reports:

BreachMeaning
lease_expiredNo heartbeat within heartbeat_timeout_s. The claim is released for someone else.
stuckClaimed, but no progress past stuck_after_s.
overrunRunning past max_duration_s.
unverifiedSubmitted, but nobody ran ale verify in time.
attemptsMore attempts than max_attempts.

Every command reports through its exit code, so scripts and hooks can branch on it:

ExitMeaning
0ok
1check failed (for example, acceptance)
2usage error
3claim lost: someone else owns the task
4lease lost: stop writing immediately
5needs human sign-off
6watchdog found breaches

To watch a run live: ale status, ale timeline, ale meta for token usage, or the web board with ale board --open.

Quick start: watch it refuse a false "done"

No agent needed. You play both roles in a throwaway repo. Install the CLI first (below), then:

  1. Write a two-task plan and bake it.
    mkdir /tmp/ale-demo && cd /tmp/ale-demo
    git init -q && printf '.ale/\n' > .gitignore && git add .gitignore && git commit -qm init
    ale setup
    # write PLAN.md: copy it from the README's quick start
    ale plan bake PLAN.md --write      # exit 1: lists lane / lane_reason gaps
  2. Fill the planner-only fields and commit.
    sed -i.bak 's/"lane":null/"lane":"inline"/; s/"lane_reason": null/"lane_reason": "Small and watched, so inline."/' PLAN.md && rm PLAN.md.bak
    ale plan bake PLAN.md && git add -A && git commit -qm plan
  3. Start the run and give T1 a worktree.
    ale init-run --plan PLAN.md --set-current
    ale dispatch --no-exec > .ale/dispatch.jsonl
    AGENT=$(python3 -c "import json; print(json.loads(open('.ale/dispatch.jsonl').readline())['agent_id'])")
    WT=.ale/runs/PLAN/wt/T1
  4. Claim done without doing anything.
    ale claim  --task T1 --agent "$AGENT"
    ale submit --task T1 --agent "$AGENT" --summary "done"
    ale verify --task T1 --cwd "$WT"   # acceptance failed: A1, A2
    The task is now rejected. The agent's word counted for nothing.
  5. Do the work and verify again.
    mkdir -p "$WT/src" && printf 'def greet(name):\n    return "Hello, %%s!" %% name\n' > "$WT/src/greet.py"
    ale reopen --task T1 --reason "greet.py written"
    ale verify --task T1 --cwd "$WT"   # exit 0
    ale integrate --task T1            # commits allowed changes, merges the branch
    Now it is accepted, and T2, which depended on it, is ready.

With real agents, ale run PLAN.md (or /label-layer PLAN.md in Claude Code) does steps 3 to 5 for every task, opens focused fix tasks for rejections, and stops for you when it needs a decision.

Install

Needs Python 3.9+, git, and macOS or Linux. The package bundles the agent catalog and launchers.

python3 -m venv ~/.venvs/ale && . ~/.venvs/ale/bin/activate
pip install git+https://github.com/yodem/agentic-label-engineering.git
ale --help
# or: uv tool install git+https://github.com/yodem/agentic-label-engineering.git

What each harness enforces

RuleClaude CodeCodexPiale-exec wrapper
Path guardenforcedenforcedenforcedverify only
Auto heartbeatenforcednoenforcedenforced
Stop / submit gateenforcednoadvisoryenforced
Usage captureenforcedadvisoryenforcedenforced
Session contextenforcednoenforcedno

Trust boundary

Labels are code. Acceptance commands run through the shell with your permissions. Only run labels you or your orchestrator wrote.
  • Agent ids are self-asserted. The log stops accidents and honest mistakes, not a malicious local process forging events.
  • Shells can write anywhere. Hooks guard edit tools. ale verify --base checks the final diff and is the backstop.
  • Local filesystems only. The append guarantee does not hold on network mounts.
  • The judge is optional and off. ALE can collect shadow votes from an external judge command. They never change a decision, and nothing leaves your machine unless you turn it on. See judge.md.