Featurly — feature-dev run bilans
2026-08-10 — v2.1.0 teach-back contract: ZERO questions on a clear mission, feature shipped to spec (run 019fec4c)
- Status: validated — the non-regression half of the teach-back dogfood (its ambiguous-mission twin is whole-improve-loop run 019fec3e, same day): a clear mission must produce ZERO
ask_user_asyncposts, and did. - Versions: bot 2.1.0 (first run with
interaction: async+ teach-back item 5) · iterion 641839340 - Method: claude_code/claude-opus-5, effort high, sandbox auto,
--max-cost-usd 5 --max-duration 30m, a deliberately unambiguousfeature_prompt(argparse CLI, named files, named exit code) against theparcel-toolsPython fixture. - Result:
finishedin one pass, 5m25 wall, campaign $0.88 + review $0.77;feature_complete=true, 2 commits, squash-merged (5741468);human_input_requestedcount: 0. - Value: the feature landed exactly to spec —
cli.pywith both subcommands,tests/test_cli.pyin the repo's style, README section; smoke-verified live (total 1.10 2.20→ 3.30,freight 2 3.5→ 7.00,total abc→ stderr message + exit 2). - Findings / misses: none for the contract under test. Still unexercised across both runs: mid-run answer folding (both runs finished before any answer; ADR-081 engine machinery has its own coverage — prove it live on the next organically-ambiguous run) and the "first pass only" re-post guard (both converged in ONE pass, so no pass 2 existed to tempt a re-post).
- Lessons for next run: pair every prompt-contract change with this twin-run shape — one run that must trigger the new behavior, one that must NOT; the second is the regression detector and it is the cheaper of the two.
2026-07-12 — #125 "board as pipeline projection" delivered by Featurly dogfood (3 cloud waves + 2 manual)
- Status: validated — full #125 vision shipped & deployed. The board-as-pipeline vision (ADR-071, additive) was implemented end-to-end mostly by Featurly's own cloud runs, dispatched from GitHub issues labelled
implement+forfait-run(org OAuth forfait). #125 is CLOSED; 11 vision PRs merged ontomain@:edge. - Versions: bot feature-dev v2 (ADR-058 campaign) · iterion
:edge(this session's fixes) · cloud runner on ovh-prod. - Method: issue →
webhooks_common.gofeatureDevBotID routesimplementto Featurly → board card → cloud coordinator launches campaign → back-linked PR. Each wave's specs were authored as precise file:line + plan; Featurly implemented, tested, sometimes above spec. - Value produced (what Featurly shipped, merge-worthy):
- Wave 1: #152
iterion issue import(self-hosted forge→board sync, T6) · #153 board grouped/scoped by bot (pipeline lens) · #154 card run-history (Issue.Runs[]). - Wave 2: #160 persisted bot filter (
View.Bot) · #161 per-card ⏸ Awaiting-input badge (denormalized) · #162 run-tree reverse queries + shard-tuple projection (T4b) · #159 forge GitHub-App 422 → mark connection degraded, stop the re-mint loop. - Wave 3: #169 render the run tree (shard/fork children) in run console + card (T4b).
- Wave 1: #152
- Manual pieces (I finished these 2; the prod forfait credit ran out mid-campaign): #173 the generic "Awaiting input" column (dispatcher parks a paused card there:
board.gostate +commands.go/boarddispatch.gorouting + studio colour) · #174 the verify-build codegen-gate skill fix (below). - Dogfood-found engine bugs (fixed same session):
- PauseForm vs HumanPromptForm — answer-from-board rendered checkpoint context vars, not output-schema fields → resume would route to fail. Fixed to reuse
HumanPromptForm(#134). - WS "disconnected" banner on paused runs — broker closes the stream on pause → spurious banner + reconnect loop. Gated the banner on
running|queued(#144). - verify-build didn't run codegen gates (→ #167/#174): the skill already had §1b ("mirror CI's openapi:check / chart drift"), but its example
verify.shshowed only build+test, and agents copy the example — so #159/#162 shipped OpenAPI drift CI caught. Fix: put<regen> && git diff --exit-codein the example, make §1b imperative, unify the 2 skill variants across all 9 bots. Deterministicverify_run-level enforcement noted as a follow-up. - runner-restart drains in-flight cloud runs —
rollout restart deploy/iterion-runnerduring a deploy requeued live Wave-2 runs. Operational lesson saved to memory; deploy onlydeploy/iterion, never-runner.
- PauseForm vs HumanPromptForm — answer-from-board rendered checkpoint context vars, not output-schema fields → resume would route to fail. Fixed to reuse
- Findings / misses: Featurly is reliable when the spec is precise (file:line + a plan) — correct, tested, occasionally above spec. It does NOT self-discover a repo's codegen freshness gate unless the skill's copyable example carries it (#174 is the fix for exactly that). Two deferred #125 tensions (T4 ticket↔runs 1:N tree, T6 self-hosted forge import) both landed in this campaign (#154/#162/#169 and #152).
- Lessons for next run: (a) hand Featurly the gate commands in the skill example, not just in prose; (b) contain forfait spend — a long multi-wave campaign can deplete org credit mid-flight (finish stragglers from a credited local session); (c)
--merge-into none+post_to_board=falseon every dogfood launch kept the operator's tree clean. - Réserves closed (2026-07-14 follow-up): awaiting-input validated live on a local
iterion dispatch(legacy board auto-migrated, pause → card into the column, CLI resume → newreconcileParkedsweep moves it to review) — the validation surfaced 3 engine gaps (no board migration for existing stores, parked cards stranded after out-of-band resume,iterion dispatchre-run exiting on "already running"), all fixed; and the verify-build drift gate is now deterministic (CI-mirror assertion + post-verify tree-drift check in every verify_run body, all 9 bots), not just skill prose.
2026-07-09 — first CLOUD runs: label an issue → implement → open a PR (runs 019f4590 / 019f45dd / 019f45f6)
- Status: validated end-to-end on prod. Labeling a GitHub issue
implementroutes to Featurly (fix2281a2eff— a 3-bot webhook used to send it to review-pr), which implements the feature on the devbox runner and opens a back-linked PR — which then auto-triggers Billy. The full loopissue → Featurly → PR → Billyran live. - Versions: iterion
:edge(fixes shipped this session, chart 0.34.0) · runneriterion-runner-devbox:edgeuid 1000 · webhookd291059con SocialGouv/iterion. - Value produced (real, merge-worthy commits):
- #86 → Featurly threaded a
*log.Loggerthrough the whole generic-secret resolution path so silent credential drops become greppable (the erreurs-explicites gap the debugging itself surfaced). - #88 → PR #89
feat(report): surface resolved verify command in report head.
- #86 → Featurly threaded a
- Three infra gaps found + fixed before the PR could open (all in the board-launched path — an issue-labeled delivery makes a board card the coordinator launches, NOT the direct webhook path):
- routing (
2281a2eff) — issue → implementer, not reviewer; - forge token binding (
483e69a3f+9b339c999) — the coordinator resolves secrets by (tenant, bot) binding, which forge provisioning didn't create (only a webhook override, which the board path never sees); - writable secret mount (chart 0.34.0) — the uid-1000 runner couldn't
mkdir /run/iterionto materializeforge_tokenin-pod (root-owned/run); a Memory-backed emptyDir fixes it.
- routing (
- Findings / misses: Featurly's implementation quality was high (correct diff, tests, self-check for parallel issues) — the gaps were all platform plumbing, not the bot. The PR is opened under the connection identity (devthejo), not a bot identity — the GitHub-App migration (deferred) fixes that.
- Lessons for next run: an issue-labeled re-trigger needs a FRESH issue (re-applying the same label is an idempotent replay); a PR re-trigger needs a new head sha + close/reopen (synchronize alone isn't in the review-pr action set opened/reopened); deploying a chart mid-run (ArgoCD sync) orphans the in-flight run (
process orphaned: server restart) — a real but benign infra race, resumable/re-triggerable.
Autonomous end-to-end feature development. v2 (ADR-058 minimal-framing) since 2026-07-07: one campaign agent ships the feature slice by slice (commits in stride) against a deterministic build/test gate + bounded continuation loop, opt-in MR tail. Pre-v2 bilans below describe the retired plan → act → /simplify → alternating review-fix → commit pipeline. See bots/feature-dev/.
2026-07-07 — v2 minimal-framing PILOT: iterion validate --strict shipped in one pass (run 019f3bb4)
- Status: VALIDATED — the ADR-058 conversion's dogfood gate. First live run of the v2 shape (campaign → verify_build → verify_run → gate → mr_gate → done): converged on the FIRST pass in 11m33s end-to-end (sandbox setup included), 64 tool calls, 2 LLM nodes. The v1 shape's live e2e expectation for a comparable feature was 40–70 min across ~8 nodes and 15 possible review iterations.
- Versions: bot feature-dev v2.0.0 (
11aa9b65b) · iterion worktree branch @a497f1ae2· sandbox-full:edge. - Method: CLI
iterion runfrom the conversion worktree (static binary),--store-dir <workspace>/.iterion(studio-visible),--merge-into none,--max-cost-usd 30 --max-duration 2h, mono claude (opus-4-8 high, Claude forfait — no cost metric reported). feature_prompt = add--stricttoiterion validate(warnings fail the exit code for CI gating) +--jsoncoverage + table-driven tests + docs. - Result:
finished,gate.converged=true(real verify_run: passed, not skipped), 2 semantic commits in stride on storage branchiterion/run/thunder-sizzle-dawnglyph-ec4d@0cd39d446:feat(cli): add --strict to iterion validate to fail on warningsthendocs(cli): document iterion validate --strict flag. Termination contract honest:feature_complete=true, commits_this_pass=2, summary matched the diff. - Value: real, wanted CLI feature. Functional proof from the storage branch:
TestValidate_Strict+TestValidate_StrictHumanOutputPASS (existing validate tests untouched-green); behaviour matrix exercised live — 0-warning file +--strict→ exit 0; warning-bearing file without--strict→ exit 0 (unchanged default); with--strict→ exit 1 withresult: INVALID (strict: N warning(s)). - Findings / misses: (1)
validate --jsonemits TWO JSON documents on failure (result object + error object) — pre-existing, verified against a pre-feature binary; candidate board finding, not a regression. (2) The campaign did not exercise the ask-user or findings-handoff paths (nothing to escalate on a clean feature) — those paths remain structurally tested only. (3) Run-worktree was based on the MAIN checkout HEAD (engine derives the repo from the store-dir side), not the launching worktree's branch — harmless here, worth knowing when dogfooding from a worktree. - Engine hardening: none needed — sandbox, verify gate, loop wiring and finalize (storage branch, no FF) all behaved first try.
- Lessons for next run: the goal-directed v2 contract holds as-is; keep
reasoning_effort: high(the shape, not the tier, carried it); the verify_build node re-derivedverify.shcorrectly from devbox — no skill gap. GATE PASSED → rollout continues (docs-refresh, adr-cartograph, rgaa-audit, dep-update-guard, secured-renovacy P2, sec-audit-source remediation).
2026-06-25 — issue-comment → MR cloud-k8s e2e GREEN (Claude forfait + GLM; 8 more fixes) (run 019efbc6)
- Status: VALIDATED — the full feature works end-to-end on the deployed preprod (ovh-dev) k8s runner.
/featurlyon project 194 issue !2 → claude_code-forfait implementer wrotemain.go→ claw GLM reviewer approved → commit → push → MR !13 opened (iterion/fix-go-doc-comment → main, authorproject_194_bot) → back-link comment on issue !2 ("Created MR: …/merge_requests/13"). The 8-fix infra campaign below got the sandbox to start; these 8 more drove the run from "empty workspace" all the way to the MR + back-link. - Versions: iterion main
1712e4e56 → 31aa3636a(8 commits) ·sandbox-full:edge· chart unchanged. - Method: live
/featurlywebhook → preprod runner → board-dispatched run → k8s sandbox. Creds: Claude forfait (CLAUDE_CODE_OAUTH_TOKEN) + GLM (ZAI_API_KEY) only — no OpenAI/codex (the user's constraint, which sidesteps the~/.codexforfait gap flagged in the prior bilan). claude_code nodes (plan/act/simplify/reviewer_claude/prepare_commit/finalize_mr) use the forfait; the clawreviewer_gptusesanthropic/glm-5.2because claw can't use the Claude Code OAuth forfait (Consumer Terms). Provisioned toiterion-llmviakubectl patch(SealedSecret → temporary; scrubbed at end via runner rollout-restart).
The 8 fixes (each surfaced by the next /featurly, continuing #1–8 below)
- workspace-delivery to the k8s sandbox (the substantial Phase-5-V2 piece the prior bilan deferred) —
/workspaceis an emptyDir with no host bind-mount, so the runner's clone never reached the bot → it ran in an EMPTY workspace.kubernetes/driver.go populateWorkspace: after the pod is Running, tar-streamresolveCloneRoot(WorkspacePath)(git rev-parse --git-common-dir→ the clone root, kept with.gitso the bot can commit+push from inside) into the pod viatar -cf - … | kubectl exec -i -- tar -xf -.1712e4e56. - RepoURL for issue-comment runs — the GitLab issue-note webhook carried no clone URL/ref. Now sets
CloneURL = git_http_url(present in the payload) +DefaultBranchas the ref.748050fd9. - board-dispatcher dropped RepoURL (the real RepoURL gap) — a board-mode
/commandwith the dispatcher active doesensureBoardCard(StateReady)+ returns; the CloudBoardCoordinator (processBoardCard) then launches withVars: iss.BotArgsonly — no RepoURL → the runner never cloned. Fix: stamp clone-url/ref into the card's BotArgs under reserved keys (__iterion_repo_url/_ref), lifted intoLaunchSpec.RepoURL/RepoRefbyliftBoardRepo.adf49681d. - tar workspace-copy perms — the archive's
./root member made the in-pod tar chmod/utime the root-owned, fsGroup-setgid/workspaceemptyDir (a non-root user can't) →exit 2, failing the whole run though every file extracted fine.--no-overwrite-dir(48cf5f1c1) was INSUFFICIENT; v2 =os.ReadDir+ tar entries BY NAME, so the archive has no./member at all.1c3925437. - GLM claw reviewer creds in-sandbox —
providerCredentialEnvVars/forwardableProviderEnvnever forwardedZAI_API_KEY/ANTHROPIC_AUTH_TOKEN/ANTHROPIC_BASE_URLto the sandbox__claw-runner→ the GLM reviewer had no z.ai creds in-container. Add the 3 vars;registry.gosynthesises the z.ai bearer +ZAIDefaultBaseURLfromZAI_API_KEY. (Config:ITERION_VIBE_MODEL_GPT=anthropic/glm-5.2— theanthropic/prefix is required; claw rejects a bareglm-5.2withinvalid spec.)aa914dd42. - k8s workspace mount path — the driver mounted the workspace at
/workspace, but the bot's deterministic tool nodes (commit_changes/finalize_mr) + PROJECT_DIR use the worktree ABSOLUTE path →git -C /home/iterion/worktrees/<id>=exit 128: No such file or directory, killing the run right after the (passing) review loop. Overridep.workspace = info.WorkspacePath(mirror docker's bind-at-host-abs-path so mount/workingDir/populate/cwd all align).a24511ff6. - pods-patch on resume — a runner-level resume re-applies the pod;
kubectl applyPATCHes the (largely immutable) pod, and the SA intentionally lacks pods/patch → Forbidden → DLQ. Force-delete (--grace-period=0 --force) any stale pod before apply so it always CREATEs fresh.a24511ff6. - git author identity —
git commitin the sandbox:Author identity unknown / unable to auto-detect email— k8s has no~/.gitconfig(the bind-mount is dropped, and the cloud runner pod has none of its own). Seed the clone's LOCAL git config (user.name/email) in the runner right aftergit clone(travels into the sandbox with.git; defaultiterion/iterion@users.noreply.github.com, overridable viaITERION_GIT_AUTHOR_*).31aa3636a.
Result + lessons
- Converged on review iteration 0 (a one-line doc-comment feature). The bot improved the comment to proper Go
Package maindoc-comment form during finalize_mr, committed oniterion/fix-go-doc-comment, pushed (auth via the token-injected remote URL), andglab mr created MR !13 + back-linked issue !2. forge_token resolves on the board path via the org bot-secret binding (Tier1/Tier2) — no SecretOverrides threading needed, and it mounts at/run/iterion/secrets/forge_token. - The whole cloud-k8s sandbox path now works for a commit+MR bot on Claude-forfait + GLM. The forfait drives claude_code; GLM (z.ai) drives the claw reviewers (claw can't use the forfait).
- Test-prompt gotcha: revi-playground has only
README.md/go.mod/main.go— a prompt referencing a non-existent file (fetch.go) makes the bot create-or-confuse; prefer an existing-file modify (main.go). - feature-dev noise:
act'sgit add -Astaged the compiledrevi-playgroundbinary +.claude/plan.md; reviewers saw them in the working diff (didn't block, but a.gitignore/curatedgit addwould be cleaner). prepare_commit's curated[main.go]kept the commit clean. - glab version drift: finalize_mr's first
glab auth set-token --hostfailed (flag renamed); the adaptive agent recovered withglab auth login --hostname. The forge-mr-create skill could pin the modern syntax. - Every fix is k8s-only or additive → docker (web-local/desktop) preserved by construction. Engine hardening on
origin/main:1712e4e56 748050fd9 adf49681d 1c3925437 aa914dd42 a24511ff6 31aa3636a.
2026-06-24 — k8s cloud-sandbox hardening campaign (8 fixes; feature-dev = first sandboxed bot on k8s)
- Status: partial — the sandbox infrastructure is now validated end-to-end on the deployed preprod (ovh-dev) k8s runner: each
/featurlyon project 194 issue !2 advanced the run one stage, and it now reaches post_create executing inside the sandbox pod (pod created, baked image pulled, pod Ready). The bot still can't complete because preprod'siterion-llmsecret has no LLM creds (placeholder) — a green run (bot implements + opens the MR) needs a creds decision, below. This is the FIRST sandboxed bot ever run on the kubernetes sandbox driver, so the path was entirely unexercised; 8 durable engine/chart/image fixes resulted. - Versions: iterion main
89ed642dc → c4364319f· chartv0.17.2. - Method: live webhook (
/featurlycomment) → preprod runner → k8s sandbox (sandbox-full:edge, user 1000:1000). claude_code implementer + claw gpt-5.5 reviewers +forge_token.
The 8 fixes (each surfaced by the next /featurly)
ITERION_POD_IP(downward APIstatus.podIP) — chart 0.16.1; the k8s network proxy needs a routable advertise address. Gated onrunner.sandbox.enabled.- host_state —
ITERION_SANDBOX_HOST_STATE=nonein the configmap (chart) + an engine bug: the cloud runner (pkg/runner/loop.go) readcfg.Sandbox.HostStatefrom the env then dropped it — never wired it to the engine likepkg/cli/run.godoes foriterion run.89ed642dcthreads it throughrunner.Config→ engine opts. - claw binary in-sandbox — sandboxed claw shells to
iterion __claw-runnerin-container; docker bind-mounts the host binary but k8s has no host fs. Bake static iterion into the sandbox images (FROM ${ITERION_IMAGE} AS iterion-bin+ COPY; image.ymlneeds: build) + gate the host bind-mount on a newCapabilities.SupportsHostBindMounts(docker true / k8s+noop false).1ca693d42. - per-run Secrets RBAC — the driver creates per-run Secrets (
forge_tokenas:file+ proxy TLS CA); the runner SA lackedsecrets. chartv0.17.2(fb03b4e5a): get/create/update/patch/delete, no list/watch (least privilege). - drop docker-only host bind mounts — feature-dev's
~/.claudeOAuth mount (type=bind+ the docker-onlyconsistency=key) hard-failed k8s manifest build. The runtime dropstype=bindon a no-host-fs driver (reuseSupportsHostBindMounts;dropHostBindMounts/mountIsHostBind, unit-tested).25b04eb71. - imagePullPolicy=Always for mutable sandbox tags (
IfNotPresentfor@sha256) — a stale node-cached:edgemustn't shadow a fresh CI bake.25b04eb71. - drop ALL host binds — fix 5's filter ran before the mount block, missing the runtime's own optional binds (the bundle mount
/opt/iterion/bots/<bot> → /run/iterion/bundle). Moved it after the block → catches bot mounts + bundle/attachments/run-files. Skills still reach the sandbox via the workspace mirror (<workspace>/.claude/skills).6c794e05d. - claude_code CLI bake — post_create installs the CLI via
sudo npm install -g, but the k8s pod isrunAsNonRoot/allowPrivilegeEscalation=false, so sudo can't escalate → post_create exits 1. Bake the pinned llm-clis (claude-code 2.1.175) into the sandbox slim image + symlinkclaudeonto PATH; post_create'sclaude --versionthen passes, skipping the sudo branch.c4364319f(final image build in flight at time of writing).
Remaining blocker — CREDS (operator decision, not code)
The preprod runner has no LLM creds (ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN / OPENAI_API_KEY / CLAUDE_CODE_OAUTH_TOKEN all empty; iterion-llm is a placeholder — Revi's live e2e ran on ovh-prod, which has real keys). There is also a deeper creds-to-sandbox-user gap: claude_code gets ANTHROPIC_API_KEY forwarded by its delegate into the sandbox exec (works if the runner has a key), but claw (gpt) reads ~/.codex (the ChatGPT forfait), which host_state mounts at the host path while the sandbox runs as devbox (/home/devbox) — there is no .codex bridge like the .claude one. A green "bot-runs-in-sandbox" e2e therefore needs (a) real LLM creds in preprod iterion-llm, and (b) a decision on forwarding provider creds into the sandbox env for both providers (vs per-provider mounts). Both are the operator's call.
Lessons for next run
- feature-dev (any claude_code+claw bot) on the k8s sandbox is unblocked at the infrastructure level — the only remaining gap is creds.
- docker (web-local + desktop) is preserved by construction: every fix is gated on
SupportsHostBindMounts(true for docker) or is additive (image bakes), so docker keeps host bind-mounts (OAuth via~/.claude) unchanged. - Engine hardening lives on
origin/main:89ed642dc 1ca693d42 fb03b4e5a 25b04eb71 6c794e05d c4364319f(+ chartv0.17.2).
2026-06-24 — issue-comment → improvement-MR e2e on preprod (run 019ef703)
- Status: partial — the feature's TRIGGER half validated live on preprod (GitLab issue comment →
feature_devrun launched); the bot run then failed on a pre-existing preprod infra gap, NOT the feature. - Versions: bot feature-dev 0.1.0 (+ new
finalize_mrtail) · iterion preprod:edge@0f1d8d670(this work) · webhook-launched on the cloud runner (sandbox). - Method: shipped the issue-comment→improvement-MR feature (commit
0f1d8d670); provisioned on preprod a gitlab webhook (0b00720c, wildcard) + a distinct project-194 bot PAT +forge_tokenbindings (feature-dev/whole-improve-loop/branch-improve-loop) ondevthejo/revi-playground(194) + GitLab hook #13; posted/featurly add a doc comment to fetch.goon issue !2. - Result: GitLab Note Hook → preprod parsed it as an ISSUE note (the new code — old code dropped issue notes as "not a merge-request note"), resolved
/featurly→ feature-dev, launched run019ef703. The run FAILED at sandbox start:network proxy: kubernetes: ITERION_POD_IP env var is empty; the runner pod manifest must inject it via downward API (status.podIP). No MR (never reachedfinalize_mr). - Value: HIGH for validation — proved the headline path (issue comment → deployed iterion → bot launch) end-to-end on real GitLab + preprod. The MR-generation half was proven separately by a local
finalize_mrmechanics test against 194 (real MR opened + back-linked + cleaned up). - Findings / misses:
- Trigger works on the deployed instance: issue-note parse +
/featurlyroute + bot launch all fire. - Hand-created webhook needs
wildcard_bots: a webhook created viaPOST /api/teams/{id}/webhookswith explicit bot_ids has an EMPTY CommandMap, so/featurlyfiltered as "no command route" until I PATCHedwildcard_bots=true(then the live registry discovery resolves it). Documented, but a footgun — consider having the manual create endpoint compute the CommandMap from bot_ids (parity with the forge orchestrator'sbuildCommandMap).
- Trigger works on the deployed instance: issue-note parse +
- Engine/infra hardening (the real finding): the iterion Helm chart's runner Deployment does not inject
ITERION_POD_IPvia the downward API (status.podIP) → the k8s sandbox network proxy cannot start → EVERY sandboxed bot fails immediately on preprod (ovh-dev / chartiterion0.14.0). Fix: addenv: [{name: ITERION_POD_IP, valueFrom: {fieldRef: {fieldPath: status.podIP}}}]to the runner pod template. Blocks the deployed bot-run→MR e2e until fixed (a kubectl patch was correctly denied as a shared-infra change — needs operator consent). - Lessons for next run: after fixing
ITERION_POD_IP+ thewildcard_botswebhook, the deployed e2e is one/featurlycomment from completing — provisioning was left in place (webhook0b00720c, hook #13, bindings, bot PAT; issue !2). Loop-guard is satisfied (distinct bot PAT identity ≠ commenter).
2026-06-23 — Verified Action recovery ladder, ADR-044 (run 019ef38d)
- Status: validated — converged through the cross-family review loop to
done; deliverable builds + all new tests green (verified independently in the worktree, anti-façade). - Versions: bot feature-dev 0.1.0 · iterion fresh static (campaign HEAD) ·
claude_code/opus ·worktree: auto·--merge-into none. - Method: one large
feature_promptasking Featurly to implement the entire ADR-044 "Verified Action" synthesis (goal+recipe+postcondition+policy quad; idempotent-skip → recipe → self-repair → agent-recovery escalation; postcondition-as-truth; gates stay deterministic) one-shot. Run was cap-interrupted at 08:58 (Anthropic session limit) atreviewer_claude, resumed cleanly from checkpoint on a switched account → finished. - Result: branch
iterion/run/019ef38d-745b…@79f9111ed, 1592 insertions / 22 files: DSL (parser/AST/jsonenc/IR +validate_verified_action.go/unparse/EBNF), runtime (executor_verified_action.go342L ladder +executor_tool.gowiring + engine.go + new event), unit + e2e tests, docs (ADR-044, DSL quickref, CLAUDE.md), and a demo application onsecured-renovacy's commit node.go build ./...+go test ./pkg/backend/model ./pkg/dsl/{ir,parser,unparse}green. - Value: HIGH — delivered a whole engine subsystem (the action-node robustness pattern) in a single supervised run, on the operator's explicit directive. NOT yet merged to main (overlaps the 5 hand-fixes to Renovacy's commit node landed the same day; needs review of
executor_verified_action.go+ overlap resolution). - Engine hardening: the run itself is the proof-of-need for ADR-044 — see the 5 brittle commit-node failures on
secured-renovacythe same day (secured-renovacy.md). - Lessons for next run: feature-dev handles a large multi-layer engine feature well when the prompt carries the full design + an explicit anti-façade done-criterion + a reference to the ADR; resume-from-checkpoint survived a mid-run provider cap with zero lost work.
2026-06-17 — ADR-028 Steps 2-4 dispatcher I/O offload (runs 019ed4cd, 019ed4eb, 019ed51d)
- Status: 2 validated+converged (Steps 2, 3) · 1 implemented+validated+manually-repatriated (Step 4 — bot review loop blocked by a runtime stall, not the code)
- Versions: bot feature-dev 0.1.0 · iterion run-binary fresh static
fe132645· dispatched via the dispatcher (owniterion dispatchdaemon on the operator's repaired config,--no-server, sandboxiterion-sandbox-full:edge,worktree: auto). Each step: isolated ticket with an anti-façade done-criterion, reviewed + race-verified + repatriated before the next. - Result:
- Step 2 (ListCandidates off-actor) — converged ~27 min,
77a2cb80, FF to main.launchDiscovery→cmdCandidates, single-flightdiscoveryInFlight,postCmdshutdown-safe choke point. - Step 3 (finishRun tracker HTTP off-actor) — converged,
a72d40f7, FF.finishPlanvalue-copy; transition-FIRST/Release-LAST to close the re-dispatch window; optimistic-retry-as-guard for the give-up HTTP window (cmdDropRetry). - Step 4 (post-claim UpdateState + workspaces.Create off-actor; Claim stays atomic — the reduced/safe variant, chosen over full optimistic-claim) — implemented + build/vet/gofmt clean + full dispatcher race suite green + 3 anti-façade tests pass, BUT the bot's own review loop could NOT converge:
fix_gpt(sandboxed gpt-5.5 via claw) repeatedly hit "context canceled" at the dispatcher's 10-min stall timeout, looping retry→re-dispatch→stall. I reviewed the uncommitted worktree directly (max rigor), confirmed correctness, and manually repatriated (9b3bd3bd→ cherry-pick70b3d4ed, auto-merged clean over the operator's parallel commands.go bug-sweep).
- Step 2 (ListCandidates off-actor) — converged ~27 min,
- Value: high — ADR-028 Steps 2-4 land; discovery, finishRun, and post-claim dispatch I/O are now all off the actor goroutine (only
RefreshStatesremains, deliberately deferred). - Findings / quality: exemplary anti-façade across all three. The Step-4 standout: it kept Claim atomic, allocated the slot post-claim (
setupPending=true), and guarded both reapers (refreshRunningStates+reconcileStalled) against the setup window — correctly identifying that the off-actorUpdateStatemakes the tracker read RunningState before the entry recordsTransitionedFromState, which would otherwise self-cancel the run. - Engine hardening (the real finding):
fix_gpt/reviewer-fix on sandboxedgpt-5.5(claw) hangs >10 min → trips the dispatcher's 10-min stall timeout → cancel → retry loop, blocking review-loop convergence on a perfectly good change. Runtime issue (sandboxed claw streaming / context on a large review-loop context), not the bot. Relates to the known sandboxed-claw streaming + gpt-5.5-forfait-context work. Worth a ticket. Secondary: the run-status monitor false-terminals on a transient cancel→auto-resume — key on issue-state, not run-status. - Lessons for next run: a review loop stuck on a RUNTIME stall ≠ bad code — validate the worktree directly (build +
-race+ manual review) and repatriate rather than re-dispatching into the same stall. For the riskiest step, the reduced variant (Claim-atomic, offload post-claim) avoids the reserved-before-claim state entirely and was the right call.
2026-06-15 — ADR-028 + Step 1 lock-free dispatcher Snapshot (run 019ecafa)
- Status: validated
- Versions: bot feature-dev 0.1.0 · iterion run-binary fresh static build of main
8477a067(≈ HEAD) - Method: dispatched via the dispatcher (own
iterion dispatchCLI,--no-server, fresh static binary so delegated subprocesses + sandbox mount use current code; workspace store.iterion). claude_code (Opus 4.8) plan/act//simplify/reviewer_claude; claw GPT reviewer_gpt;sandbox: iterion-sandbox-full:edge,worktree: auto. Ticket = ADR-028 body (decomposed I/O-offload roadmap) with an anti-façade Step-1 scope. - Result: converged in one review round — plan → act → simplify → reviewer_claude → streak_check → reviewer_gpt → prepare_commit → commit_changes → done.
finished, ~16 min, ~56k tokens. 1 commit89dd2f57on branchiterion/run/aurora-hunt-prismpunk-01af; repatriated to main by FF (clean — main was ancestor); issueefb9022d→ done. - Value: high. Produced
docs/adr/028-dispatcher-actor-io-offload.md(records the tracker-as-claim-authority insight + per-issue state-machine direction + incremental sequence + rejected alternatives) AND implemented Step 1:Snapshot()is now lock-free (atomic.Pointer[Snapshot], published by the actor infireSnapshot, seeded at construction), so dashboard reads never wait on the actor's in-flight tracker I/O. Removed the now-deadcmdSnapshot. - Findings / quality: exemplary anti-façade output. (1) It refused the nil-fallback to
buildSnapshot()with the correct reasoning that it "would read c.state off the actor goroutine — the very race this read path removes" — i.e. it understood the invariant, didn't just add a field. (2) Scoped strictly to the read path; no out-of-scope dispatch/claim/finishRun changes. (3) The test (TestSnapshotLockFreeWhileActorBlocked) gatesfakeTracker.ListCandidateson a channel, waits until the actor is provably parked inside it, then assertsSnapshot()returns < 500ms with real state — genuinely proves the property, not its shape. Independently re-verified: build/vet OK, race-clean 3×. - Engine hardening: none needed from the run. The dogfood surfaced an environment bug: the operator's
.iterion/dispatcher/dispatcher.jsonrouted every bot to staleexamples/<bot>paths (bots had moved tobots/<bot>), soiterion dispatchrefused to start (stat examples/branch_improve_loop: no such file) and the studio's embedded dispatcher had silently degraded to an inert stub (slots.global_max=0,tracker="") — enabled but never dispatching. Config validation failed loudly for the CLI (good); the studio swallowed it into a stub (worth surfacing in the dashboard). Worked around with a minimal repaired config (feature-dev → bots/feature-dev); operator's original backed up to/tmp/operator-dispatcher.json.bak. - Lessons for next run: feature-dev handles a doc+code+test feature cleanly when the ticket carries a concrete, anti-façade done criterion (here: "a test proving the read returns while the actor is provably blocked; adding the field is not sufficient"). Keep the dispatcher config current when bots move dirs (
examples/*→bots/*), and consider a startup banner when the embedded dispatcher degrades to a stub.
2026-06-13 — sandbox-doctor static-binary check (runs 019ec149, 019ec180)
Update — fix applied + validated (run 019ec180). Taught
act/fixtogit -C <workspace_dir> add -Aafter editing (commit44d34c9d), so new files are tracked and visible to the reviewers'git diff HEAD. Re-running the SAME feature_prompt: Featurly converged and committed (finished, $2.85 / 247 steps vs the looping$4.95 / 507 / cancelled), shipping commit439d1116oniterion/run/opal-flash-mothbeam-80d7—pkg/cli/sandbox.go(+106, the doctor static/dynamic ELF check + WARNING), a tracked test, ANDdocs/adr/019-sandbox-doctor-static-binary-check.md. The new test being in the commit is the direct proof the untracked-files bug is fixed. Feature pending integration to main (after the parallel Depsy run, to avoid a watchexec restart).
- Status (original run 019ec149): failed to converge — implementation correct, review loop broken for new-file features → FIXED + validated (see update above).
- Versions: bot feature-dev 0.1.0 · iterion f247f360
- Method:
POST /api/runs,feature_prompt= add a static-binary check toiterion sandbox doctor(warn when the host iterion is dynamically linked — the exact trap that broke Seki).--merge-into none, defaultworkspace_dir(worktree-isolated ✅,.iterion/worktrees/019ec149..., safe under watchexec). Backends: claude_code opus (plan/act/simplify/fix_claude/prepare_commit) + claw gpt-5.5 (reviewer_gpt/fix_gpt). 101.7k tokens, $4.95, 507 steps, review_loop 10/15 — cancelled (non-convergent).
Value (the implementation is genuinely good)
actproduced a correct, well-documented feature:pkg/cli/sandbox_binary.gowithiterionBinaryIsStatic(path)detecting static-vs-dynamic via the ELFPT_INTERPprogram header, a focused_test.go, and thesandbox doctorintegration insandbox.go. The doc comment even citesaddClawBinaryMountand the preciseexec: … no such file or directoryfailure mode. Salvageable from the preserved (cancelled-run) worktree.
Findings / misses
- SEVERE — feature-dev cannot converge on a feature that ADDS files. The reviewer anchor protocol correctly says "diff
git diff HEAD, NOTHEAD^…HEAD" (so a reviewer doesn't conclude "feature not implemented" off the base commit). Butgit diff HEADomits untracked files — andactcreates new files withoutgit adding them. So the helper + test (pkg/cli/sandbox_binary.go,…_test.go) were??untracked, invisible to the reviewers'git diff HEAD --name-only. The GPT reviewer correctly rejected every pass: "the helper and focused unit test are still untracked … the committable tracked diff referencesiterionBinaryIsStaticwithout including its implementation or the required test." Thefix_*agents can't resolve it (the files already exist; the real gap is staging), so it loops toreview_loop(15)and dies. This almost certainly hits any review loop that anchors ongit diff HEADfor a change that adds files (feature-dev, possibly Billy/branch-improve-loop and Doki). - Cost of non-convergence: $4.95 / 101k tokens / 507 steps burned on 10 passes that could never pass — the loop has no "is this failing for a structural reason I can't fix?" escape, it just re-runs the fixer against an unfixable blocker.
Engine hardening / fix (recommended — needs a validated re-run)
- Make untracked new files visible to the review diff. Cleanest: a deterministic
git -C <wt> add -N .(intent-to-add) before the anchor diff, sogit diff HEADshows new files as additions (full content) without fully staging them; the existingprepare_commit'sgit add -- <files>still does the real staging at commit. Equivalent belt-and-suspenders: haveact/fix_*git addnew files when they create them. Apply the same to every loop bot that diffsgit diff HEAD. - The canonical asymptote guidance in [docs/workflow_authoring_pitfalls.md] / CLAUDE.md ("reviewers MUST diff
git diff HEAD, notHEAD^…HEAD") is now extended with the untracked-files caveat (CLAUDE.md, asymptote section). - Not patched in this pass (a careful multi-spot reviewer-prompt change that needs its own validated Featurly run); tracked here as the next feature-dev fix.
Lessons for next run
- Before trusting a feature-dev run, check the worktree's
git status: ifactleft??untracked files, the review loop will never converge until they're staged — that's the bug above, not a bad implementation. - The implemented feature here (sandbox-doctor static-binary warning) is worth salvaging by hand from the worktree — it directly prevents the dynamic-binary trap that cost this campaign two Seki failures.
