Flywheel Vision — Dev-Loop Today, User Pipeline at v1.0
A living north-star document. It explains why the Flywheel exists and what it is becoming. The operational how lives inroles/flywheel.md; the
per-run what lives in docs/flywheel-brief.md. When the role or
the brief ever drifts from this doc, the vision wins — fix the role or the brief.
TL;DR
The Flywheel is a dev-loop for building Overdeck itself, not yet a feature for end users. Today it runs real work through Overdeck’s own pipeline so that substrate defects surface as the work flows, and we fix them at the root as they appear. At v1.0, the substrate is reliable enough that the same loop becomes a user-facing delivery pipeline: users get features and fixes into the pipeline and the system delivers them end-to-end, with only human UAT and merge in the loop. This document names the mental model, the vocabulary, and the measurable criteria for “v1.0-ready.” It is the why behind every rule in the role and the brief.The mental model
Today: the Flywheel as substrate dev-loop
The Flywheel takes whatever work is moving through Overdeck and watches it flow. When a stage misbehaves — a review specialist fails to block obvious issues, a test agent loops, a merge gate confuses itself, a planning artifact lands at the wrong path — the failure is a defect in Overdeck, not in the work. We call those substrate bugs. The work issue is almost a vehicle. The real product of a revolution is:- A merged PR for the work itself (the issue gets done).
- Zero or more substrate-bug issues filed against Overdeck.
- Zero or more substrate-bug fixes merged (often before the original work merges).
v1.0: the Flywheel as user-facing pipeline
When the substrate is solid, the Flywheel becomes a user-facing capability. Users (developers using Overdeck for their own projects) get features and bug fixes into the pipeline and the system delivers them end-to-end. The pipeline reliably picks up backlog work; plans, implements, reviews, and tests it; surfaces merge-ready work for human UAT; and merges when the human approves. At v1.0, the only required human input is UAT (optional, per-issue) and merge approval. Users who trust the system more flip per-issue toggles to skip UAT for low-risk classes of work; users who trust it less keep UAT required for everything. The substrate doesn’t change either way — only the human checkpoints do.What changes between today and v1.0
The Flywheel doesn’t disappear at v1.0. It changes role — from “exposes our own defects” to
“drives the user’s pipeline” — the same orchestration loop with different defaults and a much
smaller intervention rate.
Definitions
- Substrate bug — a defect in Overdeck itself: a broken route, a misbehaving role prompt, a
wrong gate condition, a flaky lifecycle transition, a UI that lies about state. Substrate
bugs are filed against
eltmon/overdeckand fixed through the normal Overdeck pipeline. - Work-code bug — a defect in the application code an issue agent is producing. Work-code bugs are normal pipeline operation, not a v1.0 signal — review and test agents catch them.
- Pipeline run — one issue’s journey from
starttomerged(or toparked/cancelled). The Flywheel may have many issues in flight; each is its own run for measurement. - Pipeline pass — a single end-to-end completion: plan → work → review → test → merge. A run may take multiple passes if review or test rejects work and the agent re-submits.
- Operator intervention — any human action on a run beyond clicking merge: filing a substrate bug, telling an agent something, pausing/resuming, restarting a review cycle, editing a workspace file, manual merge resolution. UAT is not an intervention — it’s the expected human checkpoint.
- Substrate-bug discovery rate — substrate bugs filed per pipeline run, over a rolling window. The headline v1.0 metric.
The autonomy model
Three properties govern how much the Flywheel does on its own. The first is a permanent security invariant; the other two are operator toggles (persisted in theapp_settings SQLite
table, user-machine level) that move the human checkpoints without changing the substrate.
Author + assignee invariant (security-critical, always on)
The Flywheel includes an issue in inventory and suggestions only if at least one of:author.loginiseltmon(project owner), ORauthor.loginispanopticon-agent[bot](the Overdeck GitHub App), ORassignees[].logincontainseltmon(operator has personally assigned the issue).
eltmon not among assignees — is out of scope. When
auto-pickup is ON, this filter is the only safeguard between an attacker filing a malicious
issue and the Flywheel autonomously running an agent against it. It applies always,
regardless of any toggle, and has no relaxation flag.
Auto-pickup from backlog (auto_pickup_backlog, default OFF) — the master autonomy switch
- OFF (dev-loop today): the Flywheel works only in-flight work plus emergency
blocks-mainunblockers. Routine backlog stays off the start list; the operator feeds specific items in by Releasing them. The Flywheel still plans ready backlog aggressively, so a deep awaiting-release queue is always there for the operator. - ON (v1.0 posture): the toggle is a blanket release. The Flywheel auto-starts
ready && plannedbacklog in sequencer-priority order, bounded byroles.flywheel.maxAgents. The gates that still apply arevetoed(absolute stop),parked/needs-design/needs-discussion,objection(an AI relevance hold), and the relevance-vet.
isAutoPickable in
src/lib/backlog/pickup.ts — shared by the dashboard forecast
and the Flywheel so they can never disagree:
ready && planned && (released || auto_pickup_backlog) && !parked && !vetoed && !objection && !inPipeline && !epic
UAT-required-before-merge (require_uat_before_merge, default ON)
“Human merges” and “human UATs” are independent constraints. When true (today’s default),
merge cannot proceed until a human marks UAT passed. When false, the merge gate only requires
the configured CI checks (overdeck/review + overdeck/test); the human still clicks merge.
The “auto-merge” side — automatically clicking MERGE after a cooldown when every gate is green
— composes orthogonally:
The model we want for v1.0 is “relax UAT for trusted work, but a human always owns the merge.”
v1.0 readiness criteria
These are the thresholds for declaring Overdeck’s substrate “v1.0-ready,” anchored to industry benchmarks (see How we measure). They are draft until refined against ~1 month of real Flywheel telemetry. We are v1.0-ready when all of the following hold for 30 consecutive days:
Why “all of the above” and not a weighted score: a single metric can be gamed (fewer runs →
fewer bugs) or hidden (intervene constantly → success rate looks fine). The combination forces
honest measurement. Why 30 consecutive days: enough runs for signal without a year-long wait;
afterward we re-affirm v1.0 or name the bottleneck criterion.
What “v1.0” does not mean: not “no bugs ever” (criterion 1 leaves an error budget); not “no
human in the loop” (humans still UAT and approve merges); not “all features done” (v1.0 is
about substrate reliability — feature work continues after).
How we measure
The targets triangulate from three industry analogues, because “substrate-bug discovery rate per pipeline run” has no direct published benchmark:- Discovery rate ↔ defect escape rate. Capers Jones’ Defect Removal Efficiency: median DRE ~85% (15% escape), best-in-class >95% (<2% escape). Substrate bugs surfacing during runs are the “escapes” of the substrate’s manufacturing process.
- Pass success rate ↔ CI/CD SLO. Google SRE and Shopify treat CI/CD as a 99%-availability service with a 1% error budget; we scope the 99% to substrate-attributable failures.
- Intervention rate ↔ agent-autonomy benchmarks. Top SWE-bench Verified systems sit at 80–94%; our <5% intervention target is stricter than the agent layer because the substrate is supposed to be the reliable part.
The path there
The work that gets us to v1.0 is tracked by thev1.0-required GitHub label, which is the
canonical, always-current list — this doc deliberately does not inline it, because a snapshot
here goes stale while the label stays true:
Live critical path: open v1.0-required issues
The minimum bar underneath every criterion is substrate-bug provenance + telemetry (so the
seven criteria can be measured at all) and metric-aware prioritization (so the bottleneck
criterion bubbles to the top of the sequencer rather than P-level and age alone). Substrate bugs
declare affected v1.0 criteria via the Flywheel-Affects-Criterion trailer; pan flywheel weights --json
computes a per-bug weight from live telemetry, and the Flywheel orders substrate-hardening
suggestions by that weight within the tier. See docs/FLYWHEEL.md for the trailer format and the
weight model. The Flywheel and the backlog sequencer prioritize substrate-hardening epics first:
a stable substrate is the prerequisite for everything else on the path.
Open questions
Decisions we haven’t made and don’t need to make today, but should revisit before v1.0:- Per-issue UAT override. Once the global
require_uat_before_mergeflag exists, do we add a per-issue override (abypass-uatlabel or an xBRIEF field)? Defer until after the global toggle has run for a while. - Substrate-bug label. A dedicated
substrate-buglabel for visibility, separate from the commit-trailer provenance? Probably yes for visibility — but it never replaces provenance (labels can be edited; trailers can’t). - What counts as a pipeline run for measurement? Does a parked/cancelled issue count? We may scope the denominator to runs that reach at least the work stage.
- When to retire this doc. At v1.0 it transitions from “north star” to “history of how we got here.” Keep it as the record of the dev-loop → v1.0 transition; don’t delete it.
Appendix: where this fits in the docs
roles/flywheel.md— the durable doctrine: root-cause rules and the load-bearing rails (author/assignee gate, pickup gate, veto, saturation cap). The tick loop itself is thepan-flywheelskill.docs/flywheel-brief.md— the per-run scope contract: scope and this run’s config. The what-this-run.docs/FLYWHEEL.md— technical reference for the contract, lifecycle, and dashboard surfaces.docs/FLYWHEEL-STATE.md— durable cumulative memory authored by the orchestrator across runs.packages/contracts/src/flywheel.ts— theFlywheelStatusschema the orchestrator emits each tick.