Skip to main content

Flywheel Vision — Dev-Loop Today, User Pipeline at v1.0

A living north-star document. It explains why the Flywheel exists and what it is becoming. The operational how lives in roles/flywheel.md; the per-run what lives in docs/flywheel-brief.md. When the role or the brief ever drifts from this doc, the vision wins — fix the role or the brief.

TL;DR

The Flywheel is a dev-loop for building Overdeck itself, not yet a feature for end users. Today it runs real work through Overdeck’s own pipeline so that substrate defects surface as the work flows, and we fix them at the root as they appear. At v1.0, the substrate is reliable enough that the same loop becomes a user-facing delivery pipeline: users get features and fixes into the pipeline and the system delivers them end-to-end, with only human UAT and merge in the loop. This document names the mental model, the vocabulary, and the measurable criteria for “v1.0-ready.” It is the why behind every rule in the role and the brief.

The mental model

Today: the Flywheel as substrate dev-loop

The Flywheel takes whatever work is moving through Overdeck and watches it flow. When a stage misbehaves — a review specialist fails to block obvious issues, a test agent loops, a merge gate confuses itself, a planning artifact lands at the wrong path — the failure is a defect in Overdeck, not in the work. We call those substrate bugs. The work issue is almost a vehicle. The real product of a revolution is:
  1. A merged PR for the work itself (the issue gets done).
  2. Zero or more substrate-bug issues filed against Overdeck.
  3. Zero or more substrate-bug fixes merged (often before the original work merges).
Every substrate bug is a step toward v1.0. The Flywheel exists to expose them faster than we’d find them with manual usage — and to fix them at the root, so each revolution leaves the substrate permanently better. A workaround is a failed tick.

v1.0: the Flywheel as user-facing pipeline

When the substrate is solid, the Flywheel becomes a user-facing capability. Users (developers using Overdeck for their own projects) get features and bug fixes into the pipeline and the system delivers them end-to-end. The pipeline reliably picks up backlog work; plans, implements, reviews, and tests it; surfaces merge-ready work for human UAT; and merges when the human approves. At v1.0, the only required human input is UAT (optional, per-issue) and merge approval. Users who trust the system more flip per-issue toggles to skip UAT for low-risk classes of work; users who trust it less keep UAT required for everything. The substrate doesn’t change either way — only the human checkpoints do.

What changes between today and v1.0

The Flywheel doesn’t disappear at v1.0. It changes role — from “exposes our own defects” to “drives the user’s pipeline” — the same orchestration loop with different defaults and a much smaller intervention rate.

Definitions

  • Substrate bug — a defect in Overdeck itself: a broken route, a misbehaving role prompt, a wrong gate condition, a flaky lifecycle transition, a UI that lies about state. Substrate bugs are filed against eltmon/overdeck and fixed through the normal Overdeck pipeline.
  • Work-code bug — a defect in the application code an issue agent is producing. Work-code bugs are normal pipeline operation, not a v1.0 signal — review and test agents catch them.
  • Pipeline run — one issue’s journey from start to merged (or to parked/cancelled). The Flywheel may have many issues in flight; each is its own run for measurement.
  • Pipeline pass — a single end-to-end completion: plan → work → review → test → merge. A run may take multiple passes if review or test rejects work and the agent re-submits.
  • Operator intervention — any human action on a run beyond clicking merge: filing a substrate bug, telling an agent something, pausing/resuming, restarting a review cycle, editing a workspace file, manual merge resolution. UAT is not an intervention — it’s the expected human checkpoint.
  • Substrate-bug discovery rate — substrate bugs filed per pipeline run, over a rolling window. The headline v1.0 metric.

The autonomy model

Three properties govern how much the Flywheel does on its own. The first is a permanent security invariant; the other two are operator toggles (persisted in the app_settings SQLite table, user-machine level) that move the human checkpoints without changing the substrate.

Author + assignee invariant (security-critical, always on)

The Flywheel includes an issue in inventory and suggestions only if at least one of:
  • author.login is eltmon (project owner), OR
  • author.login is panopticon-agent[bot] (the Overdeck GitHub App), OR
  • assignees[].login contains eltmon (operator has personally assigned the issue).
Any other state — third-party author and eltmon not among assignees — is out of scope. When auto-pickup is ON, this filter is the only safeguard between an attacker filing a malicious issue and the Flywheel autonomously running an agent against it. It applies always, regardless of any toggle, and has no relaxation flag.

Auto-pickup from backlog (auto_pickup_backlog, default OFF) — the master autonomy switch

  • OFF (dev-loop today): the Flywheel works only in-flight work plus emergency blocks-main unblockers. Routine backlog stays off the start list; the operator feeds specific items in by Releasing them. The Flywheel still plans ready backlog aggressively, so a deep awaiting-release queue is always there for the operator.
  • ON (v1.0 posture): the toggle is a blanket release. The Flywheel auto-starts ready && planned backlog in sequencer-priority order, bounded by roles.flywheel.maxAgents. The gates that still apply are vetoed (absolute stop), parked/needs-design/ needs-discussion, objection (an AI relevance hold), and the relevance-vet.
A single predicate is the source of truth — isAutoPickable in src/lib/backlog/pickup.ts — shared by the dashboard forecast and the Flywheel so they can never disagree: ready && planned && (released || auto_pickup_backlog) && !parked && !vetoed && !objection && !inPipeline && !epic

UAT-required-before-merge (require_uat_before_merge, default ON)

“Human merges” and “human UATs” are independent constraints. When true (today’s default), merge cannot proceed until a human marks UAT passed. When false, the merge gate only requires the configured CI checks (overdeck/review + overdeck/test); the human still clicks merge. The “auto-merge” side — automatically clicking MERGE after a cooldown when every gate is green — composes orthogonally: The model we want for v1.0 is “relax UAT for trusted work, but a human always owns the merge.”

v1.0 readiness criteria

These are the thresholds for declaring Overdeck’s substrate “v1.0-ready,” anchored to industry benchmarks (see How we measure). They are draft until refined against ~1 month of real Flywheel telemetry. We are v1.0-ready when all of the following hold for 30 consecutive days: Why “all of the above” and not a weighted score: a single metric can be gamed (fewer runs → fewer bugs) or hidden (intervene constantly → success rate looks fine). The combination forces honest measurement. Why 30 consecutive days: enough runs for signal without a year-long wait; afterward we re-affirm v1.0 or name the bottleneck criterion. What “v1.0” does not mean: not “no bugs ever” (criterion 1 leaves an error budget); not “no human in the loop” (humans still UAT and approve merges); not “all features done” (v1.0 is about substrate reliability — feature work continues after).

How we measure

The targets triangulate from three industry analogues, because “substrate-bug discovery rate per pipeline run” has no direct published benchmark:
  1. Discovery rate ↔ defect escape rate. Capers Jones’ Defect Removal Efficiency: median DRE ~85% (15% escape), best-in-class >95% (<2% escape). Substrate bugs surfacing during runs are the “escapes” of the substrate’s manufacturing process.
  2. Pass success rate ↔ CI/CD SLO. Google SRE and Shopify treat CI/CD as a 99%-availability service with a 1% error budget; we scope the 99% to substrate-attributable failures.
  3. Intervention rate ↔ agent-autonomy benchmarks. Top SWE-bench Verified systems sit at 80–94%; our <5% intervention target is stricter than the agent layer because the substrate is supposed to be the reliable part.
Sources: DORA 2024 · Capers Jones, DRE · Google SRE Workbook · Shopify Engineering · SWE-bench Verified · Cognition / Devin report. Caveats: there is no published intervention-rate benchmark for internal dev tooling (the <5% is derived, not measured); DORA tier cutoffs are ±20% bands; the discovery-rate framing is novel. This is why the criteria are draft and revisited after ~1 month of data.

The path there

The work that gets us to v1.0 is tracked by the v1.0-required GitHub label, which is the canonical, always-current list — this doc deliberately does not inline it, because a snapshot here goes stale while the label stays true:
Live critical path: open v1.0-required issues
The minimum bar underneath every criterion is substrate-bug provenance + telemetry (so the seven criteria can be measured at all) and metric-aware prioritization (so the bottleneck criterion bubbles to the top of the sequencer rather than P-level and age alone). Substrate bugs declare affected v1.0 criteria via the Flywheel-Affects-Criterion trailer; pan flywheel weights --json computes a per-bug weight from live telemetry, and the Flywheel orders substrate-hardening suggestions by that weight within the tier. See docs/FLYWHEEL.md for the trailer format and the weight model. The Flywheel and the backlog sequencer prioritize substrate-hardening epics first: a stable substrate is the prerequisite for everything else on the path.

Open questions

Decisions we haven’t made and don’t need to make today, but should revisit before v1.0:
  • Per-issue UAT override. Once the global require_uat_before_merge flag exists, do we add a per-issue override (a bypass-uat label or an xBRIEF field)? Defer until after the global toggle has run for a while.
  • Substrate-bug label. A dedicated substrate-bug label for visibility, separate from the commit-trailer provenance? Probably yes for visibility — but it never replaces provenance (labels can be edited; trailers can’t).
  • What counts as a pipeline run for measurement? Does a parked/cancelled issue count? We may scope the denominator to runs that reach at least the work stage.
  • When to retire this doc. At v1.0 it transitions from “north star” to “history of how we got here.” Keep it as the record of the dev-loop → v1.0 transition; don’t delete it.

Appendix: where this fits in the docs

  • roles/flywheel.md — the durable doctrine: root-cause rules and the load-bearing rails (author/assignee gate, pickup gate, veto, saturation cap). The tick loop itself is the pan-flywheel skill.
  • docs/flywheel-brief.md — the per-run scope contract: scope and this run’s config. The what-this-run.
  • docs/FLYWHEEL.md — technical reference for the contract, lifecycle, and dashboard surfaces.
  • docs/FLYWHEEL-STATE.md — durable cumulative memory authored by the orchestrator across runs.
  • packages/contracts/src/flywheel.ts — the FlywheelStatus schema the orchestrator emits each tick.