Flywheel at scale: native feedback, personas, many workers, offline

Status: proposal · 2026-09-12

Flywheel is a small, deterministic engine that plugs in under any frontier lead. The lead (Claude Code, Codex, OpenCode, a human, or any agent that can run a shell) plans and judges; flywheel drives the worker agents and keeps the record. Phase 1 workers are OpenCode only; other worker CLIs come in later phases behind the same adapter seam.

1. Goals

Non-goals for now: a hosted service; replacing CI; landing work without review; non-OpenCode workers in phase 1.

2. Friction observed so far

Friction Evidence Addressed by
Every dispatch is hand-built, and silent hangs look like slowness stdin hang, forgotten flags, scraping session ids; ~40 hand-written dispatches per feature flywheel run (#20)
Watchers die at host time limits five watcher deaths at ~1 h while workers were healthy flywheel watch (#22), restartable from the event log
Provider and account failures cost the most time ~3 h on a key limit, then a 402, then a consent gate doctor (#23), circuit breaker, approved fallbacks (§4.4)
A global plugin swapped the worker’s agent 21 reads of one file, zero edits workers always run --pure
A shared working tree couples tasks gates fail on neighbours’ edits; finished work waits on the slowest worker a worktree per task, isolated review (#24), a landing queue (§4.3)
The lead reads every diff review caught every security defect, but it does not scale reviewer persona, automated gates, escalation rules (§4.5)
Exploring cannot be told from lost 53 steps of reads before any edit; 45 steps chasing a non-existent API plan check-in, exploring state, drift signal
State in a binary DB or a mutable snapshot no diff, no merge, no history the event log (#11)
The model setting resets on upgrade MODEL= lives in the skill .flywheel/config.json (#12)
Cross-OS surprises exec bit, CRLF, process names, a hanging gh upload CI matrix, Windows guidance, per-asset uploads
Feedback depends on someone writing a file by hand one consumer kept a learnings log; others would not native signals and learnings (§6)

3. Personas

A persona is a role, not a model. One agent can hold several; at small scale the lead is also planner, foreman, reviewer and steward. In phase 1 every persona except the worker can be any agent; the worker is OpenCode.

Persona Does Does not Typical model Skill
Lead owns the goal, plans waves, sets policy (limits, budgets, escalation), reviews escalations implement frontier flywheel
Planner turns a spec into briefs with owns:/needs:/exclusive:, choke points, gates dispatch frontier or mid flywheel-planner
Foreman runs a shard of workers: run, watch, classify, retry by policy, first-pass mechanical review, escalate plan, implement mid, or the CLI alone flywheel-foreman
Worker executes one brief and reports evidence plan, commit cheap (OpenCode) flywheel-worker
Reviewer judges a finished task: owns: check, gates in isolation, traps, domain checklists; verdict pass / correct / reject / escalate implement the fix mid flywheel-reviewer
Steward triages signals into learnings, deduplicates, proposes upstream feedback, keeps learnings.md healthy change code cheap or mid flywheel-steward
Operator installs, configures, assigns personas to agents, decides who merges and publishes — human or any flywheel-operator

Personas talk only through repository state (the event log, briefs, learnings) and flywheel commands. No agent-to-agent chat is required, so any mix of lead vendors works.

4. Architecture for scale

flowchart TD
    L["Lead (any frontier agent, 1-3)"] --> P["Planner"]
    P --> Q[("Briefs + event log")]
    L --> F1["Foreman A"]
    L --> F2["Foreman B"]
    F1 --> W1["OpenCode workers (worktree each)"]
    F2 --> W2["OpenCode workers (worktree each)"]
    W1 --> R["Reviewer"]
    W2 --> R
    R -->|pass| LQ["Landing queue"]
    R -->|escalate| L
    LQ --> M["Integration branch"]
    Q -. signals .-> S["Steward"]
    S --> LE[("learnings")]

4.1 Deterministic core, judgment at the edges

The CLI does everything mechanical: dispatch, watch, classify, retry by policy, isolate, gate, land, record. Agents do planning, review and the content of corrections. A foreman can be the CLI alone (flywheel watch plus policy) for shards whose tasks need no judgment until review.

4.2 State

4.3 Isolation and landing

4.4 Throughput controls

4.5 Review at scale

4.6 Proving scale without spending tokens

A sim worker adapter replays recorded OpenCode runs with a configurable mix of latency, output caps, stalls and provider errors. CI runs a wave of 1,000+ simulated tasks through next, run, watch, review and land, and asserts no double dispatch, no lost events and bounded lead escalations. It is the only non-OpenCode adapter in phase 1, and it never calls a model.

5. Leads and workers

5.1 Skills, for any lead

5.2 Workers: OpenCode in phase 1

5.3 Config

.flywheel/config.json (#12) grows to:

{
  "workers": [
    { "name": "cheap", "adapter": "opencode", "model": "openrouter/deepseek/deepseek-v4-flash-0731",
      "max_parallel": 20, "fallbacks": [ { "model": "opencode-go/deepseek-v4-flash", "approved": true } ] },
    { "name": "local", "adapter": "opencode", "model": "ollama/qwen3:8b", "max_parallel": 1 }
  ],
  "limits": { "per_host": 32, "budget": { "wave_cost_usd": 25 } },
  "feedback": { "upstream": "suzworx/flywheel", "submit": "ask" }
}

6. Native feedback

6.1 Signals (automatic)

The CLI appends a signal event whenever a run or task is not effective:

Each signal carries its kind, task and evidence (run file, session, counts). No agent has to remember to write it.

6.2 Learnings (curated)

6.3 Hard rules

6.5 Measuring flywheel itself

flywheel stats: first-pass review rate, corrections per task, signals per 100 runs, time lost by signal kind, cost per landed task. The trend is the health metric for flywheel.

7. Phases

Phase Scope
0 (in flight, #10) event log (#11), config (#12), skill feedback (#13-#18)
1 OpenCode run/status/watch (#20-#22) behind the adapter seam, plus sim; signals and flywheel feedback; persona skills; AGENTS.md bootstrap for leads; offline mode (§8)
2 worktree per task, isolated review (#24), local landing queue (reshapes #25), limits, breaker, budgets, doctor (#23), next (#27), leases, sharded log, the 1,000-task simulated wave in CI
3 more worker adapters (Claude Code, Codex, generic), multi-machine through git-merged shards, the optional coordinator, stats

8. Offline

What needs a network today, and how each part works without one:

Part Online today Offline
flywheel CLI none: Go stdlib, local files, git works as is; builds offline (no module dependencies)
State, review, landing none (local git) works as is; push and PRs wait until online
Worker model OpenRouter / OpenCode Go an OpenCode provider pointing at a local OpenAI-compatible server (Ollama, LM Studio)
OpenCode itself install; provider packages may download on first use; @latest plugins fetch at start install and run each provider once while online; workers run --pure, so plugins never fetch
Lead frontier API (Claude, Codex) needs its API; fully offline means a local lead too, which is much weaker, or a human lead
Skills npx skills add downloads copy the skill folders
CI, releases, upstream feedback GitHub run the gates locally; publish and submit feedback later from the outbox

Making it first-class:

9. Open questions