flywheel metrics
flywheel stats --metrics [--window 24h|7d|30d] [--json] computes the factory’s lean metrics from the
event log alone (#583). This page is the contract:
every number the stats panel, the :pulse dashboard and the :metrics view show is defined here,
from the ledger, with its unit. The engine is flywheel.Metrics in internal/flywheel/metrics.go;
it reads nothing but the events and the config, and the same ledger always gives the same report.
Window and buckets
A window is since <= ts < until. --window 24h ends now and has 1-hour buckets; 7d and 30d
have 1-day buckets. Every series is a list of numbers with one value per bucket, aligned to
buckets (each bucket’s start time), so a chart draws it directly. An event whose ts does not
parse counts nowhere. The text table’s trend arrow compares each value with the same metric over
the previous window of equal length: ↑ higher, ↓ lower, → equal.
Terms used below:
- landed unit: a task whose first
landedevent is in the window. - attempt span: a
dispatchedevent to the same task and attempt’sfinishedevent. A span nofinishedcloses ends at the task’s nextdispatched,lost,withdrawnorlandedevent, or at the window’s end while none came. - correction: an attempt beyond a unit’s first; a resume or a correction is attempt
c<n>.
JSON durations are nanoseconds; a distribution (Dist) is count, p50, p90, max, mean,
the percentiles by nearest rank.
Flow
| Metric | Definition | Unit |
|---|---|---|
throughput |
landed units | units |
throughput_series |
landed units per bucket, by the landing’s bucket | units per bucket |
wip |
tasks whose derived status (as flywheel state derives it from the events before the window’s end) is dispatched, running, finished or passed, and that are not stale |
units |
wip_series |
wip at each bucket’s end |
units |
stage_series |
per stage, the tasks in it at each bucket’s end that are not stale: queued (derived status planned: planned, not yet dispatched, neither landed nor abandoned), running (dispatched or running), finished, passed; the bands of the cumulative flow diagram |
units per stage |
stale |
tasks with one of those statuses whose last event before the window’s end is at or before the window’s end minus the stale threshold | units |
stale_oldest |
the window’s end minus the oldest stale task’s last event; 0 when none | duration |
lead_time |
per landed unit: first planned to first landed |
Dist |
cycle_time |
per landed unit: first dispatched to first landed |
Dist |
queue_time |
per landed unit: first planned to first dispatched |
Dist |
touch_time |
per landed unit: the sum of its attempt spans closed by a finished event |
Dist |
flow_efficiency |
mean over landed units with a positive cycle time of touch time / cycle time | ratio 0..1 |
A unit in progress with no event for the stale threshold is stale, not WIP, so a unit finished or
passed long ago and never landed or withdrawn stops inflating wip; the threshold is 7d by default
and flywheel stats --metrics --wip-stale-after <d> sets it (issue #590).
stage_series splits the flow by stage, so a widening band shows where work piles up: a wide
passed band says landing is the bottleneck, a wide queued band says dispatch is. At every
bucket running + finished + passed equals wip_series; queued is on top of WIP. In
flywheel stats --metrics --json it is an object under flow with the four keys, each one count
per bucket:
"stage_series": {"queued": [3, 1], "running": [2, 2], "finished": [1, 0], "passed": [0, 1]}
Quality
| Metric | Definition | Unit |
|---|---|---|
landed |
landed units | units |
first_pass |
landed units whose only attempt is r1 and that never got an inspected or reviewed verdict correct or rework |
units |
first_pass_yield |
first_pass / landed |
ratio |
corrections |
over landed units, the distinct attempts beyond the first | attempts |
rework_rate |
corrections / landed |
corrections per landed unit |
gates |
per gate id (the gate’s 1-based index) and gate command clipped to 60 characters: the window’s conclusive validated readings (an rc, reason neither inconclusive nor host-blocked), pass those with rc 0, rate = pass / total |
readings, ratio |
reviewed |
tasks with an agent reviewed event (persona reviewer or reviewer:<dimension>, an adapter recorded) in the window |
units |
findings, blocking |
review_finding events in the window; of those, severity blocker or major |
findings |
review_find_rate |
findings / reviewed |
findings per reviewed unit |
blocking_share |
blocking / findings |
ratio |
escapes |
landed units with a planned event after their landing (at any time) |
units |
Reliability
| Metric | Definition | Unit |
|---|---|---|
andons |
andons in the window by kind: every signal event (kind = its signal), and every finished event with reason stalled, silent, capped or rate-limited that no signal of the same name, task and attempt already records (kind = the reason) |
andons |
andon_total, andon_series |
the sum of andons; andons per bucket |
andons |
cleared |
andons on a task that a later event of the task cleared: landed, withdrawn, dismissed, or an inspected or hand-recorded reviewed verdict pass |
andons |
mttr |
per cleared andon: the andon to its first clearing event | Dist |
frozen.total |
the time inside the window the factory was suspended: from a suspended event to the first unsuspended event or its until, whichever comes first; a later suspended event while frozen moves the thaw time; one still open runs to the window’s end |
duration |
frozen.manual, frozen.auto |
the same split by who froze it; the suspended event carries no automatic flag yet, so every interval is manual and auto is 0 |
duration |
paused |
per model, sorted: the time inside the window a rate limit paused it. At every finished event on the model, the rate-limit pause rule flywheel run applies (a reset_at still ahead, or limit_utilization at least limits.rate_limit_pause_at with limit_reset_at ahead) decides whether it is paused; the pause lasts until its reset or the model’s next finished event |
duration |
Cost
| Metric | Definition | Unit |
|---|---|---|
spend |
the cost of the window’s finished and reviewed events |
USD |
spend_series |
spend per bucket |
USD per bucket |
units |
tasks with a finished event in the window |
units |
cost_per_unit |
spend / units |
USD |
cost_per_landed |
spend / landed |
USD |
tokens, steps |
the window’s finished events’ tokens and steps |
tokens, steps |
tokens_per_step |
(input + output + reasoning tokens) / steps | tokens |
by_model |
the flywheel stats --by model scoreboard computed over the window’s events only |
per model |
Capacity
| Metric | Definition | Unit |
|---|---|---|
workers[].busy_seconds |
the attempt spans inside the window of the worker: the dispatched event’s worker, else the configured worker on its model, else the default worker |
seconds |
workers[].utilization |
busy_seconds / (max_parallel × window seconds), max_parallel at least 1 |
ratio |
utilization |
every worker’s busy seconds / every worker’s capacity seconds | ratio |
idle_share |
1 - utilization, never below 0 |
ratio |
Ratios and USD are rounded to 4 places, tokens_per_step to 2, busy_seconds to 3; a ratio whose
denominator is 0 is 0.
Evidence
Every number drills to the units behind it. evidence maps a stable metric id to a list of
units, each task, value (the unit’s part in words), group (the split the metric uses) and
sort (the number ordering the list); a list is sorted by sort, highest (worst) first, then by
task. Times in a value are UTC, MM-DD HH:MM.
| Id | Covers | Units | value |
group |
sort |
|---|---|---|---|---|---|
flow.throughput |
throughput |
landed units | landed <time> |
landed |
the landing (Unix seconds) |
flow.wip |
wip |
units in progress at the window’s end; stale units too, apart | <status> for <since dispatch>; a stale unit stale: <status>, last event <age> ago |
the status; stale for a stale unit |
seconds since the first dispatch; for a stale unit seconds since its last event |
flow.lead_time, flow.cycle_time, flow.queue_time, flow.touch_time |
the Dist of that name | landed units with that time | lead 9h12m, cycle …, queued …, touch … |
landed |
the time in seconds |
flow.flow_efficiency |
flow_efficiency |
landed units with a positive cycle time | 33% (touch 30m of cycle 1h30m) |
landed |
1 - touch / cycle |
quality.first_pass_yield |
landed, first_pass, first_pass_yield |
landed units | first pass, corrected x<n>, sent back by a verdict, no r1 attempt |
first pass or corrected |
0 first pass, else corrections + 1 |
quality.rework_rate |
corrections, rework_rate |
landed units | <n> corrections |
as first-pass yield | corrections |
quality.gates |
gates |
the window’s conclusive gate readings | gate <id> rc <rc> on <attempt> |
pass or fail |
1 fail, 0 pass |
quality.review_find_rate |
reviewed, findings, review_find_rate |
the window’s review findings | <severity> <title> |
the severity | 1 blocking, 0 other |
quality.blocking_share |
blocking, blocking_share |
the window’s review findings | as above | blocking or other |
1 blocking, 0 other |
quality.escapes |
escapes |
landed units planned again | planned again <time>, landed <time> |
escaped |
seconds from landing to the new plan |
reliability.andons |
andons, andon_total, andon_series, cleared |
the window’s andons | <kind> <time>, cleared in 3h10m or , open |
the kind | seconds to clear, or open until the window’s end |
reliability.mttr |
mttr |
cleared andons | <kind> cleared in <d> |
the kind | seconds to clear |
reliability.frozen |
frozen |
units an attempt of which ran while the factory was frozen | held <d> (frozen <from> to <to>) |
frozen |
seconds held |
reliability.paused |
paused |
finishes that paused their model | <model> paused <d> from <time> |
the model | seconds paused |
cost.spend |
spend, spend_series |
units with a finished or reviewed cost in the window |
$<spend> |
landed or not landed |
USD |
cost.cost_per_unit |
units, cost_per_unit |
units with a finished event in the window |
$<spend> |
finished |
USD |
cost.cost_per_landed |
cost_per_landed |
as cost.spend |
$<spend> |
landed or not landed |
USD |
cost.tokens_per_step |
tokens, steps, tokens_per_step |
units with steps in the window | <n> tokens/step over <n> steps |
finished |
tokens per step |
cost.by_model |
by_model |
each unit’s finished cost per model | <model> $<cost> |
the model | USD |
capacity.utilization, capacity.idle_share |
workers, utilization, idle_share |
each unit’s busy time per worker | busy <d> on <worker> |
the worker | busy seconds |