The failure mode: one proxy becomes the company.
Agent systems are unusually vulnerable to Goodhart effects. A receipt, a queue row, a test pass, a model label, a conversation score or a manager approval can begin as useful evidence and quietly turn into the objective itself. The system then gets better at producing the proxy while real-world progress stalls.
What changed in the control architecture.
Historical evidence rejected a fixed prompt recipe.
We compared two distinct operational periods covering more than 2,500 complete historical agent turns. The same broad prompt family behaved very differently across regimes. Research-oriented prompts were sustained and effect-heavy in an older period, yet much shallower in a recent period. A large imperative “programmatic self-propulsion” prompt family also produced surprisingly short median turns in the recent data.
These are observational signals, not universal causal laws. The practical conclusion is stronger than choosing a single winning prompt: prompt effectiveness is contextual and nonstationary. The allocator therefore carries a posterior over treatments and reserves challenger mass instead of freezing one template.
Parallelism is a treatment too.
Historical traces contained useful work at several concurrency levels, including higher-parallelism regimes. That evidence is confounded by workload and time period, so a fixed global “one at a time” or fixed global high-concurrency rule would both overstate what the history proves. The operating policy is a distribution over parallelism treatments updated by current reasoning integrity, task parallelizability and downstream effects.
Alignment is actuation, not imitation.
A model that can predict a manager’s later decision may still make a poor decision itself. For management alignment, the candidate must make its own decision from the available situation and goal context. There is no hidden target answer. Evidence is recorded across the relevant dimensions and subsequent real-world consequences; later human text is not converted into a scalar reward.
What counts as success.
- A source commit or passing test is implementation evidence, not deployment success.
- A provider-accepted turn is execution evidence, not proof of external value.
- A public page, listing, order, deployed runtime or independently read-back effect is stronger downstream evidence.
- Cash is recorded only as cash; public activity and forecasts cannot impersonate settled revenue.
- An autonomous chain is credited only when the system itself causes the next useful action and the causal relationship is retained.