How it works

Five stages. One of them is you.

Fidelio is not a chat window that emits code. It is a pipeline with an explicit plan, an independent reviewer that is allowed to fail the work, and a deploy gate that only a human can open.

The pipeline

Every run moves through the same five stages, in the same order.

Each stage produces an artifact that outlives it. That is what makes the run auditable after the fact instead of just observable while it happens.

  1. 01Planner

    The intent becomes a written plan before anything is built.

    You describe the outcome in plain language. The planner turns it into an explicit, versioned plan — the files it intends to touch, the schema it intends to change, the behaviour it intends to add. Nothing is written yet. This is the cheapest possible place to disagree, and it is the stage most tools skip entirely.

    Output: a reviewable plan, versioned alongside the run.

  2. 02Builder

    Implementation runs against the plan, not against a vibe.

    The builder executes the approved plan. Because the plan is explicit, drift is detectable: work that wanders outside the stated scope is visible as a difference between what was planned and what was produced, rather than disappearing into a large diff nobody reads.

    Output: the change set, plus the record of how it was produced.

  3. 03Reviewer

    A separate agent grades the work — and it can fail it.

    The reviewer is a distinct stage with its own context, not the builder marking its own homework. It checks the result against the original intent and the plan, and it returns a verdict that can be negative. A review that is structurally incapable of saying no is decoration, so this one is allowed to stop the run.

    Output: a verdict with reasoning, attached to the run.

  4. 04Gate

    Nothing reaches production without a human decision.

    The gate is the point where a person decides. Fidelio surfaces the plan, the change, and the reviewer's verdict together, then waits. Certain classes of change — an untested authorization change, for example — are flagged specifically, because those are the ones where an automatic yes is most expensive.

    Output: an approval or a rejection, recorded against the run.

  5. 05Deploy

    The deploy is the consequence of the decision, not a side effect.

    Only an approved run deploys. The full chain stays attached to it afterwards: what was asked, what was planned, what was built, what the reviewer said, and who approved it. When something breaks at 2am, that chain is the difference between a diagnosis and a guess.

    Output: a deployment traceable to its authorising decision.

Why it is built this way

A review that cannot fail is not a review.

Most agent tools generate confidently and verify loosely. The failure mode is not bad code — it is plausible code that nobody checked, shipped by a process with no record of who agreed to it.

Separation

The reviewer is not the builder

Grading your own output produces agreement, not verification. The reviewer runs as its own stage with its own context so its verdict carries information.

Evidence

Artifacts outlive the run

The plan, the change, and the verdict are stored, not streamed and forgotten. You can answer 'why is this here' months later.

Authority

The gate is human by construction

Not a setting that defaults to on. An approved deploy has a person attached to it, recorded against the run.

Scope

Drift is visible

Because the plan is explicit and versioned, work outside it shows up as a difference rather than vanishing into a large diff.

Blast radius

Dangerous classes are flagged

Changes like untested authorization edits are surfaced specifically at the gate, because that is where an automatic yes costs the most.

Recovery

Traceable after the fact

Every deployment points back to the intent, the plan, the review, and the approval that authorised it.

Put an intent in and see what comes back.

Every run produces a plan, a build, an independent review, and a gate you control.