kubemend
Systems design · autonomous remediation

An agent that cannot undo itself should not be allowed to act.

So this one's only write is a git commit — and after it commits, it watches, withdraws the fix that did not hold, and keeps score.

Author
Srivatsa Kamballa
Published
August 2026
Evidence
—
Implementation
—

Every observability tool can tell you a pod is in CrashLoopBackOff. Almost none of them are trusted to do anything about it, and the reason is not model quality. Granting write access to production requires an answer to one question — what stops this doing something catastrophic at 3am — and the model is usually careful is not an answer. It is a claim about average behaviour offered in reply to a question about the worst case.

The obstacle is structural. An agent that emits shell commands or free-form manifests has an unbounded action space, and three properties then become unavailable in principle: you cannot enumerate what it might do, you cannot compute the blast radius before it runs, and there is no mechanical inverse, so recovery is itself a judgement call under pressure.

kubemend inverts the constraint. It diagnoses freely and acts narrowly, and the narrowness is enforced by code with tests rather than by instructions in a prompt. It holds no cluster credentials and issues no write to the Kubernetes API. Its only write surface is a commit to the GitOps repository that already defines the cluster; the reconciler you already run carries it the rest of the way.

Undoing the agent is git revert. Not a feature anyone had to build, correct for twenty years, and already known by every operator on the team at 3am.
the commit it wrote—
loading…
Figure 1. One incident, end to end. The commit carries the evidence that produced it, the blast radius, and the instruction that matters. Nothing here was written by a human, and the manifest edit underneath it is a single line.

A correct change can be an unreviewable change

If review is your safety mechanism, diff size is a safety property — and the obvious implementation destroys it. Parse the YAML, mutate the object, serialise it back, and you get a correct file and a useless commit: comments dropped, keys reordered, quoting normalised. A reviewer expecting to check one number is shown three hundred changed lines.

What happens next is the actual failure. Nobody reads three hundred lines of reformatted YAML at 3am. They skim it, see it is machine-generated, and approve. The review step still exists in the diagram and has stopped being a control.

So kubemend rejects the YAML library and edits the text surgically, against an indentation-path tracker. Trailing comments survive, including the exact spacing before them.

clusters/prod/checkout.yaml1 file changed
loading…
Figure 2. The entire change. The comment keeps its column, because a diff where the comment shifts is a diff a reviewer has to read twice — and every extra token of reading is a chance the review stops being real.

What it refuses

Before anything is written, the plan goes through a gate: a pure function from plan and policy to a verdict, with no model in the loop. It bounds blast radius, requires a computable undo, rate-limits itself, and excludes the control plane outright.

Below are the four plans kubemend produces from a recorded cluster shipped in the repository. Switch the policy and watch where the verdicts move — and where they do not.

policy
WorkloadAction VerdictWhy
loading…
Figure 3. Verdicts computed by the real gate and exported at build time, so this table cannot drift from the code. Loosening policy changes the destination of a plan, not permission: payments/api is promoted from review to unattended, while kube-system/coredns stays refused under both — the identical rollback that is applied elsewhere.

That refusal is a hard one. Control-plane namespaces are excluded rather than guarded by a threshold, because touching them can remove the machinery you would use to recover.

An unconfigured Policy() permits nothing. A permissive default is the difference between a bounded assistant and an unsupervised process with cluster credentials.

Autonomy is earned, not just configured

Those levels are a starting point, not a verdict. A workload's own record moves it: ten consecutive fixes that held raise it a level, one withdrawn fix lowers it. The number the log has been collecting all along is what decides.

The asymmetries are the design. Promotion is slow and demotion is instant, because promoting too eagerly hands production to an agent that has not earned it while demoting too eagerly costs a human some review. Policy still sets the bounds and evidence only moves inside them. And a spotless record buys a shorter leash, never a different rulebook, so kube-system stays refused no matter how good the history is.

Demotion stops at "a human reviews the next one". It never falls to silence. A workload whose last fix was withdrawn is exactly the one you want to see a proposal for, and an earlier version of this rule dropped it to report-only, which meant the agent went quiet at the worst possible moment.

Abstaining is the common case

—

Those are not gaps. A missing ConfigMap needs a value the agent has no business inventing; an unschedulable pod is a capacity decision; a container restarting for unclear reasons needs a human to read the logs. An agent that acted on all thirteen would be worse, not more capable — so abstention is a designed output, and the test suite asserts it rather than merely tolerating it.

Then it checks its own work

A loop that stops at committed is half a loop. The agent acted on a diagnosis that may have been wrong, and until something checks, the cluster is in a state nobody has confirmed is better than the one it replaced.

So it polls. Recovery means the motivating findings are gone and no new critical finding appeared, confirmed by two consecutive clean reads, because a rollout looks briefly healthy as it begins. A poll that cannot reach the cluster proves nothing — so it does not count toward recovery, and it breaks the streak rather than being skipped. Two clean reads either side of a blind one are not two consecutive observations.

On positive evidence that the fix did not hold, it withdraws its own commit.

Figure 4. Both outcomes, from one run of demo/run.sh against a live k3d cluster. The second scenario ships two bad releases in a row, so rolling back one still leaves the workload broken — the case where the agent's plan is reasonable and still wrong. Timings are wall-clock from commit to verdict and vary between runs; the outcomes do not.
Reverting on no evidence would be worse than not reverting. An unreachable API server is not a failed fix. Undoing on that basis turns a network blip into a second unplanned production change, so an unverifiable outcome stops and asks for a human. Silence is not success — but it is not failure either.

The number nobody publishes

Detection rates get published constantly, and nobody is stuck on whether a crash loop can be identified. The question actually gating autonomy is different: how often is it wrong when it acts?

Because kubemend verifies, that number exists. Every run — findings, plan, verdict, commit, verification, revert — is written to an append-only log, and the log reports the agent's revert rate: how often its own fix failed to hold and it had to withdraw its own commit.

Figure 5. From the same live run. The 50% is a property of the demo, not an accuracy claim — one of its two scenarios is deliberately unfixable by rollback. It shows the measurement working end to end: a failed fix detected, withdrawn, and counted. A production revert rate has not been measured, and quoting this one as such would be quoting a demo parameter.

The repeat-offender line is the same idea over time. A workload remediated every week does not have a bad release, it has a bug — and without a history, three successful rollbacks look like three successes rather than one unaddressed defect.

A revert rate is what should earn an agent the promotion from propose to apply: measured on that workload, in that cluster, by that agent. A tool that declines to measure it is asking for trust it has not earned.

Try it

The analysis path runs with no cluster and nothing installed, against a recorded snapshot of a cluster having a bad afternoon.

terminalno dependencies
git clone https://github.com/Srivatsa03/kubemend
cd kubemend && pip install -e ".[dev]"

kubemend diagnose --snapshot fixtures/broken-cluster.json
kubemend policy                    # what each policy permits

demo/run.sh                        # the full loop, real k3d cluster
kubemend log                       # what it knows about itself
kubemend serve                     # the console, on localhost
The live demo needs k3d, kubectl and a Docker daemon. CI runs that same script against a real cluster on every push, because a project claiming it works against a real cluster should not prove it with mocks.

What has not been established

No production deployment — the cluster is local, there is no traffic, and the incidents are injected rather than organic. The headline demo substitutes kubectl apply for a reconciler; deploy/ runs real Argo CD against a real GitHub repository instead, and doing that is what exposed a bug where an applied fix was committed but never pushed. The planner is deterministic; there is no model in the decision path, which is sequencing rather than limitation, so that when one is added it correlates and explains rather than deciding what happens to your cluster. Verification is single-workload: it confirms the treated workload recovered, not that the change was harmless to its dependents. Coverage is Deployments only. Version 0.2.0, alpha.