An agent that cannot undo itself should not be allowed to act.
So this one's only write is a git commit — and after it commits, it watches, withdraws the fix that did not hold, and keeps score.
Every observability tool can tell you a pod is in CrashLoopBackOff. Almost none of them are trusted to do anything about it, and the reason is not model quality. Granting write access to production requires an answer to one question — what stops this doing something catastrophic at 3am — and the model is usually careful is not an answer. It is a claim about average behaviour offered in reply to a question about the worst case.
The obstacle is structural. An agent that emits shell commands or free-form manifests has an unbounded action space, and three properties then become unavailable in principle: you cannot enumerate what it might do, you cannot compute the blast radius before it runs, and there is no mechanical inverse, so recovery is itself a judgement call under pressure.
kubemend inverts the constraint. It diagnoses freely and acts narrowly, and the narrowness is enforced by code with tests rather than by instructions in a prompt. It holds no cluster credentials and issues no write to the Kubernetes API. Its only write surface is a commit to the GitOps repository that already defines the cluster; the reconciler you already run carries it the rest of the way.
git revert. Not a feature anyone had to build, correct for twenty years, and already known by every operator on the team at 3am.
loading…
A correct change can be an unreviewable change
If review is your safety mechanism, diff size is a safety property — and the obvious implementation destroys it. Parse the YAML, mutate the object, serialise it back, and you get a correct file and a useless commit: comments dropped, keys reordered, quoting normalised. A reviewer expecting to check one number is shown three hundred changed lines.
What happens next is the actual failure. Nobody reads three hundred lines of reformatted YAML at 3am. They skim it, see it is machine-generated, and approve. The review step still exists in the diagram and has stopped being a control.
So kubemend rejects the YAML library and edits the text surgically, against an indentation-path tracker. Trailing comments survive, including the exact spacing before them.
loading…
What it refuses
Before anything is written, the plan goes through a gate: a pure function from plan and policy to a verdict, with no model in the loop. It bounds blast radius, requires a computable undo, rate-limits itself, and excludes the control plane outright.
Below are the four plans kubemend produces from a recorded cluster shipped in the repository. Switch the policy and watch where the verdicts move — and where they do not.
| Workload | Action | Verdict | Why |
|---|---|---|---|
| loading… | |||
payments/api is promoted from review to unattended, while kube-system/coredns stays refused under both — the identical rollback that is applied elsewhere.That refusal is a hard one. Control-plane namespaces are excluded rather than guarded by a threshold, because touching them can remove the machinery you would use to recover.
An unconfigured Policy() permits nothing. A permissive default is the difference between a bounded assistant and an unsupervised process with cluster credentials.
Autonomy is earned, not just configured
Those levels are a starting point, not a verdict. A workload's own record moves it: ten consecutive fixes that held raise it a level, one withdrawn fix lowers it. The number the log has been collecting all along is what decides.
The asymmetries are the design. Promotion is slow and demotion is instant, because promoting too eagerly hands production to an agent that has not earned it while demoting too eagerly costs a human some review. Policy still sets the bounds and evidence only moves inside them. And a spotless record buys a shorter leash, never a different rulebook, so kube-system stays refused no matter how good the history is.
Abstaining is the common case
—
Those are not gaps. A missing ConfigMap needs a value the agent has no business inventing; an unschedulable pod is a capacity decision; a container restarting for unclear reasons needs a human to read the logs. An agent that acted on all thirteen would be worse, not more capable — so abstention is a designed output, and the test suite asserts it rather than merely tolerating it.
Then it checks its own work
A loop that stops at committed is half a loop. The agent acted on a diagnosis that may have been wrong, and until something checks, the cluster is in a state nobody has confirmed is better than the one it replaced.
So it polls. Recovery means the motivating findings are gone and no new critical finding appeared, confirmed by two consecutive clean reads, because a rollout looks briefly healthy as it begins. A poll that cannot reach the cluster proves nothing — so it does not count toward recovery, and it breaks the streak rather than being skipped. Two clean reads either side of a blind one are not two consecutive observations.
On positive evidence that the fix did not hold, it withdraws its own commit.
demo/run.sh against a live k3d cluster. The second scenario ships two bad releases in a row, so rolling back one still leaves the workload broken — the case where the agent's plan is reasonable and still wrong. Timings are wall-clock from commit to verdict and vary between runs; the outcomes do not.The number nobody publishes
Detection rates get published constantly, and nobody is stuck on whether a crash loop can be identified. The question actually gating autonomy is different: how often is it wrong when it acts?
Because kubemend verifies, that number exists. Every run — findings, plan, verdict, commit, verification, revert — is written to an append-only log, and the log reports the agent's revert rate: how often its own fix failed to hold and it had to withdraw its own commit.
The repeat-offender line is the same idea over time. A workload remediated every week does not have a bad release, it has a bug — and without a history, three successful rollbacks look like three successes rather than one unaddressed defect.
A revert rate is what should earn an agent the promotion from propose to apply: measured on that workload, in that cluster, by that agent. A tool that declines to measure it is asking for trust it has not earned.
Try it
The analysis path runs with no cluster and nothing installed, against a recorded snapshot of a cluster having a bad afternoon.
git clone https://github.com/Srivatsa03/kubemend cd kubemend && pip install -e ".[dev]" kubemend diagnose --snapshot fixtures/broken-cluster.json kubemend policy # what each policy permits demo/run.sh # the full loop, real k3d cluster kubemend log # what it knows about itself kubemend serve # the console, on localhost
k3d, kubectl and a Docker daemon. CI runs that same script against a real cluster on every push, because a project claiming it works against a real cluster should not prove it with mocks.What has not been established
No production deployment — the cluster is local, there is no traffic, and the incidents are injected rather than organic. The headline demo substitutes kubectl apply for a reconciler; deploy/ runs real Argo CD against a real GitHub repository instead, and doing that is what exposed a bug where an applied fix was committed but never pushed. The planner is deterministic; there is no model in the decision path, which is sequencing rather than limitation, so that when one is added it correlates and explains rather than deciding what happens to your cluster. Verification is single-workload: it confirms the treated workload recovered, not that the change was harmless to its dependents. Coverage is Deployments only. Version 0.2.0, alpha.