← Blog
AI & DEVOPS·1 AĞUSTOS 2026AUG 1, 2026·11 DK OKUMA11 MIN READ

Auto-remediation incident'ı ne zaman büyütür: blast radius ve guardrail'lerWhen auto-remediation makes incidents worse: blast radius and guardrails

Self-healing altyapının demosu doksan saniyede biter, alkış alır. Aynı döngünün sağlıklı bir servisi dört dakikada bir yeniden başlattığı gün ise kimseye page düşmez — çünkü page düşmemesi zaten özellikti. İddiamız dar: self-healing'in değeri otomasyonda değil, karar sınırında yaşar. Ve bunu dört gerçek deney koşusuyla gösteriyoruz.

The self-healing demo ends in ninety seconds to applause. The day the same loop restarts a healthy service every four minutes, nobody gets paged — because not-paging was the feature. Our claim is narrow: self-healing's value lives in the decision boundary, not the automation. And we show it with four real experiment runs.

Naylalabs · MühendislikEngineering
Snapshot: Son doğrulama: 1 Ağustos 2026. Deneyler demo repo'nun remediation lab'inde koşuldu: agent-native-ops-demo (tag v0.4.0; ham timeline logları results/ altında). Aksiyonlar gerçek bir kind cluster'ında gerçek pod restart'ları; sağlık sinyali lab betiğinden — dedektör bilerek kör: sinyalin nedenini bilmiyor, gerçek dünyadaki gibi. Snapshot: Last verified August 1, 2026. Experiments ran in the demo repo's remediation lab: agent-native-ops-demo (tag v0.4.0; raw timeline logs under results/). Actions are real pod restarts on a real kind cluster; the health signal comes from the lab script — the detector is deliberately blind: it does not know the signal's cause, just like the real world.
Agent-Native Operations · 5/5. Beş bölümlük serinin 5. yazısı — 1. Stateless MCP server · 2. Tool contract'ları · 3. Agent kimliği · 4. OpenTelemetry GenAI · 5. Auto-remediation guardrail'leri Agent-Native Operations · 5/5. Part 5 of a five-part series — 1. Stateless MCP server · 2. Tool contracts · 3. Agent identity · 4. OpenTelemetry GenAI · 5. Auto-remediation guardrails

Self-healing altyapının, operasyon dünyasındaki en iyi demo'ya sahip olduğu tartışılmaz. Alert düşer, agent teşhis koyar, deployment geri alınır, grafik yeşile döner — doksan saniye, sıfır insan. Pitch kendini yazar ve 2026'da yanında rakamlar da gelir: MTTR %73 düşmüş, incident'ların çoğu otonom çözülüyormuş.

Self-healing infrastructure has, indisputably, the best demo in operations. An alert fires, an agent diagnoses, a deployment rolls back, the graph goes green — ninety seconds, zero humans. The pitch writes itself, and in 2026 it comes with numbers attached: MTTR down 73%, most incidents resolved autonomously.

Auto-remediation'a karşı değiliz — bu yazının deney düzeneği onu çalıştırıyor. Sınırsız auto-remediation'a karşıyız. Sınırı iyi çizerseniz remediation, sabahın üçündeki restart'larınızı yutar. Kötü çizerseniz — ya da sizin yerinize bir vendor'a çizdirirseniz — küçük incident'ları büyüğe çeviren, pager'ı da kapatılmış bir makine kurmuş olursunuz. Bu, serinin kimlik yazısındaki tezin en sert hali: agent'ın yetkisine production'ı düzeltmek de girdiğinde, sınırlı yetki kusursuz muhakemeyi döver.

We are not against auto-remediation — this post's experimental rig runs it. We are against unbounded auto-remediation. Draw the boundary well and remediation absorbs your 3 a.m. restarts. Draw it badly — or let a vendor draw it for you — and you have built a machine that turns small incidents into large ones, with the pager disabled. This is the hardest-edged version of the thesis from this series' identity post: when the agent's authority includes fixing production, bounded authority beats perfect judgment.

01Pitch destesindeki rakamlar ikinci bir bakışı hak ediyorThe numbers in the pitch deck deserve a second look

Bunların hiçbiri "yapmayın" demiyor. Şunu diyor: pazarlama mutlu yolu ölçüyor; siz dağılımın tamamını işleteceksiniz.

None of this says "don't". It says: marketing measures the happy path; you will be operating the whole distribution.

02Remediation'ın işleri kötüleştirdiği dört desenFour ways remediation makes things worse

1. Flapping. Remediation ateşlenir, semptom kısaca temizlenir, tetik yeniden kurulur, remediation yine ateşlenir. İmza: aynı hedefe aynı aksiyonun metronom aralıklarla uygulanması — kurması çocuk oyuncağı, neredeyse hiç kurulmayan bir alarm; çünkü aksiyonların hepsi loglara başarı diye yazılır. Aşağıdaki deneyde bunu üretiyoruz: 106 saniyede 13 restart, 12'si flapping.

1. Flapping. Remediation fires, the symptom briefly clears, the trigger re-arms, remediation fires again. Signature: the same action on the same target at metronomic intervals — a trivial alert to build, almost never built, because every action logs as a success. We produce it in the experiment below: 13 restarts in 106 seconds, 12 of them flapping.

2. Kaskad. Düzeltme yayılır: rollback trafiği kaydırır, komşu scale eder, scaler kotaya çarpar, kota alert'i kendi remediation'ını tetikler. Her aksiyon yerel olarak doğrudur — incident, bileşimin kendisidir. İmza: farklı hedefler arasında, alert-evaluation pencerenizin altında kalan aksiyon-aksiyon gecikmeleri.

2. Cascade. The fix propagates: a rollback shifts traffic, the neighbor scales, the scaler hits a quota, the quota alert triggers its own remediation. Each action is locally correct — the incident is the composition. Signature: action-to-action gaps across different targets shorter than your alert-evaluation window.

3. Alert fırtınası amplifikasyonu. Her remediation aksiyonu event üretir — dönen pod'lar, kopan bağlantılar. Naif bir dedektör kendi ilacının yan etkisini yeni hastalık diye okur ve dozu artırır. Bu, kimlik yazısındaki kaçak agent döngüsünün ops-aromalı kuzeni.

3. Alert-storm amplification. Every remediation action emits events — cycling pods, dropped connections. A naive detector reads its own medicine's side effects as new disease and increases the dose. This is the ops-flavored cousin of the runaway agent loop in the identity post.

4. Maskeleme. En sessizi, en pahalısı. Auto-restart yavaş bir leak'in üstünü aylarca örter; defect, restart'ın düzeltebileceği eşiği aşana kadar birikir — ve nihai incident, birikmiş versiyondur; üstelik bağlam buharlaşmıştır, çünkü küçükleri için kimseye page düşmedi. İmza: incident sayısı sabitken hedef başına remediation frekansının yukarı trendlenmesi. Tek metrik izleyecekseniz, bunu izleyin.

4. Masking. The quietest and the most expensive. Auto-restart papers over a slow leak for months; the defect compounds past what a restart can fix — and the incident you finally get is the accumulated version, with all context evaporated, because nobody was paged for the small ones. Signature: remediation frequency trending up per target while incident count stays flat. If you track one metric from this post, track that.

03State machine — vendor'ların atladığı iki state ileThe state machine — with the two states vendors skip

Sınırsız remediation bir reflekstir: detect → act. Sınırlı remediation, güvenliği iki üvey-evlat state'in taşıdığı bir pipeline'dır:

Unbounded remediation is a reflex: detect → act. Bounded remediation is a pipeline whose safety is carried by two under-loved states:

remediation state machineremediation state machine
Detect Diagnose Propose Policy Gatebütçe · cooldown · otonomi tavanıbudget · cooldown · autonomy ceiling Actsınırlıbounded Verifyhold: sinyal temiz + yeni alert yokhold: signal clear + no new alerts Close
gate.deny →Page human. Reddedilen öneri sessizce düşmez; insana devredilir.Page human. A denied proposal doesn't drop silently; it escalates.
verify.fail →Stop & page. "Daha iyi değil" retry sebebi değildir — dur ve devret.Stop & page. "Not better" is not a reason to retry — stop and hand over.
Şekil 1. Gate, confidence'ı değil policy'yi değerlendirir — her sorusu aksiyondan önce audit edilebilir. Verify "bitti"ye değil "daha iyi"ye bakar; onu besleyen telemetri de serinin observability tesisatının ta kendisi. Figure 1. The gate evaluates policy, not confidence — every question it asks is auditable before the action. Verify checks "better", not "done"; the telemetry feeding it is exactly this series' observability plumbing.
remediation/policy.yaml
# Bilerek sıkıcı, bilerek versiyonlu.# Deliberately boring, deliberately versioned.
budgets: { per_incident: 3, per_service_day: 10 }   # bitti → page; sessiz duraklama yok
actions:
  restart_pod:      { autonomy: full,  cooldown_s: 120, max_per_target_hour: 3 }
  scale_deployment: { autonomy: full,  bounds: { min: 1, max: 5 } }
  rollback_deploy:  { autonomy: gated, max_per_service_window: 1 }
  config_change:    { autonomy: propose_only }
  data_mutation:    { autonomy: forbidden }   # her confidence seviyesinde, sonsuza dek
verify: { hold_s: 20, signal: triggering_alert_cleared_and_no_new_alerts }
kill_switch: REMEDIATION_ENABLED                     # tek env var. Çevirmeyi tatbikatla öğrenin.# one env var. Learn to flip it by drilling it.

04Dört koşu, iki senaryo: ne çalıştırdıkFour runs, two scenarios: what we ran

Lab iki senaryo üretiyor. Dürüst arıza: bir pod gerçekten hasta, restart gerçekten düzeltir. Aldatıcı sinyal: pod'lar sağlıklı, bozuk olan ölçüm hattı — restart hiçbir şeyi düzeltmez, sadece pod döndürür. Dedektör ikisini ayırt edemez (gerçek dünyadaki gibi). İki döngü aynı sinyale bakıyor: naif (detect → act) ve guarded (yukarıdaki policy). Aksiyonlar gerçek: kind cluster'ında force-delete restart'lar.

The lab produces two scenarios. Honest failure: a pod is genuinely sick; a restart genuinely fixes it. Deceptive signal: the pods are healthy and the measurement pipeline is what's broken — a restart fixes nothing, it just cycles pods. The detector cannot tell them apart (just like the real world). Two loops watch the same signal: naive (detect → act) and guarded (the policy above). The actions are real: force-delete restarts on a kind cluster.

KoşuRunAksiyonActionsFlappingGate redGate denialsSonuçOutcome
Naif · dürüst arızaNaive · honest failure108,3 sn'de çözüldü — evet, guarded'dan hızlıresolved in 8.3 s — yes, faster than guarded
Guarded · dürüst arızaGuarded · honest failure10020,4 sn'de çözüldü — verify hold'u gerçek bir vergiresolved in 20.4 s — the verify hold is a real tax
Naif · aldatıcı sinyalNaive · deceptive signal1312asla çözülmedi, kimse page'lenmedi — 106 sn'de kesildi; kendi hâline bırakılsa sonsuza deknever resolved, nobody paged — cut off at 106 s; left alone, forever
Guarded · aldatıcı sinyalGuarded · deceptive signal10120,4 sn'de insana devrettiescalated to a human at 20.4 s

İlk satırı saklamıyoruz: dürüst arızada naif döngü daha hızlı (8,3'e karşı 20,4 sn) — verify hold'u bedava değil. Takasın tamamı üçüncü ve dördüncü satırda: aynı kör sinyale bakan naif döngü, sağlıklı bir servisi 8 saniyede bir yeniden başlatıp bunu başarı olarak logladı; guarded döngü bir kez denedi, "daha iyi değil" dedi, cooldown gate'ine çarptı ve 20 saniyede pager'ı çaldırdı. Hangi davranışı sabahın üçünde istersiniz?

We are not hiding the first row: on the honest failure, the naive loop is faster (8.3 vs 20.4 s) — the verify hold is not free. The whole trade lives in rows three and four: watching the same blind signal, the naive loop restarted a healthy service every 8 seconds and logged it as success; the guarded loop tried once, said "not better", hit the cooldown gate, and rang the pager at 20 seconds. Which behavior do you want at 3 a.m.?

Guarded döngünün aldatıcı sinyaldeki zaman çizelgesi, kelimesi kelimesine (results/lab-guarded-deceptive.log):

The guarded loop's timeline on the deceptive signal, verbatim (results/lab-guarded-deceptive.log):

timeline — guarded · aldatıcı sinyal (gerçek çıktı)timeline — guarded · deceptive signal (real output)
{"t_s":0,   "event":"fault.injected","note":"ölçüm hattı bozuk — pod'lar sağlıklı, sinyal yalancı""measurement pipeline broken — pods healthy, signal lying"}
{"t_s":0.1, "event":"detect",  "target":"mcp-server-9bc554744-975lp","health":0.2}
{"t_s":0.1, "event":"propose", "action":"restart_pod"}
{"t_s":0.1, "event":"gate.allow","action":"restart_pod"}
{"t_s":0.1, "event":"act",     "action":"restart_pod","n":1}
{"t_s":0.1, "event":"verify.start","hold_s":20}
{"t_s":20.4,"event":"verify.fail","note":"sinyal hâlâ kötü — aksiyon işe yaramadı""signal still bad — the action did not work"}
{"t_s":20.4,"event":"propose", "action":"restart_pod"}
{"t_s":20.4,"event":"gate.deny", "reason":"cooldown","target":"deploy/mcp-server"}
{"t_s":20.4,"event":"page.human","reason":"cooldown_hit_same_target","actions":1}

Bu on satır, yazının bütün tezidir: verify.fail retry üretmedi, gate.deny sessizce düşmedi. İkisi de aynı yere aktı — bir insanın bağlam hâlâ tazeyken haberdar edilmesine. Naif döngünün loglarında ise (aynı klasörde) 13 tane act satırı art arda duruyor; hepsi "başarılı".

These ten lines are the entire thesis: verify.fail did not produce a retry, and gate.deny did not drop silently. Both flowed to the same place — a human being informed while context was still fresh. The naive loop's log (same folder) shows 13 consecutive act lines; all "successful".

05Blast-radius matrisi: otonomi, sistemin değil aksiyonun özelliğidirThe blast-radius matrix: autonomy is a property of the action, not the system

"Remediation agent'ına ne kadar güveniyoruz?" yanlış soru — tek cevabı yok. İşleyen soru aksiyon sınıfı başına: bu aksiyon yanlışsa en kötü ne olur ve geri alınabilir mi?

"How much do we trust the remediation agent?" is the wrong question — it has no single answer. The workable question is per action class: what is the worst case if this exact action is wrong, and can it be undone?

Aksiyon sınıfıAction classYanlışsa en kötüWorst case if wrongGeri almaUndoOtonomi tavanıAutonomy ceiling
Tek pod restart'ıRestart one podKısa kapasite düşüşüBrief capacity dipKendi kendini geri alırSelf-undoingTam — cooldown + hedef başına tavanFull — with cooldown + per-target cap
Sınırlı bantta scaleScale within boundsMaliyet, hafif thrashCost, minor thrashÖnemsizTrivialTam — sınırlar zaten guardrailFull — the bounds are the guardrail
Deployment rollback'iDeployment rollbackDüzeltilen bug geri gelirReintroduces the fixed bugRoll forwardGate'li — pencere başına bir; sonrası pageGated — one per window, then page
Config/flag değişikliğiConfig/flag changeSessiz davranış kaymasıSilent behavior driftAncak versiyonluysaOnly if versionedYalnızca öneri — config declarative olana dekPropose-only until config is declarative
Failover / region kaydırmaFailover / region shiftÖlçekte kaskadCascade at scaleYavaş, riskliSlow, riskyYalnızca öneri, insan uygularPropose-only, a human executes
Veri mutasyonu / silmeData mutation / deletesGeri dönüşsüzIrreversibleYokNoneAsla. Review'la da olmaz, confidence skoruyla da.Never. Not with review, not with confidence scores.

Kendi matrisinizi yayınlayın. Matrisin en iyi işi runtime policy'si olmak değil, konuşmanın kendisi olmak: "agent incident'ları hallediyor" cümlesini "agent pod restart'lıyor ve rollback öneriyor" cümlesine zorlayan artefakt — ikincisi, bir on-call insanının gerçekten rıza gösterebileceği cümledir. Ve matrisi bütçelerle bağlayın: bunlar kimlik yazısındaki üç eksenli kotaların aynısı — remediation agent'ı, principal'ı on-call rotasyonu olan bir delegasyondan ibarettir ve her aksiyonu aynı run_id pivotuyla aynı audit akışına düşer.

Publish your matrix. Its best work is not as runtime policy but as the conversation itself: the artifact that forces "the agent handles incidents" to become "the agent restarts pods and proposes rollbacks" — the latter being a sentence an on-call human can actually consent to. And bind the matrix with budgets: they are the same three-axis quotas from the identity post — a remediation agent is just a delegation whose principal is the on-call rotation, and its every action lands in the same audit stream under the same run_id pivot.

06Otonominin kalıcı olarak bitmesi gereken yerWhere autonomy should end, permanently

Bazı sınırlar kademe kademe mezun olunacak seviyeler değildir. Geri dönüşsüz her şey (veri mutasyonu, yıkıcı temizlik), güvenliğe komşu her şey ("anomaliye" karşılık credential döndürmek, erişim değiştirmek — tebrikler, saldırganınızın artık remediation şeklinde bir silahı var) ve geri alması incident'tan yavaş her şey (region failover) — her confidence seviyesinde, sonsuza dek, yalnızca-öneri ya da yasak kalır. "Model genelde haklı" cümlesi dağılımın ortası hakkındadır; guardrail'ler kuyruklar için vardır ve incident'lar kuyruklarda yaşar.

Some boundaries are not tiers to graduate through. Anything irreversible (data mutation, destructive cleanup), anything security-adjacent (rotating credentials or changing access in response to "anomalies" — congratulations, your attacker now has a remediation-shaped weapon), and anything whose undo is slower than the incident (region failover) stays propose-only or forbidden, at any confidence, forever. "The model is usually right" is a statement about the distribution's middle; guardrails exist for its tails, and incidents live in the tails.

07Sık sorulanlarFAQ

Bu, fazladan adımlı Kubernetes self-healing'i değil mi?
K8s'in yerleşik döngüleri (restart, rescheduling, HPA) dar, iyi anlaşılmış controller'lar — sınırlı otomasyonun çalıştığının kanıtı zaten onlar. Bu yazı, çıkarsanmış teşhislerden keyfî aksiyonlar öneren yeni katman hakkında. Yenilik genişlikte; risk de genişlikte.

Isn't this just Kubernetes self-healing with extra steps?
K8s' built-in loops (restarts, rescheduling, HPA) are narrow, well-understood controllers — they are the proof that bounded automation works. This post is about the new layer that proposes arbitrary actions from inferred diagnoses. The novelty is the breadth; so is the risk.

Vendor'ımız agent'ının incident'ların %80'ini otonom çözdüğünü söylüyor.
Üç soru: hangi incident dağılımının %80'i? "Çözüldü" ne demek — stabil doğrulandı mı, alert mi kapandı? Ve geçen çeyrekte otonom bir aksiyonun ürettiği en kötü sonuç neydi? İlk iki cevap rakamı küçültür; asıl önemlisi üçüncüsü.

Our vendor says their agent resolves 80% of incidents autonomously.
Three questions: 80% of which incident distribution? What does "resolved" mean — verified stable, or alert closed? And what was the worst outcome an autonomous action produced last quarter? The first two shrink the number; the third is the one that matters.

Hiçbir şeyim yok; nereden başlarım?
Matrisin alt satırından yukarı: tam otonomi yalnızca kendini geri alan aksiyonlara, bütçe ve cooldown birinci günden, geri kalan her şey incident kanalınıza yalnızca-öneri. Üç hafta önerileri okumaktan, her tür değerlendirmeden daha çok şey öğrenirsiniz — kimsenin uygulamadığına sevindiğiniz öneriler dahil.

I have nothing today; where do I start?
From the matrix's bottom row up: full autonomy only for self-undoing actions, budgets and cooldowns from day one, everything else propose-only into your incident channel. Three weeks of reading proposals teaches more than any evaluation — including which proposals you are glad nobody executed.

LLM'in bu döngüde hiç mi yeri yok?
Teşhiste ve öneride var — incident kanıtını özetlemek bu modellerin gerçekten parladığı iş. Gate'te yok: gate policy'dir, policy koddur, kod review edilebilir. Yaratıcı yarı ile yetkili yarıyı o çizginin iki ayrı tarafında tutun.

Does the LLM belong in this loop at all?
In diagnosis and proposal, yes — summarizing incident evidence is genuinely where these models shine. In the gate, no: the gate is policy, policy is code, and code is reviewable. Keep the creative half and the authorized half on opposite sides of that line.

08Changelog

AIOpsSelf-HealingKubernetesGuardrailsSRE