<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Ranga's blog]]></title><description><![CDATA[Ranga's blog]]></description><link>https://ranga-blog.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Thu, 08 Oct 2026 13:48:28 GMT</lastBuildDate><atom:link href="https://ranga-blog.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Nobody Cares About Your Stack Until Something Breaks]]></title><description><![CDATA[Six things I learned writing and rewriting alert rules for systems that were already live, already watched, and still surprising people anytime
On Grafana alerting, PromQL, and the difference between ]]></description><link>https://ranga-blog.hashnode.dev/nobody-cares-about-your-stack-until-something-breaks</link><guid isPermaLink="true">https://ranga-blog.hashnode.dev/nobody-cares-about-your-stack-until-something-breaks</guid><category><![CDATA[alerting]]></category><category><![CDATA[#prometheus]]></category><category><![CDATA[Grafana]]></category><category><![CDATA[observability]]></category><dc:creator><![CDATA[ranga bashyam]]></dc:creator><pubDate>Sat, 19 Sep 2026 16:11:13 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6aaea132e6d52a47c061ede4/3c1b30b0-3a13-42f6-b581-ad950ec83825.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Six things I learned writing and rewriting alert rules for systems that were already live, already watched, and still surprising people anytime</h2>
<p>On Grafana alerting, PromQL, and the difference between a system that is monitored and a system that can actually speak.</p>
<p>A while back I picked up a task that sounded boring on the ticket: <em>fill the observability gaps</em>.</p>
<p>There was already an internal alerting platform. Built, deployed, service teams onboarded, dashboards everywhere. On paper we were covered. In practice, people were still opening incidents by hand, someone in a Slack channel typing "is checkout slow for anyone else?" twenty minutes before anything fired. And in other corners of the estate there was nothing at all. Just logs. Or just events. Or a service that had been running in production for two years emitting absolutely nothing except stdout.</p>
<p>That's the job most of us actually get. Not a greenfield platform. A half-instrumented estate where the tools are fine and the rules are wrong.</p>
<p>Here's the thing I keep coming back to: <strong>your system is observed exactly as well as your alert configs are written.</strong> Not as well as your stack is built. You can run Prometheus, Thanos, Mimir, OTel, Alloy, the full pipeline, and still be blind, because blindness and health look identical on a dashboard nobody is staring at.</p>
<p>These are the six things that cost me the most to learn. Written for you so they cost you less.</p>
<h2>01 Start with the manual incidents, not the shiny service</h2>
<p>When you're handed "fill the gaps," the instinct is to open a blank rule file and start writing good alerts. Don't. You already have a dataset telling you exactly where the holes are, and it's free.</p>
<p>Pull every manually-raised incident from the last three to six months and sort them into four buckets:</p>
<ul>
<li><p><strong>An alert existed and fired late.</strong> Your <code>for:</code> duration or threshold is wrong. Tuning job.</p>
</li>
<li><p><strong>An alert existed, fired on time, went nowhere.</strong> Routing problem. Labels, notification policy, a Slack channel nobody is in. The query was never the issue.</p>
</li>
<li><p><strong>An alert existed, fired on time, got ignored.</strong> This one hurts, because it means the alert is buried in noise. Go fix noise before you write anything new.</p>
</li>
<li><p><strong>No alert existed.</strong> The actual gap.</p>
</li>
</ul>
<p>The distribution across those four buckets is your roadmap, and it's almost never what you assumed. On my run, bucket four was the smallest. Most of the pain was buckets two and three, alerts that technically worked and practically didn't.</p>
<p>The second thing to do before writing anything: inventory what signal actually exists per service. Metrics? Logs only? Events only? Nothing? Because you cannot alert on what isn't emitted, and the honest first ticket for half those services was "add a counter," not "add an alert." That's not a detour. That <em>is</em> the gap.</p>
<blockquote>
<p><strong>Worth the effort</strong></p>
<p>Every manual incident is a free bug report on your alerting, written by someone who was there. It's the highest-signal input you will ever get, and most teams throw it away by closing the ticket and moving on.</p>
</blockquote>
<h2>02 Alerts overfit and underfit, exactly like models do</h2>
<p>This framing helped me more than any monitoring blog post, so I'll hand it over.</p>
<p><strong>Underfit alerts</strong> are the ones everyone complains about. Threshold too loose, window too short, scope too wide. <code>CPU &gt; 90% for 1m</code> on every pod in the fleet. It fires constantly, it's right roughly never, and after three weeks your on-call has a Slack filter that hides it. That alert is now worse than having no alert, because the org believes it's covered.</p>
<p><strong>Overfit alerts</strong> are the quiet ones and almost nobody talks about them. Someone had a bad Tuesday, did a postmortem, and wrote a rule so precisely shaped around that one incident that it will never fire again:</p>
<pre><code class="language-plaintext"># tuned to a single incident that will never repeat in this exact shape
avg by (pod) (rate(container_cpu_usage_seconds_total{
    namespace="payments",
    pod=~"checkout-v2-.*",
    node="node-17"
}[3m])) &gt; 0.937
</code></pre>
<p>It looks fantastic in the repo. It is dead weight. Worse, it produces false confidence "we have an alert for that" and nobody checks that the alert is structurally incapable of firing.</p>
<p>The tell is easy to check. Look at your firing history and find every rule with zero firings in six months. Each one is either overfit, broken, or genuinely guarding something you never breach. It is almost never the third.</p>
<p>The middle ground is boring and correct: <strong>alert on a symptom your user actually feels, scoped to a level you can actually act on.</strong> Not CPU. Not one pod. Error ratio per service, latency per endpoint, queue not draining, data gone stale.</p>
<blockquote>
<p><strong>Careful here</strong></p>
<p>Adding labels to a noisy query <em>feels</em> like tuning. Sometimes it's just overfitting with extra steps. Ask yourself honestly: am I narrowing the scope because this alert genuinely only matters in this scope, or because I want the pages to stop? Those are very different commits with very similar diffs.</p>
</blockquote>
<h2>03 Noise isn't too many alerts. It's the same truth told too many times</h2>
<p>This distinction took me embarrassingly long. Two hundred pods failing the same way is not two hundred problems. It's one problem, reported two hundred times. Fixing noise is mostly about collapsing repetition, and PromQL gives you the tools if you're deliberate about it.</p>
<h3>Aggregate to the level you'd actually act on</h3>
<pre><code class="language-plaintext"># noisy: one series per pod. 200 pods = 200 notifications
rate(http_requests_total{status=~"5.."}[5m]) &gt; 0

# what you actually want to know
sum by (cluster, service) (rate(http_requests_total{status=~"5.."}[5m]))
  /
sum by (cluster, service) (rate(http_requests_total[5m]))
  &gt; 0.05
</code></pre>
<p>Aggregating up throws away the pod label, and that's fine. The pod label belongs in the dashboard you open <em>after</em> you're paged, not in the page itself.</p>
<h3>Use <code>group_left()</code> to carry context without fanning out</h3>
<p>The usual objection to aggregating is "but then I lose the team and ownership labels I need for routing." You don't. Join them back from a metadata metric:</p>
<pre><code class="language-plaintext">(
  sum by (cluster, service) (rate(http_requests_total{status=~"5.."}[5m]))
    /
  sum by (cluster, service) (rate(http_requests_total[5m]))
    &gt; 0.05
)
* on (cluster, service) group_left(team, tier, runbook)
  service_metadata{env="prod"}
</code></pre>
<p>One series per service, carrying <code>team</code>, <code>tier</code> and <code>runbook</code> so your Grafana notification policy can route it properly. This single pattern removed more noise for me than any threshold change, because it let me aggregate aggressively without losing the routing I needed.</p>
<h3>Match <code>for:</code> to how long the system actually takes to heal itself</h3>
<p>Here's the one that got me, and it's the reason for the card at the top of this page.</p>
<p>An AC power reclaim in the datacenter. Power gets pulled and restored on a rack, nodes go down, nodes come back, retries fire, things settle. Normal, planned, understood by everyone who has been there longer than me. Total time to settle: ten to fifteen minutes.</p>
<p>My rule had <code>for: 5m</code>.</p>
<p>So every reclaim produced a burst of pages that resolved themselves, then fired again as the next batch of nodes cycled, then resolved, then fired. Four events that mattered turned into somewhere north of twenty notifications, in the middle of the night, all of them technically correct and all of them useless.</p>
<p>The fix is not clever. If the system genuinely recovers on its own in fifteen minutes, then five minutes is not an incident, it's a stage of normal operation, and your rule is reporting on the weather.</p>
<pre><code class="language-plaintext">- alert: NodeUnreachableAfterPowerEvent
  expr: up{job="node-exporter"} == 0
  for: 20m
  keep_firing_for: 10m
  labels:
    severity: critical
    team: infra
  annotations:
    summary: "{{ $labels.instance }} unreachable for 20m"
    description: "Past the normal reclaim window. Check power and BMC before assuming host failure. {{ $value }} instances affected."
    runbook_url: "https://runbooks.internal/node-down"
</code></pre>
<p><code>keep_firing_for</code> is the underrated half of that. It holds the alert open for a grace period after the condition clears, which kills the fire/resolve/fire/resolve flapping that makes an incident channel unreadable.</p>
<blockquote>
<p><strong>This is where you get the heat</strong></p>
<p>If you're junior and you're about to widen a page window from 5m to 20m, do not do it alone. Get your senior, your manager, or the service owner to say yes, in writing, in a thread you can find later.</p>
<p>Because the query is the easy part. The hard part is the day something <em>doesn't</em> self-heal in twenty minutes, and the first question in the postmortem is "who widened this?" You want that thread. I promise you want that thread. Nobody blames you for asking; everybody remembers when you didn't.</p>
</blockquote>
<h3>Then let Grafana collapse the rest</h3>
<p>Query-level work gets you most of the way. Notification policies do the last mile. Group by the labels that define the incident, not the ones that define the instance:</p>
<ul>
<li><p><code>group_by: [alertname, cluster, service]</code> : fifty firing series arrive as one notification.</p>
</li>
<li><p><code>group_wait</code> : hold briefly so related alerts from the same event land together instead of trickling in.</p>
</li>
<li><p><code>repeat_interval</code> :if you set this to five minutes you have built a very expensive way to teach people to mute you.</p>
</li>
<li><p><strong>Mute timings</strong> for known windows. Nightly batch spikes consumer lag every night at 01:00? That's a schedule, not an incident. Mute it and alert on the batch failing instead.</p>
</li>
</ul>
<h2>04 Writing from scratch: watch the system before you threshold it</h2>
<p>For the corners with nothing at all, the temptation is to copy a starter rule pack and move on. Resist for about two weeks. That's roughly what it costs to write alerts that survive contact with production.</p>
<p>The order that worked for me:</p>
<p><strong>Find the symptom, not the cause.</strong> Ask the service team one question: "what does a user notice when this is broken?" The answers are things like <em>payments fail</em>, <em>the queue stops draining</em>, <em>the nightly reconciliation didn't run</em>, <em>the dashboard is showing yesterday's data</em>. None of those are CPU. Cause-based alerts multiply endlessly because there are infinite ways to break; symptom-based alerts stay small because there are only a few ways to be broken.</p>
<p><strong>Get the metric emitted if it doesn't exist.</strong> Often the whole ticket. A counter and a last-success timestamp will carry you further than any amount of log parsing.</p>
<p><strong>Then sit in Grafana Explore for a while.</strong> Thirty days back. What does a normal Tuesday look like? What does deploy day look like? What does month-end look like? What does that 03:00 spike turn out to be? Pick your threshold from <em>your</em> data, not from a blog post that confidently said five percent.</p>
<p><strong>Prefer thresholds that age well.</strong> Static numbers rot. <code>disk &gt; 85%</code> is meaningless on a 20 TB volume and screaming on a 50 GB one. Alert on trajectory instead:</p>
<pre><code class="language-plaintext"># will we run out in the next 4 hours, and are we already low?
predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4*3600) &lt; 0
and
node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
  / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} &lt; 0.15
</code></pre>
<p>That second clause matters. <code>predict_linear</code> alone will page you because someone wrote a large temp file for ten minutes.</p>
<h3>The alert that never fires because your exporter died</h3>
<p>This is the one I'd tattoo on people. A rule evaluating over a metric that has stopped arriving does not fire. It evaluates to nothing and sits there, silently, permanently green.</p>
<p><strong>Green because healthy and green because blind look exactly the same in Grafana.</strong> There is no visual difference. You will not notice.</p>
<pre><code class="language-plaintext"># is the thing that watches the thing still alive?
absent(up{job="payments-exporter"} == 1)

# did the nightly job actually succeed, or did it just stop reporting?
time() - batch_job_last_success_timestamp_seconds{job="nightly-recon"} &gt; 26*3600
or
absent_over_time(batch_job_last_success_timestamp_seconds{job="nightly-recon"}[26h])
</code></pre>
<p>Every critical path needs one of these. It's the cheapest alert you will ever write and it is the one that catches the failure mode nobody plans for.</p>
<h2>05 A dashboard panel is not an alert rule, and copy-pasting one into the other will hurt you</h2>
<p>Everyone who moves from Grafana dashboards into Grafana alerting does this once. Here's your once, for free.</p>
<p><strong>Template variables don't exist in alert rules.</strong> Your panel query is full of <code>$namespace</code> and <code>$__rate_interval</code>. Those are dashboard-time substitutions. An alert rule has no dashboard and no time picker.</p>
<pre><code class="language-plaintext"># in a panel: fine
sum(rate(http_request_duration_seconds_count{namespace="$namespace"}[$__rate_interval]))

# in an alert rule: pin everything explicitly
sum by (namespace) (rate(http_request_duration_seconds_count[5m]))
</code></pre>
<p><strong>Panels are range queries; rules are instant queries on a schedule.</strong> A panel shows you a spike because you're looking at a window. A rule evaluating every minute with <code>for: 5m</code> may never see that spike at all. If the panel and the alert disagree, the panel is usually not the one that's wrong, but it's also not the one that pages anybody.</p>
<h3>Choose your rule type on purpose</h3>
<ul>
<li><p><strong>Grafana-managed rules</strong> give you the expression pipeline (query → reduce → threshold), multi-datasource conditions, and screenshots in notifications. Great for teams who live in the UI.</p>
</li>
<li><p><strong>Data-source-managed rules</strong> (the Prometheus or Mimir ruler) live as YAML in the same repo as your recording rules. Reviewable, diffable, and this is the part people forget <em>they keep evaluating when Grafana is down.</em></p>
</li>
</ul>
<p>Think carefully about where the rules that watch your monitoring stack should live. Putting the alert for "Grafana is unhealthy" inside Grafana is a joke that stops being funny at 02:00.</p>
<h3>No Data and Error handling is not a footnote</h3>
<p>It's the setting at the bottom of the form that people scroll past, and the default is to treat it as alerting. Which means a scrape gap, a datasource timeout, or a Mimir hiccup becomes a page and worse, it becomes a page on <em>every rule at once</em>, which is how a five-minute blip turns into an incident channel with four hundred messages in it.</p>
<p>Decide per rule:</p>
<ul>
<li><p>Payment path stops reporting? <strong>No Data should page.</strong> Silence there is the emergency.</p>
</li>
<li><p>A batch job that only runs 09:00–18:00? No Data overnight is <strong>normal</strong>. Set it to OK or you've just built a nightly alarm clock.</p>
</li>
<li><p>Error state on a flaky remote datasource? Usually OK, with a separate dedicated alert on the datasource itself.</p>
</li>
</ul>
<h3>Precompute anything expensive</h3>
<p>If your alert query takes eight seconds to evaluate on a quiet afternoon, it will time out under load, which is precisely when you need it to work. Push the heavy aggregation into a recording rule and let the alert read a cheap series:</p>
<pre><code class="language-plaintext"># recording rule: evaluated once, reused by every alert and panel
- record: slo:http_error_ratio:rate5m
  expr: sum by (cluster, service) (rate(http_requests_total{status=~"5.."}[5m]))
        / sum by (cluster, service) (rate(http_requests_total[5m]))

# fast-burn alert: budget being consumed 14.4x too quickly,
# confirmed over both a short and a long window so one blip can't trigger it
- alert: CheckoutErrorBudgetFastBurn
  expr: slo:http_error_ratio:rate5m{service="checkout"} &gt; 14.4 * 0.001
        and slo:http_error_ratio:rate1h{service="checkout"} &gt; 14.4 * 0.001
  for: 2m
</code></pre>
<p>The two-window trick is the same noise-reduction idea from section three, applied to SLOs: a short window makes it fast, a long window makes it real, and requiring both means a thirty-second blip doesn't wake anyone.</p>
<h2>06 An alert with no owner and no next step is just a notification</h2>
<p>You can write the best PromQL of your career and still ship something worthless. The query is maybe sixty percent of an alert. Here's the rest.</p>
<p><strong>Labels are routing, and routing is the whole point.</strong> <code>team</code>, <code>service</code>, <code>env</code>, <code>severity</code>. Get one wrong and your perfect rule fires into a Slack channel that was archived last quarter. That's not a monitoring failure anyone will catch, it looks exactly like the alert working. Test the route, don't assume it.</p>
<p><strong>Annotations are what the tired person reads.</strong> At 2 a.m. nobody parses a PromQL expression. Write the summary as if you're waking up a colleague who has never seen this service:</p>
<ul>
<li><p><code>summary</code> : what broke and where, in one line, with the labels interpolated.</p>
</li>
<li><p><code>description</code> : the actual numbers. <code>{{ $value }}</code>, <code>{{ $labels.instance }}</code>. "Error ratio 12%" beats "error ratio high" every single time.</p>
</li>
<li><p><code>runbook_url</code> : even if the runbook is three bullets in a wiki. Three bullets beats opening Grafana and guessing.</p>
</li>
</ul>
<p><strong>Severity has to mean something.</strong> If everything is critical, nothing is. My working rule: critical means a human gets woken up and there is something that human can do right now. If you can't name the action, it isn't critical, it's a warning, and warnings can wait for business hours.</p>
<p><strong>Review on a schedule, and delete things.</strong> Once a month, three lists: what fired the most, what fired and got closed with no action taken, and what never fired at all. The middle list is your noise. The last list is your overfitting. Deleting a dead alert is real engineering work, even though it feels like you're removing coverage. You're removing the <em>illusion</em> of coverage, which is more dangerous than the gap.</p>
<blockquote>
<p><strong>The social part nobody puts in the runbook</strong></p>
<p>When you change someone else's alert, tell them. You are modifying the thing that wakes them up at night. Even a one-line message in their channel "widening the node-down window to 20m because of DC reclaims, shout if that's wrong for you" buys you enormous goodwill and occasionally saves you from a mistake, because they know something about their service that you don't.</p>
</blockquote>
<hr />
<h2>What this is actually for</h2>
<p>By the end of that quarter the numbers were fine. Noise down, coverage up, manual incidents way down. But the thing that stuck with me wasn't a metric.</p>
<p>You can build the most beautiful observability platform in the world. Prometheus scraping cleanly, Thanos and Mimir holding years of history, OTel traces threaded end to end, Alloy collecting everything, pipelines tuned, a model summarizing your incidents. Genuinely impressive work, and I'd enjoy building every piece of it.</p>
<p>And if, at 02:14, checkout is failing and nobody gets told none of it counted. Not one line.</p>
<p>Everything else in observability is how you <em>investigate</em> a problem. Alerting is how you <em>find out</em> there is one. It's the only part of the stack that reaches outward and taps a human on the shoulder, and it's the only part anyone outside your team will ever experience.</p>
<p>So write the boring rules first. Get the windows right, get the routing right, get the runbook links in. Then go build the rest of the platform, knowing your system can speak when it needs to.</p>
<hr />
<p>If you take one thing: go find every rule that hasn't fired in six months, and check whether it <em>can</em>. That afternoon will teach you more about your alerting than this whole article.</p>
]]></content:encoded></item></channel></rss>