Skip to main content
All of the catalog
Scenario

Observability & SRE

Full observability stack with log aggregation, distributed tracing, alerting rules, and SLO dashboards.

ObservabilityVerifiedk3dkind
Definition on GitHub

What you'll do

  • Aggregate application logs in Loki and query them from Grafana
  • Export traces through a collector (Grafana Alloy) into Tempo, instead of pointing the app at the backend
  • Correlate the three signals: jump from a log line to its trace, and from a span back to that pod's logs
  • Drive the stack with real load from k6 and compare client-side against server-side request metrics
  • Alert on high error rate / latency and track SLOs on dashboards

Stages

  1. 1log-aggregation

    Ship and store application logs

    lokipromtail
  2. 2tracing

    Trace collector (Alloy) and trace backend (Tempo)

    tempoalloywire-app-to-collector
  3. 3alerting-and-slo

    Alert rules and SLO dashboards

    alerting-rulesslo-dashboards

Prerequisites

These are installed into the lab cluster for you — listed so you know what the scenario actually depends on.

ingressmonitoring/metricsmonitoring/grafanago-api

The incident field notes

One real Kubernetes failure a week — the symptom, the commands that found it, and the fix. Written from actual lab runs, not from memory.

You'll get the Kubernetes Incident Response Field Guide, plus occasional emails about new scenarios, posts and paid offerings such as courses and workshops. Unsubscribe any time.