SLO and Error Budget Designer (YAML)

💻 PROGRAMMING · DevOps advanced ⭐ 92

Design user-centric SLIs, SLOs, error budgets, burn-rate alerts, and an error budget policy in YAML.

Заполните поля:

Значения подставятся в промпт ниже.

role: >
 You are a senior Site Reliability Engineer who designs Service Level Objectives
 (SLOs) that match what users actually experience. You favor a few meaningful
 objectives over many vanity metrics, and you turn every SLO into an error
 budget policy and alerts a team can act on.

task: >
 Design user-centric SLIs, SLOs, error budgets, burn-rate alerts, and an error
 budget policy for the service described in the inputs. Then return the result
 in the exact YAML output schema below.

inputs:
 service_name: "${service_name:checkout-api}"
 what_users_do: "${user_journeys:browse cart, apply coupon, place order, view order status}"
 architecture: "${architecture:Node.js API behind a load balancer, PostgreSQL, Redis cache, Stripe for payments}"
 traffic: "${traffic:about 40 requests per second at peak, strong evening and weekend peaks}"
 telemetry_available: "${telemetry:load balancer access logs, Prometheus metrics, OpenTelemetry traces}"
 current_pain: "${pain:slow order placement during sales, on-call fatigue from noisy CPU alerts}"
 business_constraints: "${constraints:payments must never be double-charged; marketing runs flash sales monthly}"
 compliance_window_days: ${window_days:28}

method:
 - step: Map critical user journeys
 detail: List 3-5 journeys from what_users_do and rank them by business impact. Ignore internal-only endpoints unless a user waits on them.
 - step: Choose SLIs
 detail: For each top journey, pick availability, latency, or correctness/freshness SLIs written as "good events / valid events". Name the exact data source and say where it is measured (load balancer, server, or client).
 - step: Set targets
 detail: Propose an SLO target and justify it against current pain, dependencies (a service cannot beat its critical dependencies), and cost. Prefer 99.5-99.9% unless there is strong evidence otherwise. Never 100%.
 - step: Compute error budgets
 detail: Convert each target into allowed bad events and allowed bad minutes for the compliance window, and show the arithmetic.
 - step: Design burn-rate alerts
 detail: Use multi-window, multi-burn-rate alerting (for example a 1h/5m window pair at 14.4x paging and a 6h/30m pair at 6x paging, plus a 3d/6h pair at 1x as a ticket). Write example PromQL or pseudo-queries using the telemetry the team actually has.
 - step: Write the error budget policy
 detail: Define concrete actions at 50%, 75%, and 100% budget consumed (for example a release freeze except fixes, or a reliability sprint), who decides, and how exceptions work.
 - step: Retire noise
 detail: Name the current alerts that should be downgraded or deleted once SLO alerts exist, such as raw CPU alerts.

output_format:
 type: yaml
 rules:
 - Return only valid YAML that follows the schema below. No prose before or after it and no code fences.
 - Use null for unknown values and record every guess in assumptions.
 - Keep each description under 25 words.
 schema:
 service: string
 compliance_window_days: integer
 assumptions: [string]
 user_journeys:
 - name: string
 business_impact: high | medium | low
 slos:
 - id: string (e.g. SLO-1)
 journey: string
 sli_type: availability | latency | correctness | freshness
 sli_definition: "good events / valid events, in plain words"
 good_event: string
 valid_event: string
 data_source: string
 measured_at: load_balancer | server | client | synthetic
 target_percent: number
 latency_threshold_ms: integer or null
 error_budget:
 allowed_bad_ratio: number
 allowed_bad_events_estimate: integer
 allowed_bad_minutes: number
 math: string
 burn_rate_alerts:
 - name: string
 long_window: string
 short_window: string
 burn_rate: number
 budget_consumed_at_trigger_percent: number
 action: page | ticket
 query_example: string
 error_budget_policy:
 thresholds:
 - budget_consumed_percent: integer
 actions: [string]
 decision_owner: string
 exception_process: string
 alerts_to_retire:
 - alert: string
 reason: string
 dashboards:
 - panel: string
 purpose: string
 review_cadence: string
 open_questions: [string]

constraints:
 - Maximum 4 SLOs. If more seem necessary, explain the merge choice in assumptions.
 - Every SLO must be measurable with the telemetry the team says it has; otherwise list the missing instrumentation in open_questions.
 - Do not use CPU, memory, or other resource metrics as SLIs.
 - Latency SLIs use a threshold ("requests faster than 800 ms"), not an average.
 - Treat payment correctness (no double charge) as its own SLO or as an explicit invariant in assumptions.
#slo#sre#error-budget#yaml#alerts

Кураторская подборка fizoni.com · структура, примеры использования и рабочие сценарии

Похожие промпты

💡 Понравился этот промпт?

📦 Prompts for Programmers — 448 промптов

Вместо одного промпта — целый набор по теме. Готовые процессы, структура, быстрый старт. Код, рефакторинг, дебаг, архитектура, ревью, SQL и API.…

Забрать набор — $2.99 🎁 6 бесплатных промптов

Всего $2.99 · мгновенный доступ · оплата картой

🎁 Забери 6 лучших промптов бесплатно

Понравился этот промпт? Мы собрали ещё 6 отборных — для бизнеса, кода и продуктивности. Плюс доступ к полной библиотеке 2000+ промптов.