Skip to content

LOOPIN-74 — Add Grafana Dashboards, Prometheus Alert Rules and Operational Runbooks #157

Description

@shaig-mahmudov

Type: Observability / Operations
Priority: Medium

Description

Create version-controlled Grafana dashboards and actionable Prometheus alerts for the Loopin API.

Scope

  • Configure Grafana with Prometheus as a data source.
  • Add a Loopin API dashboard.
  • Include request rate, error rate, latency, JVM, HikariCP, WebSocket, notification, AI and media metrics.
  • Add alert rules for critical operational conditions.
  • Add a short runbook for each alert.
  • Store dashboard and provisioning files in the repository.

Acceptance Criteria

  • Grafana loads the Loopin dashboard automatically.
  • Dashboard panels display data from Prometheus.
  • Alerts cover readiness failure, high 5xx rate, high latency, database pool saturation and background-operation failures.
  • Each alert includes a clear description and remediation note.
  • Dashboards and alert rules are version controlled.
  • No dashboard depends on high-cardinality identifiers.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

  • Status
    Backlog

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions