Zurück zum Ranking

SquadcastHub/awesome-sre-tools

A curated list of Site Reliability and Production Engineering Tools

sredevopssite-reliability-engineeringproductionavailabilitymonitoringpost-mortemreliability-engineeringreliabilitymonitoring-toolsservice-level-agreementservice-level-objective
Sterne-Wachstum
Sterne
1.5k
Forks
243
Wochenwachstum
Issues
0
5001k
Apr. 2020Mai 2022Juni 2024Juli 2026
README

Awesome Site Reliability Engineering Tools Awesome

A curated list of Site Reliability and Production Engineering tools - Maintained by Raghu Chinnannan and Squadcast

Contents

Development

Source Code Management

Project Management & Issue Tracking Software

Bug / Defect Tracking Software

Code Editors and IDEs

Continuous Testing

Continuous Integration

Build

Integration

Continuous Delivery

Deployment

Infrastructure orchestration

Container

Container Registry

Container Orchestration

Continuous Monitoring

  • AWS CloudWatch
  • DebugBear
  • net-benchmark - DNS/HTTP/SSL benchmarking with CSV, Excel, PDF, and JSON exports.
  • Prometheus
  • StackDriver
  • Sensu
  • Sentry
  • CopperEgg
  • Crashlytics
  • Kapacitor
  • loggly
  • logmatic
  • Logstash
  • MongoDB Atlas
  • MongoDB Cloud Manager
  • NewRelic
  • ReleaseRun Vulnerability Scanner
  • Papertrail
  • PageGuard - Free all-in-one website health scanner. Core Web Vitals, SEO, WCAG 2.1 accessibility, and best practices. AI-generated action plan. No signup required.
  • Pingdom
  • ServerDensity
  • Zabbix
  • InsightOps
  • AppSignal
  • API Status Check - Centralized dashboard tracking real-time status and outages for 1,000+ popular APIs and services (AWS, Stripe, GitHub, Twilio, etc.). Monitor third-party dependencies, get instant outage alerts, reduce MTTR.
  • Grafana
  • VictoriaMetrics
  • Chaos Genius
  • Cloud Waste Scanner - Detects cloud waste and helps DevOps/platform teams identify quick cloud cost optimization opportunities.
  • Thanos
  • Mimir
  • Hydrozen.io - Uptime monitoring & Statuspages
  • SSL Certificate Monitor - Open-source SSL/TLS certificate expiry monitoring tool with email alerts
  • DNS Propagation Checker - Open-source DNS propagation monitoring tool with global DNS server coverage
  • whatbroke.today - AI-powered outage aggregator tracking 100+ cloud services with Telegram alerts
  • Prismix - Real-time status dashboard for 75+ AI services (OpenAI, Anthropic, Gemini, Mistral, etc.) with starring, email/webhook alerts, 30-day uptime history, and embeddable SVG badges.
  • Steampipe.io - Universal SQL interface to any cloud API
  • Better Stack
  • Netdata
  • DoctorGPT - Brings GPT into production for application log error monitoring
  • Dynatrace
  • Datadog
  • DevHelm - Developer-first uptime monitoring with HTTP, DNS, TCP, ICMP, and heartbeat checks, dependency intelligence for 80+ providers, hosted status pages, incident management, and a full developer surface (CLI, SDKs, Terraform provider, MCP server).
  • Elastic APM
  • Healthchecks.io
  • OnlineOrNot - Uptime monitoring for websites, APIs, and cron jobs, with integrated status pages.
  • Uptrack - Uptime monitoring with 30-second checks on free tier, consecutive-check alert confirmation to cut false positives, hosted status pages, and a built-in MCP server for AI agents.
  • Streamdal - Code-Native Data Privacy - embed privacy controls in your application code to detect and monitor PII. Streamdal
  • Dash0 - OpenTelemetry Native Observability, built on CNCF Open Standards such as PromQL, Perses and OTLP with full cost control. Supporting Metrics, Traces and Logs with full custom dashboarding and alerting capabilities.
  • CICube - AI DevOps monitoring platform by monitoring your CI workflows, detect anomalies, and provide actionable fixes.
  • Middleware - A Full-Stack Cloud Observability Platform designed to empower developers and organizations to monitor, optimize, and streamline their applications and infrastructure in real-time.
  • Shipfox - Boost GitHub Actions speed by 2x and cut costs by up to 75%, with smarter caching, deep CI insights, and zero-config setup.
  • Ingero - eBPF-based GPU causal observability agent. Traces CUDA APIs and host kernel events to build causal chains explaining GPU latency. Includes MCP server for AI-assisted incident investigation.
  • cloud-audit - AWS security auditing CLI that runs 17 checks across IAM, S3, EC2, VPC, and RDS with built-in remediation engine generating AWS CLI commands and Terraform snippets.
  • FlareWarden - Uptime, content, and dependency monitoring with multi-region verification, status pages, and incident management.
  • Phare - Shockingly good uptime monitoring, alerts, incident management, and status pages.
  • API Status Check - Real-time status monitoring dashboard for 250+ developer APIs including AWS, Stripe, GitHub, and OpenAI. Free, no signup required.
  • LynxDB - Lightweight columnar log analytics database for SRE workflows, with a pipe-style query language inspired by SPL for investigating production logs.
  • KubeStellar Console - Open-source multi-cluster Kubernetes dashboard with AI-powered operations, MCP server bridging kubeconfig to LLM agents, and real-time observability across edge and cloud clusters. CNCF Sandbox. KubeStellar Console
  • Apitally - API monitoring, analytics, and request logging for REST APIs, with lightweight open-source SDKs for Python, Node.js, Go, .NET, and Java.
  • Riftmap - Cross-repo infrastructure dependency discovery and change impact analysis for multi-repo environments using Terraform, Docker, Helm, and more.
  • Oack - HTTP monitoring with TCP kernel telemetry, 6-phase latency breakdown, Server-Timing header capture, Cloudflare CDN enrichment, and built-in incident management with on-call scheduling.
  • OpenClaw Monitor - Real-time AI agent monitoring dashboard for OpenClaw agents. Track Gateway status, sessions, token usage & trends.
  • agenttrace - TUI observability for AI coding agents. Track cost, tokens, tool failures, latency, anomalies, health, diffs, and CI gates across Claude Code, Codex CLI, Gemini CLI, Aider, and Cursor exports.
  • sunwatch - Crypto-paid uptime monitoring for side projects. Pay per monitor with USDC on Base; webhook alerts on down/up state changes.
  • Drumbeats - Cron, heartbeat, and HTTP uptime monitoring for background jobs and services, with concurrent-job (run_id) correlation, duration/hang alerts, LOG pings for mid-run progress, incident management, and status pages. One curl ping instruments a job; no agent or SDK. Free tier: 50 monitors, 200K Beats/mo, all notification channels.
  • OpenChainBench - Continuous monitoring of blockchain RPC providers, bridges and oracles. Multi-region Prometheus probes of latency, tx-landing success and finality with public dashboards and open methodology.
  • Respan - Observability platform for LLM and AI agent applications, with tracing, evals, prompt management, and a gateway across 250+ models.
  • Faultline - Deterministic CI failure analysis CLI that classifies build logs into explainable failure types with evidence and fix steps.
  • Oh Dear - Monitoring for uptime, performance, broken links, SSL certificates, and DNS, with hosted status pages.
  • Yorker - OpenTelemetry-native synthetic monitoring with HTTP and Playwright browser checks, monitoring-as-code via YAML and CLI, and enriched OTLP export to any OTel backend.

Incident Management / Incident Response / IT Alerting / On-Call

IT Service Management

Incident Communication

Internal Developer Portal

AI SRE Tools & SRE Copilots

  • Sherlocks.ai
  • Resolve.ai
  • Deductive.ai
  • Ingero - eBPF-based GPU causal observability agent. Traces CUDA APIs and host kernel events to build causal chains explaining GPU latency. Includes MCP server for AI-assisted incident investigation.
  • IncidentFox (open source)
  • metoro.io
  • Ops AI by Middleware
  • tailscale-mcp - MCP server with 52 tools for managing Tailscale tailnets from AI assistants like Claude Code and Cursor.
  • KubeStellar Console - AI-powered multi-cluster Kubernetes management console with MCP server (kc-agent) for AI-assisted cluster operations, pod inspection, deployment management, and real-time observability across distributed environments.
  • Cynative - Deep research agent for your infra - sandboxed, read-only, covers AWS, GCP, Azure, Kubernetes, GitHub and GitLab.
  • Aurora - Open source (Apache 2.0) AI SRE agent that autonomously investigates incidents and performs root cause analysis across AWS, Azure, GCP, and Kubernetes. Self-hosted via Docker Compose or Helm, works with major LLM providers or local models via Ollama.
  • Anyshift - AI SRE built on a versioned resource graph of your infrastructure, for root cause analysis and predicting the impact of changes before they ship.
  • Radar - Open source Kubernetes visibility tool with a built-in MCP server for AI-assisted cluster operations — topology, service traffic, events, logs, and a 31-check best-practices audit.
  • KnoxOps - AI-native ops agent that gives agents production-safe execution with human review and a built-in knowledge graph.
  • NudgeBee - Unified AI agentic platform for cloud ops, offering AI SRE, AI FinOps, AI Kubernetes Ops, and AI CloudOps assistants that automate alert triage, root-cause analysis, and cost optimization.
  • Hyground - Self-hosted AI SRE agent that goes beyond on-call incident resolution.

Stargazers over time

Stargazers over time

Licence

Shield: CC BY 4.0

This work is licensed under a Creative Commons Attribution 4.0 International License.

CC BY 4.0

Ähnliche Repositories
bregman-arie/devops-exercises

Linux, Jenkins, AWS, SRE, Prometheus, Docker, Python, Ansible, Git, Kubernetes, Terraform, OpenStack, SQL, NoSQL, Azure, GCP, DNS, Elastic, Network, Virtualization. DevOps Interview Questions

PythonPyPIOtherdevopsaws
83.3k19.8k
awesome-foss/awesome-sysadmin

A curated list of amazingly awesome open-source sysadmin resources.

Otherawesomeawesome-list
sysadmin.awesome-selfhosted.net
34.7k2.1k
milanm/DevOps-Roadmap

DevOps Roadmap for 2026. with learning resources

Apache License 2.0awsazure
newsletter.techworld-with-milan.com
19.9k3.4k
dastergon/awesome-sre

A curated list of Site Reliability and Production Engineering resources.

Creative Commons Zero v1.0 Universalsite-reliability-engineeringproduction
sre.xyz
13.4k1.8k
kubeshark/kubeshark

eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.

GoGo ModulesApache License 2.0kubernetesgolang
kubeshark.com
12k542
upgundecha/howtheysre

A curated collection of publicly available resources on how technology and tech-savvy organizations around the world practice Site Reliability Engineering (SRE)

JavaScriptnpmCreative Commons Zero v1.0 Universalsite-reliability-engineeringsre
9.8k887
runatlantis/atlantis

Terraform Pull Request Automation

GoGo ModulesApache License 2.0terraformdevops
runatlantis.io
9.2k1.3k
mxssl/sre-interview-prep-guide

Site Reliability Engineer Interview Preparation Guide

studypreparation
9k2.3k
Tracer-Cloud/opensre

Build your own AI SRE agents. The open source toolkit for the AI era.

PythonPyPIApache License 2.0ai-srealerting
discord.com/invite/opensre
8.9k1.2k
isno/theByteBook

⭐ 【出版书籍】京东购买链接 https://item.jd.com/14531549.html 深入讲解内核网络、Kubernetes、ServiceMesh、容器等云原生相关技术。经历实践检验的“大规模分布式系统”开发指南。

JavaScriptnpmcontainerdistributed-systems
thebyte.com.cn
8.5k624
linkedin/school-of-sre

At LinkedIn, we are using this curriculum for onboarding our entry-level talents into the SRE role.

HTMLOthersrelinux
linkedin.github.io/school-of-sre/
8.1k738
k8sgpt-ai/k8sgpt

Giving Kubernetes Superpowers to everyone

GoGo ModulesApache License 2.0devopskubernetes
k8sgpt.ai
8k1k