Best Open Source observability Libraries
A curated list of the most popular GitHub repositories tagged with observability. Select any project to visualize its architecture and dive into the codebase using Codyn's AI engine.
#1netdata/netdata
The fastest path to AI-powered full stack observability, even for lean teams.
#2SigNoz/signoz
SigNoz is an open-source observability platform native to OpenTelemetry with logs, traces and metrics in a single application. An open-source alternative to DataDog, NewRelic, etc. 🔥 🖥. 👉 Open source Application Performance Monitoring (APM) & Observability tool
#3mlflow/mlflow
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
#4apache/skywalking
APM, Application Performance Monitoring System
#5elastic/kibana
Your window into all of your data
#6mikeroyal/Self-Hosting-Guide
Self-Hosting Guide. Learn all about locally hosting (on premises & private web servers) and managing software applications by yourself or your organization. Including Cloud, LLMs, WireGuard, Automation, Home Assistant, and Networking.
#7openobserve/openobserve
OpenObserve is an open-source observability platform for logs, metrics, traces, and frontend monitoring. A cost-effective alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.
#8openzipkin/zipkin
Zipkin is a distributed tracing system
#9kubesphere/kubesphere
The container platform tailored for Kubernetes multi-cloud, datacenter, and edge management ⎈ 🖥 ☁️
#10VictoriaMetrics/VictoriaMetrics
VictoriaMetrics: fast, cost-effective monitoring solution and time series database
#11Effect-TS/effect
Build production-ready applications in TypeScript
#12Tracer-Cloud/opensre
Build your own AI SRE agents. The open source toolkit for the AI era.
#13upgundecha/howtheysre
A curated collection of publicly available resources on how technology and tech-savvy organizations around the world practice Site Reliability Engineering (SRE)
#14hyperdxio/hyperdx
Resolve production issues, fast. An open source observability platform unifying session replays, logs, metrics, traces and errors powered by ClickHouse and OpenTelemetry.
#15highlight/highlight
highlight.io: The open source, full-stack monitoring platform. Error monitoring, session replay, logging, distributed tracing, and more.
#16apache/hertzbeat
An AI-powered next-generation open source real-time observability system.
#17grafana/mimir
Grafana Mimir provides horizontally scalable, highly available, multi-tenant, long-term storage for Prometheus.
#18micrometer-metrics/micrometer
An application observability facade for the most popular observability tools. Think SLF4J, but for observability.
#19parca-dev/parca
Continuous profiling for analysis of CPU and memory usage, down to the line number and throughout time. Saving infrastructure cost, improving performance, and increasing reliability.
#20rajnandan1/kener
Stunning status pages, batteries included!
#21openlit/openlit
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs.
#22parseablehq/parseable
Parseable is an open source, unified infrastructure observability platform built in Rust on a data lake architecture. It tracks logs, metrics, traces, and events across apps, agents, and systems, reducing storage costs by up to 90% through columnar telemetry compression.
#23seakee/CPA-Manager-Plus
A self-hosted CPA / CLIProxyAPI management panel and AI gateway observability dashboard for requests, usage, cost, quota, failures, and account health.
#24carverauto/serviceradar
Open-Source Network Management, Monitoring, ITOM, and Security Analytics
#25DataDog/dd-trace-py
Datadog Python APM Client
#26ongridio/ongrid
An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.
#27databufflabs/databuff
AI-native OpenTelemetry APM with multi-agent root-cause analysis across traces, metrics, and service topology
#28rajudandigam/agent-inspect
Local execution trees for TypeScript AI agents. agent-inspect helps you understand what happened inside an AI agent run — locally. It turns manual steps, tool calls, LLM calls, structured logs, failures, durations, and run metadata into readable execution trees you can inspect from the terminal. It is built for TypeScript/Node.js developers..
#29Dynatrace/dynatrace-operator
Automate Kubernetes observability with Dynatrace
#30Vinix24/vnx-orchestration
Governance-first orchestration for Claude Code, Codex, and Gemini CLI — parallel workers, receipts, quality gates, and full provenance.