AgenticBook

OpenLIT

github.com

Open-source LLM observability and evaluation platform with an OpenTelemetry-native SDK. Auto-instruments 20+ LLM providers and vector databases. Ships with a self-hosted UI and Grafana dashboard templates for production monitoring.

Open SourceOpenTelemetryGrafana

More in Observability, Evaluation & Testing

SigNozTracing & Observability

Open-source alternative to Datadog with native OpenTelemetry support. Provides distributed tracing, metrics, and logs in a single platform. LLM-specific dashboards make it easy to monitor inference latency and token consumption in production.

GoOpen SourceSelf-hosted
HumanloopEvaluation & Benchmarking

Prompt management and evaluation platform for building reliable AI features. Provides a prompt playground, version control, and A/B testing for prompts. Evaluation suite includes human feedback collection and automated LLM-as-judge scoring.

CloudPromptsA/B Testing
TraceloopTracing & Observability

OpenTelemetry-native LLM observability platform that auto-instruments popular frameworks with a single line of code. Provides tracing without vendor lock-in, exporting to any OpenTelemetry-compatible backend. Free open-source SDK with optional cloud dashboard.

OpenTelemetryOpen SourceAuto-Instrument
Weights & Biases WeaveTracing & Observability

W&B's toolkit for tracking LLM application calls, inputs, and outputs alongside traditional ML experiments. Integrates deeply with the W&B ecosystem for comparing model versions and prompt iterations. Ideal for teams already using W&B for ML experiments.

W&BPythonML Ops
Arize PhoenixEvaluation & Benchmarking

Open-source AI observability and evaluation library for LLM and computer vision. Provides a local UI for inspecting traces, running evals, and exploring embedding clusters. Works offline and ships with a comprehensive set of LLM evaluators.

PythonOpen SourceLocal
Confident AIEvaluation & Benchmarking

Cloud platform for running, tracking, and regression-testing LLM evaluation pipelines. Works with DeepEval for running evals and provides a dashboard for tracking quality over time. Helps teams catch prompt regressions before they reach production.

CloudRegressionCI/CD