AgenticBook

Weights & Biases

wandb.ai

Developer-first MLOps platform for experiment tracking, dataset versioning, and model management. Free tier for personal projects with 100 GB storage.

MLOpsTrackingVersioning

More in Observability, Evaluation & Testing

Weights & Biases WeaveTracing & Observability

W&B's toolkit for tracking LLM application calls, inputs, and outputs alongside traditional ML experiments. Integrates deeply with the W&B ecosystem for comparing model versions and prompt iterations. Ideal for teams already using W&B for ML experiments.

W&BPythonML Ops
Arize PhoenixEvaluation & Benchmarking

Open-source AI observability and evaluation library for LLM and computer vision. Provides a local UI for inspecting traces, running evals, and exploring embedding clusters. Works offline and ships with a comprehensive set of LLM evaluators.

PythonOpen SourceLocal
TraceloopTracing & Observability

OpenTelemetry-native LLM observability platform that auto-instruments popular frameworks with a single line of code. Provides tracing without vendor lock-in, exporting to any OpenTelemetry-compatible backend. Free open-source SDK with optional cloud dashboard.

OpenTelemetryOpen SourceAuto-Instrument
OpenLITTracing & Observability

Open-source LLM observability and evaluation platform with an OpenTelemetry-native SDK. Auto-instruments 20+ LLM providers and vector databases. Ships with a self-hosted UI and Grafana dashboard templates for production monitoring.

Open SourceOpenTelemetryGrafana
SigNozTracing & Observability

Open-source alternative to Datadog with native OpenTelemetry support. Provides distributed tracing, metrics, and logs in a single platform. LLM-specific dashboards make it easy to monitor inference latency and token consumption in production.

GoOpen SourceSelf-hosted
HumanloopEvaluation & Benchmarking

Prompt management and evaluation platform for building reliable AI features. Provides a prompt playground, version control, and A/B testing for prompts. Evaluation suite includes human feedback collection and automated LLM-as-judge scoring.

CloudPromptsA/B Testing