TruLens
trulens.org
Evaluation and tracking library for LLM applications using feedback functions as metrics. Measure groundedness, answer relevance, context relevance, and custom metrics. Integrates with LangChain and LlamaIndex. Dashboard UI for visualizing evaluation trends over time.
More in Observability, Evaluation & Testing
W&B's toolkit for tracking LLM application calls, inputs, and outputs alongside traditional ML experiments. Integrates deeply with the W&B ecosystem for comparing model versions and prompt iterations. Ideal for teams already using W&B for ML experiments.
Open-source AI observability and evaluation library for LLM and computer vision. Provides a local UI for inspecting traces, running evals, and exploring embedding clusters. Works offline and ships with a comprehensive set of LLM evaluators.
OpenTelemetry-native LLM observability platform that auto-instruments popular frameworks with a single line of code. Provides tracing without vendor lock-in, exporting to any OpenTelemetry-compatible backend. Free open-source SDK with optional cloud dashboard.
Open-source LLM observability and evaluation platform with an OpenTelemetry-native SDK. Auto-instruments 20+ LLM providers and vector databases. Ships with a self-hosted UI and Grafana dashboard templates for production monitoring.
Open-source alternative to Datadog with native OpenTelemetry support. Provides distributed tracing, metrics, and logs in a single platform. LLM-specific dashboards make it easy to monitor inference latency and token consumption in production.
Prompt management and evaluation platform for building reliable AI features. Provides a prompt playground, version control, and A/B testing for prompts. Evaluation suite includes human feedback collection and automated LLM-as-judge scoring.