AgenticBook

SWE-bench

swe-bench.github.io

The definitive benchmark for evaluating AI coding agents on real GitHub issues. Contains 2,294 tasks from 12 popular Python repositories. Measures ability to generate patches that pass existing test suites. Every major coding agent reports SWE-bench scores.

BenchmarkSWEPython

More in Observability, Evaluation & Testing

Weights & Biases WeaveTracing & Observability

W&B's toolkit for tracking LLM application calls, inputs, and outputs alongside traditional ML experiments. Integrates deeply with the W&B ecosystem for comparing model versions and prompt iterations. Ideal for teams already using W&B for ML experiments.

W&BPythonML Ops
Arize PhoenixEvaluation & Benchmarking

Open-source AI observability and evaluation library for LLM and computer vision. Provides a local UI for inspecting traces, running evals, and exploring embedding clusters. Works offline and ships with a comprehensive set of LLM evaluators.

PythonOpen SourceLocal
TraceloopTracing & Observability

OpenTelemetry-native LLM observability platform that auto-instruments popular frameworks with a single line of code. Provides tracing without vendor lock-in, exporting to any OpenTelemetry-compatible backend. Free open-source SDK with optional cloud dashboard.

OpenTelemetryOpen SourceAuto-Instrument
OpenLITTracing & Observability

Open-source LLM observability and evaluation platform with an OpenTelemetry-native SDK. Auto-instruments 20+ LLM providers and vector databases. Ships with a self-hosted UI and Grafana dashboard templates for production monitoring.

Open SourceOpenTelemetryGrafana
SigNozTracing & Observability

Open-source alternative to Datadog with native OpenTelemetry support. Provides distributed tracing, metrics, and logs in a single platform. LLM-specific dashboards make it easy to monitor inference latency and token consumption in production.

GoOpen SourceSelf-hosted
HumanloopEvaluation & Benchmarking

Prompt management and evaluation platform for building reliable AI features. Provides a prompt playground, version control, and A/B testing for prompts. Evaluation suite includes human feedback collection and automated LLM-as-judge scoring.

CloudPromptsA/B Testing