LMSYS Chatbot Arena
chat.lmsys.org
Crowdsourced benchmark for LLM evaluation through blind A/B tests. Over 1 million human comparisons powering the Elo-based leaderboard. Identifies which models perform best on instruction following and reasoning — critical data for choosing the right LLM backend for agent applications.
More in Agent Community, Directories & Research
Official collection of reference applications and production-ready LangChain templates. Covers RAG, extraction, chatbots, agents, and more — each deployable with a single command. The fastest way to bootstrap a production LangChain application.
Open-source LLM application development platform that combines BaaS and LLMOps. Provides a visual workflow builder, RAG pipeline, agent framework, and model management in one product. One of the fastest-growing open-source AI application platforms.
OpenAI's official marketplace for custom ChatGPT agents (GPTs). Browse thousands of community-built specialized agents for coding, research, creative writing, and analysis — or publish your own GPT to reach millions of ChatGPT users without building a standalone product.
Community hub for sharing and discovering Superagent-based AI assistants and agent configurations. Browse pre-built agents for sales, support, and research use cases. Fork and deploy community agents directly to the Superagent cloud.
Transformers Agents is a multi-modal agent API built into the Hugging Face ecosystem. Provides a natural language interface for calling 100,000+ HF models as tools. Integrates with the Hub model, dataset, and Space ecosystem for end-to-end AI agent workflows.
Comprehensive benchmark for evaluating LLM-based agents across 8 distinct environments including code, databases, web browsing, and games. Reveals major performance gaps between open and closed models on real-world agentic tasks. The de-facto standard for comparing agent capabilities.