๐Ÿค–AI RepoIndex
๐Ÿ”ง AI ProductivityOpen SourceFree

Phoenix (Arize)

AI observability โ€” tracing, evaluation, experimentation

4.4/ 5โญ 10,000 GitHub starsPythonApache-2.0

๐Ÿ“‹ Overview

Phoenix is an open-source AI observability platform that provides tracing, evaluation, and experimentation for LLM applications. Unlike Langfuse (which focuses on tracing and evals), Phoenix emphasizes experimentation โ€” A/B testing prompts, comparing model versions, and analyzing output quality at scale. It integrates with OpenTelemetry, LangChain, LlamaIndex, and OpenAI SDK. The platform provides trace exploration, span analysis, and built-in evaluation frameworks for common tasks (hallucination detection, relevance scoring).

โœจ Key Features

  • โ€ข10K+ GitHub stars โ€” AI observability with experimentation
  • โ€ขTracing: capture prompts, responses, latency, token usage
  • โ€ขEvaluation: LLM-as-judge, heuristic, human evals
  • โ€ขExperimentation: A/B test prompts and models
  • โ€ขOpenTelemetry integration
  • โ€ขLangChain, LlamaIndex, OpenAI SDK support
  • โ€ขTrace exploration and span analysis
  • โ€ขStatistical significance testing
  • โ€ขSpan analysis

๐ŸŽฏ The Problem It Solves

Teams need to compare prompt versions, evaluate model outputs, and debug production LLM behavior. Phoenix provides a unified platform for tracing, evaluation, and experimentation โ€” going beyond observability to active improvement.

๐Ÿ”ง How It Works

Phoenix integrates via OpenTelemetry or SDK โ€” wrap your LLM calls to automatically capture traces. The evaluation framework runs LLM-as-judge, heuristic, and human evaluations against your traces. The experimentation module lets you A/B test prompts and models, with statistical significance testing. The dashboard provides trace exploration, span analysis, and quality metrics.

๐Ÿš€ Installation & Quick Start

Installation

See website

Quick Start

  1. See documentation

โœ… Pros

  • โ€ขStrong experimentation features
  • โ€ขOpenTelemetry integration
  • โ€ขMultiple evaluation methods
  • โ€ขActive development
  • โ€ขApache 2.0 licensed
  • โ€ขStatistical significance testing
  • โ€ขSpan analysis

โŒ Cons

  • โ€ขSmaller community than Langfuse
  • โ€ขNo cloud tier
  • โ€ขEvaluation framework less mature
  • โ€ขDocumentation lags
  • โ€ขLimited deployment options

๐Ÿ’ฌ Practitioner Verdict

โ€œPhoenix is a strong Langfuse alternative โ€” the experimentation features are genuinely differentiated, and the OpenTelemetry integration is excellent. The trade-off: smaller community than Langfuse, no cloud tier, and the evaluation framework is less mature. For teams that want open-source AI observability with experimentation, Phoenix is the default.โ€
1

Self-Hosted (Free)

Open source, MIT/Apache licensed. Run it yourself.

โญ Star & Clone on GitHub

Free forever. Your infrastructure, your data.

3

Deployment Options

Ways to run Phoenix (Arize) in production

๐Ÿ“Š Specifications

Language
Python
License
Apache-2.0
Platform
Linux, macOS, Windows
Supported Models
REST API, CLI

๐Ÿ’ฐ Pricing Reality

100% free Apache 2.0. No paid tier. Self-hosted only.

๐Ÿ‘ฅ Community Health

Stars10,000
Forks1,250
Contributors200
Health Score7/10

๐Ÿท๏ธ Tags

Open SourceFree