Skip to main content
Evaluations help you measure and improve LLM output quality through datasets, automated scoring, and experiments.

Score an Observation

Add quality scores to any traced observation.

Score a Trace

Apply scores to the entire trace (multiple observations).

Create a Dataset

Create evaluation datasets to systematically test your LLM.

Run Dataset Evaluation

Iterate through dataset items and score results.

LLM-as-Judge Scoring

Use an LLM to evaluate output quality.

Batch Scoring

Efficiently score multiple traces/observations.

A/B Test Prompts

Compare different prompts using dataset evaluation.

Integration Patterns

Next: Combining features

Evaluations Guide

Reference: Full evaluations documentation