Scores Data Model
This page describes the data model for score-related objects in Langfuse. For an overview of what scores are and when to use them, see the Scores overview. For datasets, experiment runs, and function definitions, see the Experiments data model.
For detailed reference please refer to
- the Python SDK reference
- the JS/TS SDK reference
- the API reference
Scores
Scores are the data object to store evaluation results. They are used to assign evaluation scores to traces, observations, sessions, or dataset runs. Scores can be added manually via annotations, programmatically via the SDK/API, or automatically via LLM-as-a-Judge evaluators.
Scores have the following properties:
- Each Score references exactly one of
Trace,Observation,Session, orDatasetRun - Scores are either numeric, categorical, boolean, or text (see Score Types)
- Scores can optionally be linked to a
ScoreConfigto ensure they comply with a specific schema
Score object
Prop
Type
Common Use Cases
| Level | Description |
|---|---|
| Trace | Used for evaluation of a single interaction. (most common) |
| Observation | Used for evaluation of a single observation below the trace level. |
| Session | Used for comprehensive evaluation of outputs across multiple interactions. |
| Dataset Run | Used for performance scores of a Dataset Run. |
Score Config
Score configs are used to ensure that your scores follow a specific schema. Using score configs allows you to standardize your scoring schema across your team and ensure that scores are consistent and comparable for future analysis.
You can define a ScoreConfig in the Langfuse UI or via our API. Configs are immutable but can be archived (and restored anytime).
A score config includes:
- Score name
- Data type:
NUMERIC,CATEGORICAL,BOOLEAN,TEXT - Constraints on score value range (Min/Max for numerical, Custom categories for categorical data types, 1-500 characters for text)
ScoreConfig object
Prop
Type
Last updated on