Experiments Data Model
This page describes the data model for experiment-related objects in Langfuse. For an overview of how these objects work together, see the Concepts page. For score and score config objects, see the Scores data model.
For detailed reference please refer to
- the Python SDK reference
- the JS/TS SDK reference
- the API reference
How experiments are created
You create experiment runs in Langfuse through one of these paths:
| Path | Use when |
|---|---|
| Experiments via SDK | Python or JS/TS Experiment runner |
| Experiments via UI | Prompt or model experiments from the dataset page |
| Experiments via OpenTelemetry | Direct OTEL ingestion: other languages, custom OTLP pipelines, or re-ingesting experiment traces |
To read experiment runs, items, and scores after they exist, use the Experiments API. There is no public REST endpoint for creating new experiment runs; the legacy POST /api/public/dataset-run-items path is deprecated.
Objects
Datasets
Datasets are a collection of inputs and, optionally, expected outputs that can be used during Dataset runs.
Datasets are a collection of DatasetItems.
Dataset object
Prop
Type
DatasetItem object
Prop
Type
DatasetItemMediaReference object
Dataset item media references point from a stored media token in input, expectedOutput, or metadata to a signed media download URL.
Prop
Type
The nested media object contains mediaId, contentType, contentLength, url, and urlExpiry. The url is a signed download URL and should be used before its expiration date. To refresh the signed URL, refetch the dataset.
DatasetRun (Experiment Run)
Dataset runs are used to run a dataset through your LLM application and optionally apply evaluation methods to the results. This is often referred to as Experiment run.
DatasetRun object
Prop
Type
DatasetRunItem object
Prop
Type
Langfuse currently assumes that experiments do not contain repetitions: each dataset item appears once per experiment. Accordingly, reads surface at most one experiment item per dataset item within an experiment. Repetition support is tracked in #5855.
Most of the time, we recommend that DatasetRunItems reference TraceIDs directly. The reference to ObservationID exists for backwards compatibility with older SDK versions.
End to End Data Relations
An experiment can combine a few Langfuse objects:
DatasetRuns(or Experiment runs) are created by looping through all or selectedDatasetItems of aDatasetwith your LLM application.- For each
DatasetItempassed into the LLM application as an Input aDatasetRunItem& aTraceare created. - Optionally
Scores can be added to theTraces to evaluate the output of the LLM application during theDatasetRun.
See the Concepts page for more information on how these objects work together conceptually. See the observability core concepts page for more details on traces and observations. See the Scores data model for more details on score and score config objects.
Function Definitions
When running experiments via the SDK, you define task and evaluator functions. These are user-defined functions that the experiment runner calls for each dataset item. For more information on how experiments work conceptually, see the Concepts page.
Task
A task is a function that takes a dataset item and returns an output during an experiment run.
See SDK references for function signatures and parameters:
Evaluator
An evaluator is a function that scores the output of a task for a single dataset item. Evaluators receive the input, output, expected output, and metadata, and return an Evaluation object that becomes a Score in Langfuse.
See SDK references for function signatures and parameters:
Run Evaluator
A run evaluator is a function that assesses the full experiment results and computes aggregate metrics. When run on Langfuse datasets, the resulting scores are attached to the dataset run.
See SDK references for function signatures and parameters:
For detailed usage examples of tasks and evaluators, see Experiments via SDK. For ingesting experiment traces without the SDK, see Experiments via OpenTelemetry.
Local Datasets
With Langfuse v4 and the current SDKs, experiments on local data appear under Experiments without a hosted dataset. Each task execution also creates a trace for debugging. See Compare experiments.
Last updated on