Alpha

Finterion is now in active Alpha — early access is open.

QuantOS
Investing-algorithm-framework
.obtf
backtesting
Introducing the Open Backtest Format (.obtf)
Backtesting & Validation

Introducing the Open Backtest Format (.obtf)

Open Backtest Format OBTF Open Backtest Format is a versioned, framework and engineagnostic, studiesfirst ondisk format for storing the evidence produced by al
mduyn

mduyn

•

September 7, 2026

Summarize with AI
Share this article

Open Backtest Format

OBTF (Open Backtest Format) is a versioned, framework- and engine-agnostic, studies-first on-disk format for storing the evidence produced by algorithmic trading backtests in a single, self-contained file.

An OBTF file can contain the complete context and results of an quantitative trading algorithm, including studies, universes, metrics, trades, orders, portfolio snapshots, execution assumptions, and Monte Carlo test results.

The goal of OBTF is to make backtest results portable, reproducible, and comparable across frameworks and engines. By storing results in a standardized format, the performance of a single algorithm can be compared across different studies, datasets, and configurations. It also provides the foundation for comparing different algorithms under the same study conditions.

OBTF is an open effort to standardize how algorithmic trading research and backtest results are stored and exchanged. Rather than tying results to a specific framework or execution engine, the format defines a common structure that different projects can adopt and implement independently.

The first open-source framework to adopt OBTF is the investing-algorithm-framework.

OBTF v1.0 will be released soon, together with a reference specification and adoption guide for other frameworks, engines, and research platforms. The project can be found here

The Problem: Results Without Their Research Context

Many backtesting workflows save results one backtest at a time. That is convenient while running experiments, but awkward when evaluating a strategy as a whole.

Current systems scatter the evidence for the performance of algorithms accross a wide variety of files. For example, a notebook holds the equity curve, a CSV holds the trades, and a separate configuration holds the fees. A proprietary dashboard may bring them together, but only inside the platform that produced them.

That fragmentation creates three practical problems:

  • Portability: sharing or archiving results often requires custom exports and engine-specific parsing. The evidence should be usable beyond its original dashboard.
  • Comparison: two headline returns can hide different universes, time windows, and execution assumptions. Researchers need access to that context before treating results as comparable.
  • Lineage: disconnected reports make it difficult to follow related experiments, identify which strategy variant produced a result, or see how a strategy performs across studies.

Together, these make straightforward questions surprisingly difficult to answer:

  • How consistently does the strategy perform across different time windows?
  • What do its aggregate metrics look like across all runs, rather than in one selected result?
  • How much does performance change across universes, assets, and market conditions?
  • Do the results remain attractive after applying realistic commissions, slippage, and fill assumptions?
  • Are returns broadly distributed, or driven by a small number of trades, assets, or unusually favorable periods?
  • Does the same strategy variant behave consistently across different studies and backtesting engines?
  • Does the evidence show a robust strategy, or merely its best run?

A backtest by any framework answers, "What happened in this run?" OBTF is designed to help preserve the evidence for the bigger question: "What have we learned about this strategy?"

The missing piece is a shared structure that keeps results connected to their research context. The goal here is to have in a single opienenated overview the evidence how a strategy performs on different assets, time-windows, parameters and if it overfits.

research-workflow Illustrative view: one universe per study, with runs across multiple windows.

OBTF does not replace the experiments. It provides a common structure for preserving their outputs together. You end up with a file containing all the evidence how your strategy performs in different scenarios.

From Individual Runs to Studies

The most important design choice is that OBTF is a studies-first structure. A study contains backtest runs, metrics, the execution context of these runs, the universe that was traded and the metadata that the user has provided.

A run is one execution over a particular time window and universe of assets. A study groups related evaluations of a strategy idea, with context such as its validation role and execution assumptions. Within each study, engine slots separate results produced by different engine classes.

For example, a obtf. bundle could contain an exploratory study using a vectorized engine on an in-sample set of backtest windows and a separate out-of-sample validation study using an event-driven engine. Each can contain multiple runs across different periods on its single universe.

study-structure

Notice where the assumptions sit: on the study. A zero-cost exploration and a realistic-cost validation can coexist without implying that they used the same execution model.

Each bundle also carry an algorithm identifier, a set of parameters that where used to initialised the algorith and an anchor_algorithm_id linking a variant to its source algorithm. This supports tracing related experiments across bundles. All these identifiers are producer-defined and therefore are the responsibility of the developer.

bundle-structure

Making Robustness Easier to Inspect

Consider a hypothetical trend-following strategy. Its initial backtest looks promising. Before taking it seriously, you want to examine several additional questions:

Evidence to inspectResearch question
Runs across different periodsDoes performance depend on one favorable market?
Separate studies for different universesDoes the result depend on a particular asset selection?
A study with realistic costsDo commission and slippage erase the apparent edge?
Held-out validation runsDoes the idea survive beyond the data used to develop it?
Monte Carlo test resultsHow unusual is the observed result under the chosen null model?

OBTF gives these outputs places within the same algorithm results. If these studies are appended to the same .obtf file, all the evindence becomes centrally stored and gives developers the opportunity to compare different studies and algorithms in a standardised way.

For example, tools can be developed for comparisons across strategies, where it can align runs from separate bundles by universe, window, and assumptions.

One File, Built From Established Technologies

The file begins with an OBTF identifier and a format-version number. Its body uses MessagePack for structured data and Zstandard for compression. Larger supported time-series fields can be stored as embedded Parquet blobs.

Those pieces travel together in a single file. Versioning provides a compatibility boundary: readers must reject newer versions they do not support.

Because the specification is open, a compatible notebook, dashboard, or ranking tool can read the bundle without being the tool that produced it. However the results such as metrics and runs are tied to one platform's presentation. Therefore its not recommended to compare .obtf files produced by different frameworks.

Portable Does Not Mean Identical

Engine-agnostic storage does not make different engines behave identically. A vectorized simulation and an event-driven simulation can make different execution assumptions; OBTF keeps their results separate within a shared structure.

There are other important boundaries:

  • Metric calculations remain producer-defined. Two stored Sharpe ratios of different frameoworks or engines are not automatically comparable without checking their definitions.
  • Detailed row schemas are not yet universal. Trades, orders, and several other record types remain producer-defined in the current design, so some cross-framework analysis still needs producer-specific knowledge. However, .obtf provides a reference archictecture for these objects.
  • A self-contained results file is not a complete reproduction environment. OBTF does not package strategy source code or raw market data. It can carry data-source references, but reproducing a run requires those external inputs and the relevant software.

The promise is a shared container and research structure, not automatic equivalence between every backtesting system.

Available Today

OBTF is now officially supported in investing-algorithm-framework (QuantOS), which hosts the reference Python implementation.

The specification is currently v0.1.0-draft. It is not yet stable, parts of the specification remain unfinished, and adoption beyond the reference implementation has not yet been established.

The ambition is to make a strategy's evaluation portable as a coherent body of evidence, rather than a collection of engine-specific exports. Broader adoption would let engine authors focus on simulation and analysis-tool authors build against a shared results structure.


Upvote1
Comments0
Views224
0 Comments
About the Author
mduyn
mduyn

Founder and CEO of Finterion. I am passionate about making algorithmic trading accessible to everyone. I regularly post about open-source projects and share my strategies on finterion.

Share this article
Table of Contents
  • Open Backtest Format

  • The Problem: Results Without Their Research Context

  • From Individual Runs to Studies

  • Making Robustness Easier to Inspect

  • One File, Built From Established Technologies

  • Portable Does Not Mean Identical

  • Available Today

About the Author
mduyn
mduyn

Founder and CEO of Finterion. I am passionate about making algorithmic trading accessible to everyone. I regularly post about open-source projects and share my strategies on finterion.


Share this article