Out-of-Sample Validation
Out-of-sample validate an ARIMA or VAR forecast of every ticker in the Toolkit instance – fit only on a training portion of the history, forecast the held-out remainder, and score the forecast against what actually happened.
Also known as: hold-out validation, train/test split validation.
See forecast_evaluation_model.get_out_of_sample_validation for the general harness this wraps. Since that function takes a raw Python callable (not serializable for e.g. the MCP-facing tool layer), this controller method instead hardcodes the choice between the two Part-1 forecasting models via the model string:
model="arima":time_series_model.get_arima_forecastis fit on each ticker’s own training-period series (p,d,q,include_constantcontrol the model, same asget_arima_forecast).model="var":time_series_model.get_var_forecastis fit on the training-period series of each ticker together withother_tickers(lagscontrols the VAR order, defaulting to every other ticker in the Toolkit instance if not given); only that ticker’s own forecast column is scored against its holdout.
No programming experience? With the Finance Toolkit MCP server, AI assistants such as Claude and ChatGPT can calculate the Out-of-Sample Validation for you. Just ask in plain English.
Calculate the Out-of-Sample Validation in Python
The Out-of-Sample Validation is available in the Econometrics module of the open-source Finance Toolkit. Install it with:
pip install financetoolkit -U
Then call get_out_of_sample_validation as shown below.
from financetoolkit import Toolkit
toolkit = Toolkit(["AAPL", "MSFT"], api_key="FINANCIAL_MODELING_PREP_KEY")
toolkit.econometrics.get_out_of_sample_validation(
period="weekly", model="arima", p=1, d=1, q=1
)
Which returns:
| AAPL | MSFT | |
|---|---|---|
| RMSE | 12.9091 | 24.2258 |
| MAE | 10.2476 | 20.6824 |
| Holdout Observations | 32 | 32 |
Parameters
get_out_of_sample_validation accepts the following parameters:
- period (str, optional): The data frequency (daily, weekly, monthly, quarterly, or yearly). Defaults to “daily”.
- column (str, optional): The historical data column to validate. Defaults to “Adj Close”.
- model (str, optional): Either “arima” or “var”. Defaults to “arima”.
- train_fraction (float, optional): The fraction of observations used for training; the remainder is the holdout. Defaults to 0.8.
- p (int, optional): The ARIMA autoregressive order (
model="arima"only). Defaults to 1. - d (int, optional): The ARIMA differencing order (
model="arima"only). Defaults to 1. - q (int, optional): The ARIMA moving-average order (
model="arima"only). Defaults to 1. - include_constant (bool, optional): Whether the ARIMA model estimates a
free intercept (
model="arima"only). Defaults to True. - lags (int, optional): The VAR order (
model="var"only). Defaults to 1. - other_tickers (list[str] | None, optional): The other assets to include in
the VAR system alongside the ticker being validated (
model="var"only). Defaults to None, meaning every other ticker in the Toolkit instance. - include_benchmark (bool, optional): Whether to include “Benchmark” among
the tickers validated (and, for
model="var", among the defaultother_tickers). Defaults to False. - rounding (int | None, optional): The number of decimals to round the results to. Defaults to None.
Related Forecast Evaluation
The Econometrics module page introduces the module, and the sidebar lists all of its functions.