Structure your Model
The structure of a model is the part you should design yourself, even when an assistant writes all the code. It decides where every calculation lives, how data flows and how easy it is to check a change. Once you have chosen a structure, write it down in your instruction file so the assistant follows it in every task. Treat the guidelines below as a reference rather than as strict rules: what matters most is a structure that is understandable and maintainable.
My preferred approach for structuring a model is the Model, View, Controller (MVC) pattern. This software design pattern is commonly used to separate user interfaces, data, and controlling logic. It emphasizes separating the software’s business logic from its presentation. This “separation of concerns” makes it easier to divide the work and to maintain the code. Models and Views should ideally function independently, while Controllers depend on Models or a combination of Models and Views. Note that while many variations of this pattern exist (see here), they all adhere to the same core principles.
The following diagrams illustrate different data flows within the MVC pattern, depending on the model’s structure and purpose.
The main advantage of the MVC pattern is its clear structure, which makes it obvious which module does what. For instance, if you need to find the Gross Margin ratio calculation or understand the data used for it, you know to look in the Model and Controller, respectively.
The Importance of Separation of Concerns for Financial Modelling
The Model, View, Controller (MVC) structure works well for financial models because they often combine different datasets and complex calculations, which requires overarching logic to manage the data flow correctly.
For example, during scenario analysis or time-based simulations, a dedicated module is needed to track the current period, scenario, dataset, etc. This is where the Controller comes in. The actual calculations should be independent of this tracking logic, which is why they are separated into the Model. Similarly, visualization is distinct from both calculation and control logic and is thus separated into the View.
This modular structure simplifies debugging, as each component has a specific purpose, making it easier to pinpoint formula or data issues. When calculation and data components are combined, it becomes harder to find where an issue comes from.
The Data Layer
The Model layer is where data manipulation occurs. This can range from complex calculations to simple data transformations. Ideally, Models should operate on various data inputs and have minimal dependencies.
For example, a Gross Margin calculation function shouldn’t require input data with specific column names. Instead, it should accept generic inputs like two Series, Arrays, Floats, or Integers. This leads to “dumb” calculation functions that contain no application-specific logic.
You can see how simple such a model is here and below:
def get_gross_margin(revenue: pd.Series, cost_of_goods_sold: pd.Series) -> pd.Series:
"""
Calculate the gross margin, a profitability ratio that measures the percentage of
revenue that exceeds the cost of goods sold.
Args:
revenue (float or pd.Series): Total revenue of the company.
cost_of_goods_sold (float or pd.Series): Total cost of goods sold of the company.
Returns:
float | pd.Series: The gross margin percentage value.
"""
return (revenue - cost_of_goods_sold) / revenue
Each function is categorized in a specific module. For example, the Gross Margin calculation belongs in the profitability_model.py module, alongside other profitability ratio functions. Similarly, other ratio categories like liquidity, solvency, efficiency, and valuation reside in their respective model files (liquidity_model.py, solvency_model.py, efficiency_model.py, and valuation_model.py, respectively).
Because these functions are so small, they are also the easiest to verify. You can put the formula from a textbook next to the code and check it line by line. This is the layer to review most carefully when an assistant adds something: the controller and view only move data around, but a wrong formula here ends up in every result.
The Visualization Layer
The View layer is responsible for presenting data. This can be a table, a graph, a dashboard, etc. In some cases, the data structure (like a DataFrame) produced by the Model might suffice as a “View,” making a dedicated View component optional.
Below is an example of what such a view could look like:
def plot_gross_margin(gross_margin: pd.Series) -> pd.Series:
"""
Plot the gross margin, a profitability ratio that measures the percentage of
revenue that exceeds the cost of goods sold.
Args:
gross_margin (pd.Series): Gross Margin of the company.
Returns:
A plot of the gross margin.
"""
gross_margin.plot(
kind="bar",
title="Gross Margin",
ylabel="Gross Margin (%)",
xlabel="Date",
color="green"
)
Similar to the Model layer, the View layer is typically organized into modules. For example, the Gross Margin plotting function belongs in the profitability_view.py module, along with other profitability visualization functions. Likewise, visualizations for other ratio categories reside in corresponding view files (liquidity_view.py, solvency_view.py, efficiency_view.py, and valuation_view.py, respectively).
The Controlling Layer
The Controller layer contains the logic that connects the Model and View. It handles application-specific logic and data preparation. For example, when calculating the Gross Margin, the Controller extracts the ‘Revenue’ and ‘Cost of Goods Sold’ data from the specific dataset and passes it to the Model function.
The Controller should perform no core calculations. Assistants regularly break this rule by computing something “just here” inside a controller method, so it is worth stating explicitly in your instructions. Its only job is to direct the data flow between the Model and View. Controller logic is often encapsulated within classes. For instance, the Gross Margin calculation might be accessed via a method within a Ratios class (see the actual code here):
class Ratios:
"""
The Ratios Module
"""
def __init__(
self,
tickers: str | list[str],
historical: pd.DataFrame,
balance: pd.DataFrame,
income: pd.DataFrame,
cash: pd.DataFrame,
quarterly: bool = False,
rounding: int | None = 4,
):
"""
Initializes the Ratios Controller Class. Remaining docstring omitted.
"""
# The necessary data initialization is done here.
def get_gross_margin(
self,
rounding: int | None = None,
growth: bool = False,
lag: int | list[int] = 1,
) -> pd.DataFrame:
"""
Calculate the gross margin, a profitability ratio that measures the percentage of
revenue that exceeds the cost of goods sold. Remaining docstring omitted.
"""
gross_margin = profitability_model.get_gross_margin(
self._income_statement.loc[:, "Revenue", :],
self._income_statement.loc[:, "Cost of Goods Sold", :],
)
if growth:
return calculate_growth(
gross_margin, lag=lag, rounding=rounding if rounding else self._rounding
)
return gross_margin.round(rounding if rounding else self._rounding)
Unlike the Model and View layers, a single Controller module might handle multiple related functionalities. This is because the Controller acts as the central coordinator or “glue.” So in this case, this function would fit in the ratios_controller.py module.
It’s also possible, and sometimes useful, to have multiple Controllers. For example, the Finance Toolkit uses a main toolkit_controller.py for initialization.
This main controller then initializes a dedicated ratios_controller.py responsible for ratio calculations. This separation prevents the main Toolkit Controller from becoming overloaded and allows the Ratios Controller to be potentially used independently. For example, in toolkit_controller.py, the Ratios Controller is initialized like this (see actual code here):
class Toolkit:
"""
The Finance Toolkit
"""
def __init__(
self,
tickers: list | str,
api_key: str = "",
start_date: str | None = None,
end_date: str | None = None,
quarterly: bool = False,
rounding: int | None = 4,
):
"""
Initializes an Toolkit object. Remaining docstring omitted.
"""
# The necessary data initialization is done.
def ratios(self) -> Ratios:
"""
The Ratios Module. Remaining docstring omitted.
"""
# The necessary data collection is done here as depicted in the graph above.
return Ratios(
tickers=tickers,
historical=self._quarterly_historical_data
if self._quarterly
else self._yearly_historical_data,
balance=self._balance_sheet_statement,
income=self._income_statement,
cash=self._cash_flow_statement,
custom_ratios_dict=self._custom_ratios,
quarterly=self._quarterly,
rounding=self._rounding,
)
The Supportive Layer
In addition to the Model, View, and Controller modules, a helpers module is useful. This module houses utility functions used across different parts of the application.
For example, consider a function that calculates the growth of a pd.Series or pd.DataFrame. Since growth calculation might be needed for ratios, technical indicators, performance metrics, etc., placing it in a central helpers module avoids code duplication. See the actual code here and an example below:
def calculate_growth(
dataset: pd.Series | pd.DataFrame,
lag: int | list[int] = 1,
rounding: int | None = 4,
axis: str = "columns",
) -> pd.Series | pd.DataFrame:
"""
Calculates growth for a given dataset. Defaults to a lag of 1
(i.e. 1 year or 1 quarter).
Args:
dataset (pd.Series | pd.DataFrame): the dataset to calculate the growth values for.
lag (int | str): the lag to use for the calculation. Defaults to 1.
rounding (int | None): the number of decimals to round the results to.
Defaults to 4.
axis (str): the axis to use for the calculation. Defaults to "columns".
Returns:
pd.Series | pd.DataFrame: _description_
"""
return dataset.pct_change(periods=lag, axis=axis).round(rounding)
Other examples of helper functions include reading data from files (e.g., XLSX, CSV) or handling common errors. Usually, a single helpers module suffices.
Structure and AI Assistants
A clear structure is what makes AI-assisted development manageable. Assistants produce the best results when a task is small and well-defined, such as “add the operating margin to profitability_model.py with a matching controller method and test”. They tend to produce tangled code when the task is vague and the codebase has no clear place for things.
The separation of concerns also determines how easy it is to review what an assistant changed. A change that only touches one model function and its test can be checked in a minute. A change that mixes data handling, calculations and plotting in one function cannot. When you review generated code, ask yourself:
- Is each piece in the right layer? Calculations belong in the Model, data selection in the Controller, presentation in the View. A calculation inside a Controller method is a sign to move it.
- Does the Model function stay generic? It should take series or numbers, not a specific DataFrame with specific column names.
- Did it reuse what exists? Assistants sometimes write a new growth calculation instead of calling the one in
helpers. Duplicated logic drifts apart over time.
Combining Everything
As discussed in the Setting up your Project page, the Model, View, and Controller components for Gross Margin calculations would typically reside in profitability_model.py, profitability_view.py, and profitability_controller.py, respectively. The helpers.py module is usually placed at the package’s root level.
The Finance Toolkit applies this approach and uses the structure shown in the final diagram:
Following this structure, here’s how you might execute the code:
from financetoolkit import Toolkit
companies = Toolkit(['AMZN', 'ASML', 'META'], api_key="FINANCIAL_MODELING_PREP_KEY")
companies.ratios.get_gross_margin()
Which returns the following dataset based on the actual financial statements:
| 2013 | 2014 | 2015 | 2016 | 2017 | 2018 | 2019 | 2020 | 2021 | 2022 | |
|---|---|---|---|---|---|---|---|---|---|---|
| AMZN | 0.2723 | 0.2948 | 0.3304 | 0.3509 | 0.3707 | 0.4025 | 0.4099 | 0.3957 | 0.4203 | 0.4381 |
| ASML | 0.3977 | 0.4264 | 0.4606 | 0.4481 | 0.4503 | 0.4311 | 0.4146 | 0.4863 | 0.5271 | 0.4965 |
| META | 0.7618 | 0.8273 | 0.8401 | 0.8629 | 0.8658 | 0.8325 | 0.8194 | 0.8058 | 0.8079 | 0.7835 |
Alternatively, the growth of the Gross Margin can be calculated as follows:
companies.ratios.get_gross_margin(growth=True)
Which returns the growth of the Gross Margin based on the same financial statements:
| 2013 | 2014 | 2015 | 2016 | 2017 | 2018 | 2019 | 2020 | 2021 | 2022 | |
|---|---|---|---|---|---|---|---|---|---|---|
| AMZN | 0.0828 | 0.1207 | 0.0621 | 0.0563 | 0.0858 | 0.0185 | -0.0347 | 0.0623 | 0.0422 | |
| ASML | 0.0723 | 0.08 | -0.0271 | 0.005 | -0.0426 | -0.0384 | 0.173 | 0.0839 | -0.058 | |
| META | 0.0859 | 0.0155 | 0.0272 | 0.0034 | -0.0386 | -0.0157 | -0.0165 | 0.0026 | -0.0303 |
With this structure in place, you are ready to start building your model. Visit Build your Model to continue!