Analysis¶
The openpytea.analysis module provides tools for understanding cost
structure and how uncertain inputs affect financial outcomes:
Cost breakdowns — prepare equipment-level and plant-level CAPEX/OPEX data
Levelized cost breakdown — split the LCOP into discounted CAPEX, OPEX, and side revenue
Cash flow diagram — track a project’s cumulative cash position over time
One-way sensitivity — vary one parameter across a range and observe the metric
Tornado diagram — rank all parameters by their ±impact on a single metric
Monte Carlo simulation — propagate all uncertainties simultaneously
All analysis functions accept a configured and calculated
Plant object. Visualization of the results is
handled separately by Plotting.
To see the outputs of all code examples below, refer to the walkthrough notebook.
from openpytea.analysis import (
direct_costs_data, fixed_capital_data,
fixed_opex_data, variable_opex_data, levelized_cost_data,
cash_flow_data, sensitivity_data, tornado_data, monte_carlo,
)
CAPEX and OPEX breakdowns¶
Cost breakdowns are produced in two steps: the analysis functions prepare structured data, and the plotting functions render it. This separation lets you reuse the data in custom visualizations or export it directly.
The five data-preparation functions and their outputs:
Function |
Output |
|---|---|
|
Equipment-level purchased and direct costs. |
|
ISBL, OSBL, D&E, contingency (and optional additional CAPEX). |
|
Each fixed OPEX component (absolute or as % of total). |
|
Each variable OPEX item. |
|
Discounted CAPEX, OPEX, and side revenue per unit of main product. |
Basic usage (single plant):
# Equipment-level CAPEX
direct_costs = direct_costs_data(plants=plant)
# Fixed capital breakdown (include additional CAPEX events)
fixed_capital = fixed_capital_data(plants=plant, additional_capex=True)
# Fixed OPEX as percentage of total
fixed_opex = fixed_opex_data(plants=plant, pct=True)
# Variable OPEX by item
variable_opex = variable_opex_data(plants=plant)
Pass the returned data to plot_stacked_bar() to visualize it — see
Plotting for details.
Comparing multiple plants¶
Pass a list of Plant objects to compare two or
more configurations side-by-side:
from copy import deepcopy
plant_b = deepcopy(plant)
plant_b.update_configuration({
"plant_name": "Scenario B",
"variable_opex_inputs": {
"electricity": {"consumption": 0.9e6, "price": 0.05},
},
})
plant_b.calculate_all()
variable_opex = variable_opex_data(plants=[plant, plant_b])
# pass to plot_stacked_bar() for a side-by-side chart
Levelized cost breakdown¶
levelized_cost_data() follows the same data +
plot_stacked_bar() pattern as the CAPEX/OPEX breakdowns above, but
mirrors the discounting logic in
calculate_levelized_cost(): capital cost, cash
cost, side-product revenue, and production are each discounted over the
project lifetime at the plant’s interest rate, then divided by discounted
production to express every component per unit of main product.
from openpytea.analysis import levelized_cost_data
from openpytea.plotting import plot_stacked_bar
lcop = levelized_cost_data(plants=plant)
fig, ax = plot_stacked_bar(lcop)
Side revenue is stored as a negative value (since it is subtracted from
the LCOP numerator), so the three components sum directly to the plant’s
LCOP: CAPEX + OPEX + Side revenue = LCOP. plot_stacked_bar() renders
it as a waterfall-style base below zero rather than stacking it like a
normal cost, so the top of the bar still reads as the true net LCOP — see
Plotting for the rendering details.
As with the other breakdowns, pass a list of plants to compare their LCOP
composition side-by-side, and pct=True to express components as a
percentage of the total instead of absolute values. Only the scalar
(non-Monte Carlo) case is supported — each plant’s project_lifetime and
interest_rate must be a single value, not a sampled array.
Cash flow diagram¶
cash_flow_data() prepares the data behind the
classic project cash flow diagram: cumulative cash position vs. time,
including the dip into debt during construction/start-up, the point of
deepest (“maximum”) investment, the break-even (pay-back) point where the
curve first crosses back above zero, and the eventual climb into profit.
It (re)runs each plant’s
calculate_cash_flow() to ensure the underlying
annual cash flow array is up to date.
from openpytea.analysis import cash_flow_data
from openpytea.plotting import plot_cash_flow
cash_flow = cash_flow_data(plant)
fig, ax = plot_cash_flow(cash_flow)
The returned dict has one entry per plant under "curves", each carrying
the cumulative curve itself ("years", "cumulative") alongside the
derived figures "max_investment", "max_investment_year",
"breakeven_year" (None if the project never recovers), and its alias
"payback_time" — useful for pulling numbers into a report without
re-deriving them from the curve:
curve = cash_flow["curves"][0]
print(f"Max investment: {curve['max_investment']:,.0f} in year {curve['max_investment_year']:.0f}")
print(f"Break-even: year {curve['breakeven_year']:.1f}")
Comparing multiple plants¶
Pass a list of plants to overlay their cumulative cash flow curves, each with its own shaded debt region and break-even line:
cash_flow_multi = cash_flow_data([plant, plant_b])
fig, ax = plot_cash_flow(cash_flow_multi, figsize=(4.5, 3))
Only the scalar (non-Monte Carlo) case is supported; if a plant’s
cash_flow has multiple rows (vectorised inputs), the first row is used.
One-way sensitivity analysis¶
sensitivity_data() varies a single parameter over
a symmetric range while holding everything else constant, then records the
selected metric at each point.
# Default metric is LCOP; vary electricity price ±50 %
sens = sensitivity_data(plants=plant, parameter="electricity", plus_minus_value=0.5)
# Specify metric and label explicitly
npv_sens = sensitivity_data(
plants=plant,
parameter="methanol", # product price
metric="NPV",
plus_minus_value=0.5,
label="Project A — NPV [USD]",
)
parameter can be any of:
A key from
variable_opex_inputs— varies that item’s priceA key from
plant_products— varies that product’s price"fixed_capital"— scales total installed CAPEX"fixed_opex"— scales total fixed OPEX"interest_rate"— discount rate"project_lifetime"— project duration"operator_hourly_rate"— labor wage
Supported metric values:
Value |
Description |
|---|---|
|
Levelized cost of the primary product (default). |
|
Net Present Value. |
|
Internal Rate of Return. |
|
Return on Investment. |
|
Simple payback time in years. |
For metrics that depend on revenue (NPV, ROI, IRR, PBT), product prices are included in the evaluation automatically.
Comparing multiple plants:
pbt_comparison = sensitivity_data(
plants=[plant, plant_b],
parameter="electricity",
metric="PBT",
plus_minus_value=0.5,
additional_capex=True, # account for mid-project CAPEX events
n_points=50,
)
Pass the result to plot_sensitivity() — see Plotting.
Tornado diagram¶
A tornado diagram evaluates every variable-cost driver and financial parameter independently at ±``plus_minus_value``, then ranks them by impact on the chosen metric.
from openpytea.analysis import tornado_data
# Default metric is LCOP
td = tornado_data(plant=plant, plus_minus_value=0.5)
# Profit-oriented metric — product prices are included automatically
td_roi = tornado_data(plant=plant, plus_minus_value=0.5, metric="ROI")
Pass td to plot_tornado() — see Plotting.
Monte Carlo simulation¶
Monte Carlo assigns probability distributions to all uncertain inputs and evaluates the plant thousands or millions of times, producing a distribution of outcomes for each financial metric.
Configuring input uncertainties¶
Variable OPEX and product price uncertainties are defined inline in the
existing variable_opex_inputs and plant_products configuration keys
by adding std, min, and max fields to each item:
plant.update_configuration({
"plant_products": {
"methanol": {
"production": 150_000,
"price": 1.75,
"std": 0.25, # standard deviation
"min": 1.25, # lower truncation bound
"max": 2.25, # upper truncation bound
},
},
"operator_hourly_rate": {
"rate": 38.11,
"std": 10.0,
"min": 20.0,
"max": 60.0,
},
"variable_opex_inputs": {
"electricity": {
"consumption": 1.4e6,
"price": 0.10,
"std": 0.035,
"min": 0.025,
"max": 0.175,
},
"natural_gas": {
"consumption": 1.0e5,
"price": 0.05,
"std": 0.03,
"min": 0.001,
"max": 0.10,
},
},
})
Project-level financial uncertainties are set through the
project_uncertainties key:
plant.update_configuration({
"project_uncertainties": {
"fixed_capital_factor": {"std": 0.30, "min": 0.25, "max": 1.75},
"fixed_opex_factor": {"std": 0.30, "min": 0.25, "max": 1.75},
"project_lifetime": {"std": 5}, # min/max auto-derived
"interest_rate": {"std": 0.03}, # min/max auto-derived
"plant_utilization": {"std": 0.05}, # opt-in; default std=0
"tax_rate": {"std": 0.10}, # opt-in; default std=0
}
})
The first four keys are active by default. plant_utilization and
tax_rate require an explicit std > 0 (or an explicit dist_id) to
be sampled. For project_lifetime, interest_rate, plant_utilization,
and tax_rate, omitted min/max are derived as ±2 × std around the
plant’s baseline value. For fixed_capital_factor and fixed_opex_factor
the default bounds are a fixed [0.25, 1.75] regardless of std unless
you set min/max explicitly. Set std=0 for any key to disable
sampling for it (the value collapses to its baseline).
Key |
Description |
Default std |
|---|---|---|
|
Multiplicative factor on total installed CAPEX. |
0.30 (30%) |
|
Multiplicative factor on annual fixed OPEX. |
0.30 (30%) |
|
Economic project life (years). |
5 years |
|
Discount / financing rate. |
0.03 (3 pp) |
|
Yearly fraction of operating time. |
0 (opt-in) |
|
Corporate tax rate. |
0 (opt-in) |
Every uncertain input above defaults to a Normal distribution built from
its std (and, if given, min/max truncation bounds). Add a
dist_id field to any uncertainty block — in variable_opex_inputs,
plant_products, operator_hourly_rate, or project_uncertainties —
to draw from a different family instead. Field names are reused across
families (loc/mean/price/rate, scale/std, shape,
minimum/min, maximum/max); which ones apply depends on
dist_id.
Under the hood every family is a frozen scipy.stats distribution, so the parameter meanings and shapes follow SciPy’s conventions. The table below links each family to its SciPy reference page for the full mathematical definition:
|
Family |
Parameters used |
Notes |
|---|---|---|---|
0 / 1 |
Fixed value |
|
No randomness; every draw equals |
2 |
|
Optional |
|
3 |
Normal (default) |
|
Optional |
4 |
|
||
5 |
|
||
6 |
|
Draws are 0 or |
|
7 |
|
Integers, inclusive of |
|
8 |
|
||
9 |
|
||
10 |
|
||
11 |
|
||
12 |
|
plant.update_configuration({
"plant_products": {
"methanol": {
"production": 150_000,
"price": 1.75,
"dist_id": 5, # Triangular; "price" doubles as the mode
"min": 1.25,
"max": 2.50,
},
},
"project_uncertainties": {
"fixed_capital_factor": {
"dist_id": 2, # Lognormal
"loc": 0.0, # mu
"std": 0.20, # sigma
},
},
})
Only Lognormal, Normal, and Bernoulli (dist_id 2, 3, 6) apply min/
max as post-hoc truncation via rejection sampling. For Uniform,
Triangular, and Discrete uniform the bounds define the distribution itself.
Weibull, Gamma, GEV, and Student’s t ignore min/max entirely (Beta
uses max as its upper scale bound instead).
For direct programmatic use outside of monte_carlo, the same families
are available via make_distribution() (returns a
frozen scipy.stats distribution) and
sample_distribution() (draws an array of samples,
with optional truncation).
Running the simulation¶
from openpytea.analysis import monte_carlo
mc_results = monte_carlo(
plant,
num_samples=1_000_000, # increase for accuracy, decrease for speed
batch_size=1_000, # samples evaluated per batch; default 1000
random_seed=42, # optional, for reproducible runs
)
monte_carlo returns a dict:
Key |
Description |
|---|---|
|
Dict of sample arrays keyed by metric: |
|
Dict mapping each sampled input’s display name to its sample array. |
|
The plant’s name. |
|
Echo of the parameters the run was executed with. |
LCOP is always computed. NPV, ROI, and PBT are only
meaningful (otherwise they stay zero-filled) when every entry in
plant_products has a "price" set. There is no "IRR" key here —
IRR is available for sensitivity_data() and
tornado_data(), but not for monte_carlo. Pass
additional_capex=True to account for mid-project CAPEX events in the
ROI/PBT calculation.
The same "metrics" and "inputs" dicts are also stored on the plant
as plant.monte_carlo_metrics and plant.monte_carlo_inputs, which is
what the plotting functions fall back to when passed a Plant directly.
# Access results
print(mc_results["metrics"]["LCOP"]) # array of LCOP samples
print(mc_results["metrics"]["NPV"]) # array of NPV samples
Visualizing results¶
Pass the plant (or mc_results) to the plotting functions:
from openpytea.plotting import plot_monte_carlo, plot_monte_carlo_inputs
# Distribution of the LCOP
fig, ax = plot_monte_carlo(plant, metric="LCOP", bins=30)
# Verify input distributions (useful for checking std/min/max settings)
fig, axes = plot_monte_carlo_inputs(mc_results, bins=40)
See Plotting for full plotting options.
Comparing multiple plants under uncertainty¶
from openpytea.plotting import plot_multiple_monte_carlo
mc_b = monte_carlo(plant_b, num_samples=1_000_000, batch_size=10_000)
fig, ax = plot_multiple_monte_carlo(
data_list=[plant, plant_b],
metric="LCOP",
bins=30,
)
See also¶
openpytea.analysis— full API referencePlotting — visualization options
Walkthrough notebook — end-to-end worked example