laurel.pipelines.evaluate_impacts package
Submodules
laurel.pipelines.evaluate_impacts.nodes module
Kedro pipeline nodes for the evaluate_impacts pipeline (Model Modules 5 & 6).
This module implements the final two model modules described in Passow & Rajagopal (2026): estimating expected electrified dwell counts per location and assembling per-substation/county charging load profiles across a fleet scaled to a target adoption level.
Pipeline overview
Module 5 — Estimate expected electrified dwells:
summarize_vehicles — compute per-vehicle performance metrics (range deaths, charging delays) and flag vehicles whose simulation results are implausible for inclusion in load profiles.
apply_delays — shift dwell start/end timestamps and update dwell durations to account for charging-induced delays.
filter_dwells_pre_prob — retain only weekday dwells; keep zero-charge and non-electrified-vehicle dwells (needed for probability estimation).
filter_locs_pre_prob — drop locations with missing required fields.
build_class_frame — enumerate the full cross-product of vehicle and location class combinations.
compute_class_dwell_counts — count observed dwells per (vehicle class × location class × electrification status).
compute_adoption_totals — derive electrified-vehicle counts per vehicle class from scenario adoption forecasts; optionally override with scenario parameters.
compute_dwell_rate_vclass — compute dwell rate (dwells/day) per vehicle class from observed data.
compute_class_probs — fuse observed dwell location distributions with adoption forecasts to estimate P(electrified | location class, vehicle class) and expected dwell counts per class via logistic regression + correction term (
ElectProbLocalizer).
Module 6 — Assemble load profiles:
filter_dwells_post_prob — drop dwells from non-electrified vehicles.
add_dwell_id — assign a sequential integer ID to each dwell row for joining with charging events.
get_dwells_nonzero — drop zero-charge dwells (creates a new DwellSet view, not in-place).
manage_charging — convert dwell records to charging event records using the configured charging manager.
slice_events — partition charging events into time-of-day windows and synthesise initialisation events for carry-over charging.
build_time_ordered_slice — sort and difference profile columns within each dwell to produce per-time-step power increments.
sample_profiles_node — core bootstrap sampling: build sparse correspondence matrices, compute inverse propensity weights, draw
n_bootstrapssamples, and return per-region load profiles and summaries.build_eval_columns — inject
group_colsinto the sharedpcolsparameter dictionary.add_region_geoms — dissolve H3 hexagon correspondences into region polygons and join geometries to results.
localize_time_from_hexes — look up time zones from H3 hex centroids and convert UTC timestamps to local time.
calc_utilization — compute charging utilisation (energy / peak × window duration) for each profile type.
compress_bootstrap_profiles — discretise time and quantile-compress bootstrap profile draws into a compact representation.
compress_bootstrap_summaries — quantile-compress and mean-aggregate bootstrap scalar summaries.
Key design decisions
Critical-day / zero-charge retention through Module 5: Zero-charge dwells and dwells from non-electrified vehicles are retained through the probability estimation step because they contribute to observed dwell distributions that inform the logistic regression. They are dropped only after class probabilities have been computed.
Inverse propensity score weighting: Observed dwells are re-weighted by the reciprocal of P(observed & electrified | location, vehicle class) so that locations with low observation propensity are up-weighted when constructing load profiles. This corrects for sampling bias in the telematics data.
Two-stage bootstrap sampling: Each bootstrap draw first allocates expected dwell counts to hexagons (stage 1), then samples charging-event profiles from the pool of observed events at that hex or its freight-activity class (stage 2). The
sample_self/sample_classflags control whether a hex can donate from its own observations, from class-level observations, or both.Dask for bootstrap parallelism: When
n_bootstraps > 1, individual draws are distributed across Dask workers viaclient.submit; large numpy arrays are scattered withbroadcast=Trueto avoid repeated serialisation.Timedelta columns: Several delay columns are stored as floats (hours) in the dwell data.
apply_delaysconverts them totimedeltafor arithmetic and drops the temporary columns afterwards.
References
Passow, F., & Rajagopal, R. (2026). Identifying indicators to inform proactive substation upgrades for charging electric heavy-duty trucks. Applied Energy (submitted March 2026).
- laurel.pipelines.evaluate_impacts.nodes.add_dwell_id(dw, params)[source]
Assign a sequential integer dwell ID to each row.
The ID is used as a join key between the dwell table and the charging events table produced by
manage_charging.- Parameters:
dw (
DwellSet) – Dwell dataset. Modified in-place and returned.params (
dict) –Pipeline parameters. Expected key:
dw_col(str): Name for the new integer ID column.
- Return type:
DwellSet- Returns:
The updated
DwellSetwith a new sequential integer ID column.
- laurel.pipelines.evaluate_impacts.nodes.add_region_geoms(results, hex_regions, params)[source]
Dissolve H3 hex correspondences into region polygons and join to results.
Converts a hex→region correspondence table into one polygon per region (by dissolving the H3 cell polygons), joins it to the results DataFrame, and converts any
timedeltacolumns to float hours for serialisation compatibility.- Parameters:
results (
DataFrame) – Per-region summary or profile DataFrame to annotate with geometries.hex_regions (
DataFrame) – DataFrame mapping H3 hex IDs to region identifiers.params (
dict) –Pipeline parameters. Expected keys:
hex_col(str): Column inhex_regionsholding H3 hex IDs.region_col(str): Column in bothhex_regionsandresultsholding the region identifier used to join.
- Return type:
GeoDataFrame- Returns:
A
GeoDataFramewith ageometrycolumn (dissolved region polygons) andtimedeltacolumns replaced by float-hours equivalents.
- laurel.pipelines.evaluate_impacts.nodes.apply_delays(dw, params)[source]
Apply charging-induced delays to dwell timestamps and durations.
The charging-choice simulation records delays as floating-point hours. This node converts those columns to
timedeltaobjects, then applies them to the dwell table so that downstream time-of-day analysis uses realised (delayed) times rather than originally scheduled times.Three adjustments are made to each dwell:
dw.startis shifted forward by the cumulative delay accumulated up to the start of this dwell.Net dwell duration is updated as:
original_duration − delay_decrease + delay_increase.dw.endis shifted forward by the cumulative delay plus any new delay added at this dwell (delay_increase).
Temporary
timedeltacolumns are dropped before returning.- Parameters:
dw (
DwellSet) – Dwell dataset with float delay columns (hours). Modified in-place and returned.params (
dict) –Pipeline parameters. Expected keys:
delay_columns(dict): Maps logical delay-role keys to column names indw.data. Required keys:cumul_hrs(cumulative delay at dwell start),dwell_hrs(original dwell duration),decrease_hrs(delay reduction at this dwell), andincrease_hrs(new delay added at this dwell).delay_unit(str): Time unit of the float delay columns (e.g."h"for hours), passed topd.to_timedelta.
- Return type:
DwellSet- Returns:
The updated
DwellSetwith adjusteddw.start,dw.end, and dwell duration column; temporary timedelta columns removed.
- laurel.pipelines.evaluate_impacts.nodes.build_class_frame(locs, vehs, params)[source]
Build a complete cross-product of vehicle and location class combinations.
Creates a DataFrame with one row for every (vehicle class × location class) pair, regardless of whether any observed dwells exist for that combination. This ensures that downstream probability and count computations include unobserved class pairs as explicit zeros rather than missing rows.
- Parameters:
locs (
DataFrame) – Locations table containing location-class columns.vehs (
DataFrame) – Vehicle table containing vehicle-class columns.params (
dict) –Pipeline parameters. Expected keys:
veh_class_cols(list[str]): Column names invehsthat define vehicle classes (e.g. primary operating distance class).loc_class_cols(list[str]): Column names inlocsthat define location classes (e.g. freight-activity cluster).
- Return type:
DataFrame- Returns:
A DataFrame with one row per class combination and one column per class dimension.
- laurel.pipelines.evaluate_impacts.nodes.build_eval_columns(pcols, group_cols)[source]
Inject the reporting group columns into the shared pipeline column dictionary.
group_colsspecifies the spatial aggregation level for load profiles (e.g. substation ID or county code). It is passed separately in the Kedro pipeline so that the same node logic can produce substation-level or county-level outputs by swapping the parameter.- Parameters:
pcols (
dict) – Shared column-name dictionary to update.group_cols (
list) – List of column names that define the reporting grouping (e.g.["substation_id"]).
- Return type:
dict- Returns:
The updated
pcolsdictionary with"group_cols"set togroup_cols.
- laurel.pipelines.evaluate_impacts.nodes.build_time_ordered_slice(events, params, pcols)[source]
Convert cumulative profile columns into per-time-step power increments.
manage_chargingproduces cumulative profile values (e.g. cumulative kWh delivered since the start of the charging session). The bootstrap sampler insample_profiles_nodeneeds differences — the incremental kWh delivered in each time step — to correctly aggregate contributions from multiple dwells.Sorts events by dwell ID and within-window time, computes first-order differences of each profile column within each dwell group (using the first row’s absolute value for the initial diff), and drops the original cumulative columns. The result is then sorted by within-window time to facilitate time-aligned aggregation during sampling.
- Parameters:
events (
DataFrame) – Sliced events DataFrame fromslice_events.params (
dict) –Pipeline parameters. Expected key:
slice_time_col(str): Column holding within-window time offset, used as the secondary sort key.
pcols (
dict) –Shared column-name dictionary. Expected keys:
dw_col(str): Dwell ID column used for grouping.profile_cols(dict): Maps logical profile keys to cumulative column names (input).diff_cols(dict): Maps the same logical keys to output difference column names; must be paired 1-to-1 withprofile_cols.
- Return type:
DataFrame- Returns:
A DataFrame sorted by
slice_time_colwith cumulative profile columns replaced by per-time-step difference columns.- Raises:
ValueError – If
profile_colsanddiff_colsdo not have the same key set.
- laurel.pipelines.evaluate_impacts.nodes.calc_utilization(summs, params, pcols)[source]
Compute charging utilisation for each profile type from bootstrap summaries.
Utilisation is defined as:
utilisation = cumulative_energy / (peak_power × window_duration)
where
window_durationis derived from the slice frequency in the configured time unit. A utilisation of 1.0 means the charger ran at peak power for the entire window; values below 1.0 reflect partial utilisation.- Parameters:
summs (
DataFrame) – Bootstrap summary DataFrame with cumulative-energy and peak-power columns for each profile type.params (
dict) –Pipeline parameters. Expected keys:
slice_freq(str): Window frequency string used to compute total window duration (e.g."1D").dur_unit(str): Time unit for window duration conversion (e.g."h"for hours).
pcols (
dict) –Shared column-name dictionary. Expected keys:
profile_cols(dict): Maps profile-type keys to column names (used to iterate over types).util_cols(dict): Maps profile-type keys to output utilisation column names.cum_cols(dict): Maps profile-type keys to cumulative-energy column names.peak_cols(dict): Maps profile-type keys to peak-power column names.
- Return type:
DataFrame- Returns:
The
summsDataFrame with one new utilisation column per profile type appended.
- laurel.pipelines.evaluate_impacts.nodes.compress_bootstrap_profiles(profs, params, pcols)[source]
Reduce bootstrap profile draws to quantile envelopes over discretised time.
Accepts either a full bootstrap profile DataFrame (legacy path) or a compact accumulator dict produced by the memory-efficient path in
sample_profiles_node(). The two paths share the same output schema.Accumulator path (
profsis adict): The accumulator maps(region, time_bin)tuples to a list-of-lists of observed non-zero profile values. Because timestamps are already atdiscrete_freqresolution (guaranteed bydiscretize_sparse_profiles()), no time-bin re-assignment is needed.n_bootstrapsis read directly fromparams["n_bootstraps"]and used as the uniform possible-observation count for zero-padding.DataFrame path (
profsis apd.DataFrame): Unchanged from the original implementation — usesAdaptiveTimeGrouperto assign time bins, computes per-bin possible-observation counts, and callsNonzeroGroupedSummarizer.summarize(). Retained for backward compatibility and as the fallback when multi-timezone DST correction is required.- Parameters:
profs (
DataFrame|dict) –Either:
A
pd.DataFramewith columns for region, time, bootstrap ID, and profile values (legacy path).A
dictmapping(region, time_bin)group keys to list-of-lists of observed values (accumulator path).
params (
dict) –Pipeline parameters. Expected keys:
time_col(str): Within-window time column.discrete_freq(str): Coarsened time-bin frequency string.slice_freq(str): Original window frequency string.bootstrap_id_col(str): Bootstrap draw ID column (DataFrame path only).n_bootstraps(int): Total bootstrap count used for zero-padding (accumulator path only).quantiles(list[float]): Quantile levels to compute.
pcols (
dict) –Shared column-name dictionary. Expected keys:
timezone_col(str): Time zone column (DataFrame path only).group_cols(list[str]): Region grouping columns.profile_cols(dict): Maps profile-type keys to value column names to quantile.
- Return type:
DataFrame- Returns:
A DataFrame indexed by (region, time-bin) with quantile columns for each profile type.
- laurel.pipelines.evaluate_impacts.nodes.compress_bootstrap_summaries(summs, params, pcols)[source]
Reduce per-bootstrap scalar summaries to quantiles and means by region.
Each bootstrap draw produces scalar summary statistics per region (total energy, peak power, utilisation). This node compresses those draws to a compact set of distributional statistics — quantiles via
NonzeroGroupedSummarizerand means — so that downstream analysis can reason about the uncertainty across bootstrap draws without storing all raw draws.- Parameters:
summs (
DataFrame) – Bootstrap summary DataFrame with columns for region, bootstrap ID, and scalar summary values (cumulative energy, peak power, utilisation).params (
dict) –Pipeline parameters. Expected keys:
bootstrap_id_col(str): Column holding the bootstrap draw ID, used to determinen_bootstraps.quantiles(list[float]): Quantile levels to compute.
pcols (
dict) –Shared column-name dictionary. Expected keys:
group_cols(list[str]): Region grouping columns.cum_cols(dict): Maps profile-type keys to cumulative-energy column names.peak_cols(dict): Maps profile-type keys to peak-power column names.util_cols(dict): Maps profile-type keys to utilisation column names.
- Return type:
DataFrame- Returns:
A DataFrame indexed by region with quantile columns (from
NonzeroGroupedSummarizer) and mean columns (suffixed_mean) for each summary statistic.
- laurel.pipelines.evaluate_impacts.nodes.compute_adoption_totals(adopts, params, pcols)[source]
Derive the number of electrified vehicles per class from adoption forecasts.
Filters the adoption dataset to the rows of interest, groups by vehicle class, and sums electrified and total vehicle counts. If
override_with_paramsis enabled, the adoption fractions from the scenario parameter file replace those derived from the forecast data (while the total fleet size is preserved), allowing sensitivity analysis without re-running the full forecast pipeline.- Parameters:
adopts (
DataFrame) – Adoption forecast DataFrame with vehicle counts by fuel type and vehicle class.params (
dict) –Pipeline parameters. Expected keys:
filter_totals(dict): Column-value pairs used to subsetadoptsbefore aggregation (e.g. year, region).totals_column(str): Column holding vehicle counts.fuel_type_col(str): Column identifying fuel type.electrified_fuel_types(list[str]): Fuel type values considered electrified (e.g.["BEV"]).override_with_params(bool): IfTrue, replace observed adoption fractions with those inadoption_fracs.adoption_fracs(dict): Scenario-specific adoption fractions by vehicle class, used whenoverride_with_paramsisTrue.
pcols (
dict) –Shared column-name dictionary. Expected keys:
veh_class_cols(list[str]): Vehicle-class column names used for grouping.electrified_col(str): Name for the electrification boolean column.
- Return type:
DataFrame- Returns:
A DataFrame indexed by vehicle class with columns for total fleet size and number of electrified vehicles.
- laurel.pipelines.evaluate_impacts.nodes.compute_class_dwell_counts(classes, dw, params)[source]
Count observed dwells per (vehicle class × location class × electrification).
Groups dwell rows by all class columns plus the electrification flag, counts rows in each group, and pivots electrification into two columns (
n_dwells_electrified,n_dwells_obs). The counts are left-joined onto the full class cross-product so that unobserved combinations appear as zeros.- Parameters:
classes (
DataFrame) – Full class cross-product DataFrame produced bybuild_class_frame.dw (
DwellSet) – Dwell dataset that includes vehicle-class, location-class, and electrification columns.params (
dict) –Pipeline parameters. Expected keys:
veh_class_cols(list[str]): Vehicle-class column names.loc_class_cols(list[str]): Location-class column names.electrified_col(str): Boolean column indw.dataindicating whether the vehicle is electrified.
- Returns:
n_dwells_obs(total observed dwells) andn_dwells_electrified(observed dwells from electrified vehicles). Unobserved combinations are filled with 0.- Return type:
The
classesDataFrame with two new integer columns
- laurel.pipelines.evaluate_impacts.nodes.compute_class_probs(cls, params)[source]
Estimate electrification probability and expected dwell counts per class.
This is the core data-fusion step of Module 5. It combines observed dwell location distributions from the telematics dataset with electrification adoption forecasts to produce, for every (location class, vehicle class) combination:
P(L|V) — probability of a dwell falling in location class L given vehicle class V, estimated from observed dwell counts.
P(L,V) — joint probability of a dwell falling in class pair (L, V), estimated from observed dwell counts.
P(E|V) — probability of a vehicle being electrified given vehicle class V, derived from the adoption forecast.
P(E|L,V) — probability of electrification given both location and vehicle class, estimated by
ElectProbLocalizer(logistic regression + correction term that reconciles the class-level P(E|V) with location- specific dwell observations).Expected electrified dwells —
P(L|V) × dwell_rate × n_vehs_in_class × P(E|L,V), representing the expected number of electrified dwells per time unit at each class combination.P(observed & electrified | L,V) — inverse propensity weight numerator used downstream to correct for observation bias.
- Parameters:
cls (
DataFrame) – Class-level DataFrame containing dwell counts (fromcompute_class_dwell_counts), fleet totals (fromcompute_adoption_totals), and dwell rates (fromcompute_dwell_rate_vclass), joined together.params (
dict) –Pipeline parameters. Expected keys:
loc_class_cols(list[str]): Must contain exactly one element (multi-column location classes are not yet supported).veh_class_cols(list[str]): Must contain exactly one element.columns(dict): Maps logical names to column names incls, includingn_dwells_obs,n_dwells_obs_elect,n_vehs_in_class,n_vehs_electrified_in_class, anddwell_rate.out_columns(dict): Output column name mapping with keysprob_obs_electandn_elect_dwells_expected.
- Return type:
DataFrame- Returns:
A DataFrame indexed by (location class, vehicle class) with columns for
prob_obs_elect(inverse propensity weight numerator) andn_elect_dwells_expected(expected electrified dwell count per class per unit time).- Raises:
ValueError – If
loc_class_colsorveh_class_colscontains more than one element.
- laurel.pipelines.evaluate_impacts.nodes.compute_dwell_rate_vclass(veh_classes, dw, vehs, params, pcols)[source]
Compute the observed dwell rate (dwells per unit time) by vehicle class.
Counts the total number of dwells per vehicle, divides by each vehicle’s observation window length (in the configured time unit, optionally excluding weekends), aggregates to the vehicle-class level, and joins the resulting rate back onto the vehicle-class cross-product frame.
The rate is used by
compute_class_probsto scale expected dwell counts from per-vehicle-per-time-unit rates to absolute expected dwell counts across the target fleet.- Parameters:
veh_classes (
DataFrame) – Vehicle-class cross-product frame to augment.dw (
DwellSet) – Dwell dataset (only dwell counts per vehicle are needed).vehs (
DataFrame) – Vehicle table containing observation start/end timestamps and vehicle-class columns.params (
dict) –Pipeline parameters. Expected keys:
dwell_count_col(str): Temporary column name for per-vehicle dwell counts.obs_dur_cols(dict): Maps"start"and"end"to the column names invehsholding observation window timestamps.time_unit(str): Time unit for the observation window (e.g."day").filter_out_weekends(bool): IfTrue, exclude weekend days from each vehicle’s observation window length.dwell_rate_col(str): Output column name for the dwell rate (dwells per time unit).
pcols (
dict) –Shared column-name dictionary. Expected key:
veh_class_cols(list[str]): Vehicle-class column names.
- Return type:
DataFrame- Returns:
The
veh_classesDataFrame with new columns for total dwell count, total observation time, vehicle count, and dwell rate appended.
- laurel.pipelines.evaluate_impacts.nodes.filter_dwells_post_prob(dw, pcols)[source]
Retain only dwells from vehicles flagged as electrified.
Applied after class probabilities have been computed. Non-electrified vehicle dwells — which were retained through
filter_dwells_pre_probto inform probability estimation — are now dropped because they should not contribute to load profiles.Zero-charge dwells from electrified vehicles are still retained here; they are needed for the dwell-based sampling in
sample_profiles_node(a zero-charge dwell at a location still informs that location’s dwell count) and are only removed inget_dwells_nonzero.Note: This function returns a new
DwellSet(copy_without_data), not an in-place modification ofdw.- Parameters:
dw (
DwellSet) – Full dwell dataset including non-electrified vehicle rows.pcols (
dict) –Shared column-name dictionary. Expected key:
electrified_col(str): Boolean column marking electrified vehicles (produced bysummarize_vehicles).
- Return type:
DwellSet- Returns:
A new
DwellSetcontaining only rows from electrified vehicles, with theelectrified_colcolumn dropped.
- laurel.pipelines.evaluate_impacts.nodes.filter_dwells_pre_prob(dw, params, pcols)[source]
Filter dwells to the subset used for probability estimation.
Retains dwells that fall within the desired time spans (optionally weekdays only) and drops rows with missing values in required columns.
Critically, this filter does not drop:
Zero-charge dwells — they contribute to the observed dwell location distribution used by the logistic regression in
compute_class_probs.Dwells from non-electrified vehicles — they are needed to estimate P(electrified | location class, vehicle class).
These rows are only removed in the later
filter_dwells_post_probstep, after class probabilities have been computed.- Parameters:
dw (
DwellSet) – Dwell dataset to filter. Modified in-place and returned.params (
dict) –Pipeline parameters. Expected keys:
filter_out_weekends(bool): IfTrue, drop dwells whose start and end both fall on Saturday or Sunday.drop_na_cols(list[str] | None): Column names for which rows with NaN values should be dropped. PassNoneto skip.
pcols (
dict) – Shared pipeline column-name dictionary (not used directly in this node but required by the Kedro pipeline signature for consistency with sibling nodes).
- Return type:
DwellSet- Returns:
The filtered
DwellSetwith out-of-scope and missing-value rows removed.
- laurel.pipelines.evaluate_impacts.nodes.filter_locs_pre_prob(locs, params)[source]
Drop locations that are missing required fields before probability estimation.
- Parameters:
locs (
DataFrame) – Locations table (one row per H3 hexagon or reporting unit).params (
dict) –Pipeline parameters. Expected keys:
drop_na_cols(list[str]): Column names for which rows with NaN values should be dropped (e.g. missing freight-activity class assignment).
- Return type:
DataFrame- Returns:
The filtered locations table.
- laurel.pipelines.evaluate_impacts.nodes.get_dwells_nonzero(dw, pcols)[source]
Return a new DwellSet containing only dwells with nonzero charging.
Zero-charge dwells were retained through the probability-estimation steps because they contribute to observed dwell location distributions. They are dropped here because a sampled dwell with zero charge does not contribute energy to a load profile even if it is selected during bootstrap sampling.
This function creates a new
DwellSetview rather than modifyingdwin-place, so the original (with zero-charge rows) remains available if needed.- Parameters:
dw (
DwellSet) – Dwell dataset containing both zero-charge and nonzero-charge rows.pcols (
dict) –Shared column-name dictionary. Expected key:
charge_col(str): Column holding simulated charge amount (kWh); rows where this is<= 0are excluded.
- Return type:
DwellSet- Returns:
A new
DwellSetcontaining only rows where charge amount is positive.
- laurel.pipelines.evaluate_impacts.nodes.get_unique_series(df, col, dropna=True)[source]
Extract the unique sorted values of a column or index level from a DataFrame.
- Parameters:
df (
DataFrame) – Source DataFrame to query.col (
str) – Name of a column or index level whose unique values to extract.dropna (
bool) – IfTrue(default), remove NaN values from the result.
- Return type:
Series- Returns:
A sorted
pd.Seriesof unique values withnameset tocol.- Raises:
RuntimeError – If
colis not found indf.columnsordf.index.names.
- laurel.pipelines.evaluate_impacts.nodes.localize_time_from_hexes(df, params, pcols)[source]
Convert UTC timestamps to local time using time zones inferred from H3 hexagons.
Optionally looks up the IANA time zone for each row from the H3 hex centroid (via
calc_time_zones_from_hexes), then converts one or more UTC timestamp columns to local time (viacalc_local_time). Daylight saving time ambiguities are introduced naturally by working with timezone-naïve local times.- Parameters:
df (
DataFrame) – DataFrame containing a hex column and UTC timestamp columns. Modified and returned.params (
dict) –Pipeline parameters. Expected keys:
get_tz(bool): IfTrue, derive time zones from hex centroids and write them topcols["timezone_col"]. Set toFalseif the time zone column is already present.time_cols_source(list[str]): UTC timestamp columns to localise.time_cols_local(list[str]): Output column names for the corresponding local-time timestamps.sort_result(bool): IfTrue, sortdfbysort_colsafter localisation.sort_cols(list[str]): Columns to sort by whensort_resultisTrue.
pcols (
dict) –Shared column-name dictionary. Expected keys:
hex_col(str): Column holding H3 hex IDs.timezone_col(str): Column to store or read IANA time zone strings.
- Return type:
DataFrame- Returns:
The updated DataFrame with local-time columns added and, optionally, a time zone column added and rows sorted.
- laurel.pipelines.evaluate_impacts.nodes.manage_charging(dw, params)[source]
Convert dwell records into time-resolved charging event records.
Instantiates the configured charging manager (selected from
_MANAGER_MAPbyparams["charging_manager"]) and callsget_events()to expand each dwell into one or more charging events with start time, duration, and power columns. Power values are rounded to reduce output file size.- Parameters:
dw (
DwellSet) – Dwell dataset with charge amounts, dwell timings, and vehicle parameters required by the charging manager.params (
dict) –Pipeline parameters. Expected keys:
charging_manager(str): Key into_MANAGER_MAPselecting the charging manager class (e.g."constant_power").input_cols(dict): Column-name mappings forwarded to the charging manager constructor.round_decimals(int): Number of decimal places to which power values are rounded.
- Return type:
DataFrame- Returns:
A DataFrame of charging events, with one row per time step per dwell, containing start time, duration, and power columns.
- laurel.pipelines.evaluate_impacts.nodes.sample_profiles_node(dw, events, locs, classes, params, pcols)[source]
Assemble per-region load profiles via inverse-propensity-weighted bootstrap sampling.
This is the core computational node of Module 6. For each bootstrap draw it:
Builds four sparse correspondence matrices:
Ga (Dwells × Hexagons): indicates which hex each observed dwell occurred in.
Cy (Hexagons × Freight-activity classes): maps each hex to its freight-activity class.
Be (Events × Dwells): maps each charging event to its parent dwell.
Rho (Regions × Hexagons): maps each hex to its reporting region (substation or county).
Computes inverse propensity weights Ω for each dwell using
P(observed & electrified | location, vehicle class)fromcompute_class_probs, normalised column-wise so that each hex (or class) receives unit total weight.Computes expected dwell counts per hex (
m_hex_expected) by distributing class-level expected counts uniformly across hexes within each class.Calls
sample_profilesfor each bootstrap draw. Whenn_bootstraps > 1, draws are distributed across Dask workers; large arrays are scattered withbroadcast=Trueto avoid redundant serialisation.Reduces each bootstrap profile DataFrame to one scalar per (region, time_bin) on the fly and accumulates into a compact dict, then returns confidence diagnostics (observed vs. expected hex dwell counts).
- Parameters:
dw (
DwellSet) – Electrified, nonzero-charge dwell dataset with dwell IDs and hex columns.events (
DataFrame) – Time-ordered per-time-step charging event increments frombuild_time_ordered_slice.locs (
DataFrame) – Locations table with hex IDs, region assignments, and freight-activity class codes.classes (
DataFrame) – Class-level probability and expected-count table fromcompute_class_probs.params (
dict) –Pipeline parameters. Expected keys:
sample_self(bool): Allow a hex to sample from its own observed dwells in stage 2 of the bootstrap.sample_class(bool): Allow a hex to sample from class-level pooled dwells in stage 2.n_bootstraps(int): Number of bootstrap draws (≥ 1).loc_group_col(str): Column inlocsholding the freight-activity class (categorical dtype required).dwell_prob_obs_elect_col(str): Column in dwell data holding the inverse propensity weight numerator P(obs & elect | L, V).n_dwells_expected_elect_col(str): Column inclassesholding expected electrified dwell counts per class.max_first_stage_options(int): Maximum number of candidate dwell indices considered in stage 1 of sampling (caps memory use).time_col(str): Within-window time column inevents.slice_freq(str): Window frequency string (e.g."1D").discrete_freq(str): Discretisation frequency for profile compression.summary_suffixes(dict): Suffix strings for cumulative and peak summary columns.bootstrap_id_col(str): Column name for bootstrap draw ID in output.master_seed(int | None): Random seed base; draw i usesmaster_seed + iif set.
pcols (
dict) –Shared column-name dictionary. Expected keys:
hex_col,group_cols,dw_col,duration_col,profile_cols,diff_cols,cum_cols,peak_cols.
- Returns:
boot_profs_accum (
dict): Compact accumulator mapping(region, time_bin)tuples to a list-of-lists of observed non-zero profile values across bootstrap draws, one inner list per profile column. Passed directly tocompress_bootstrap_profiles().boot_summs (
pd.DataFrame): Per-bootstrap, per-region scalar summaries (cumulative energy, peak power).hex_confidence (
pd.DataFrame): Per-hexagon observed and expected dwell counts for diagnostic use.debug_partition (
dict): Empty dict whenparams["save_bootstrap_profiles"]isFalse(production). WhenTrue, maps zero-padded bootstrap ID strings to the full per-bootstrap profile DataFrame captured before accumulation, for writing to thebootstrap_profiles_debug_scenariocatalog entry.
- Return type:
A four-tuple of
- Raises:
ValueError – If
pcols["group_cols"]has more than one element, or if a profile column name conflicts with an argument ofsample_profiles, or ifn_bootstraps < 1.NotImplementedError – If
locs[loc_group_col]is not a categorical dtype.
- laurel.pipelines.evaluate_impacts.nodes.slice_events(events, params, pcols)[source]
Partition charging events into time-of-day windows and handle cross-window carry-over.
A “slice” is a fixed-frequency time window (e.g. one day) within which charging load profiles are assembled. Events that begin in one window may extend into the next; this function handles that carry-over by synthesising initialisation events at the start of each new window.
The algorithm:
Optionally clip event durations to the window frequency to preserve the characteristic shape of a charging session without allowing a single event to dominate multiple windows.
Use
IntervalBeginSpreaderto detect all events that cross a window boundary and generate initialisation rows at each affected window start.Concatenate original events with initialisation rows, assign each row to its window (
slice_id_col), and compute each event’s within-window start time (slice_time_col) as a timezone-naive offset from midnight.
- Parameters:
events (
DataFrame) – Charging events DataFrame frommanage_charging.params (
dict) –Pipeline parameters. Expected keys:
slice_id_col(str): Output column for the window identifier (floored timestamp of the window start).slice_time_col(str): Output column for within-window offset (timezone-naiveTimestamprelative to epoch).source_time_col(str): Existing column holding each event’s absolute start timestamp (must be timezone-naïve).slice_freq(str): Pandas frequency string defining the window size (e.g."1D"for one day).clip_dur_to_slice_freq(bool): Whether to clip event durations to the window frequency before spreading.
pcols (
dict) –Shared column-name dictionary. Expected keys:
dw_col(str): Dwell ID column used as grouping key.duration_col(str): Event duration column.profile_cols(dict): Profile value columns to carry over into initialisation rows.
- Return type:
DataFrame- Returns:
A DataFrame of events (original + initialisation rows) with
slice_time_coladded andsource_time_col/slice_id_coldropped.- Raises:
RuntimeError – If
source_time_colcontains timezone-aware timestamps.
- laurel.pipelines.evaluate_impacts.nodes.summarize_vehicles(dw, vehs, params)[source]
Compute per-vehicle performance metrics and flag implausible simulation results.
Evaluates whether each vehicle’s simulated electrification is operationally plausible by measuring two failure modes:
Range deaths — trips where the vehicle ran out of charge (
dead_energy_col < 0). Circle trips (where the origin and destination hexagon are identical) are separated as non-addressable deaths that do not indicate a range design problem.Charging delays — extra time accumulated because the vehicle had to wait for charging. Shift-level delay is computed as the cumulative delay at the end of the shift minus the delay at the start, plus any delay added at the preceding refresh stop. Vehicle-days where simulation artefacts caused delay to drop (e.g. after a death event) are clipped to zero.
Metrics are first aggregated by shift, then summarised to the vehicle level as point statistics (counts, percentages) and distributional statistics (quantiles across shifts).
Vehicles that exceed any of the three thresholds — death rate, relative delay fraction, or absolute maximum delay — are flagged for exclusion from load-profile assembly.
- Parameters:
dw (
DwellSet) – Post-simulation dwell dataset. Must contain energy, delay, dwell duration, and refresh columns specified inparams.vehs (
DataFrame) – Vehicle-level table to augment with summary statistics and inclusion flags. Returned with new columns appended.params (
dict) –Pipeline parameters. Expected keys:
dead_energy_col(str): Column holding remaining energy after each trip; negative values indicate a death event.charge_energy_col(str): Column holding energy charged at each dwell; non-NaN values indicate resuscitation after a death.dwell_dur_col(str): Column holding dwell duration (hours).refresh_col(str): Boolean column marking refresh dwells (used to delineate shift boundaries).delay_inc_hrs_col(str): Column holding the delay increment added at each dwell.delay_hrs_col(str): Column holding cumulative delay (hours).summary_cols(dict): Mapping from logical names to output column names for shift-level metrics (e.g.max_delay,shift_delay,shift_dur,shift_dur_delayed,shift_dur_delayed_thresh,delay_frac).shift_max_dur_hrs(float): Threshold shift duration (hours) above which a delayed shift is considered to violate operational constraints.quantiles(list[float]): Quantile levels for distributional summaries.delay_frac_thresh_quantile(float): Quantile of the delay fraction distribution to report and use for thresholding.max_delay_thresh_quantile(float): Quantile of the max-delay distribution to report and use for thresholding.thresholds(dict): Exclusion thresholds with sub-keyspct_shifts_w_deaths_max(float),delay_frac_max(float), anddelay_hrs_max(float).death_rate_col(str): Column name for the death-rate statistic used in the threshold check.drop_events_col(str): Output boolean column marking vehicles to exclude from load profiles.electrified_col(str): Output boolean column (inverse ofdrop_events_col).
- Return type:
DataFrame- Returns:
The
vehsDataFrame with per-vehicle summary statistics, quantile distribution columns, and two boolean flag columns (drop_events_colandelectrified_col) appended.
laurel.pipelines.evaluate_impacts.pipeline module
Kedro pipeline definition for the evaluate_impacts pipeline.
Wires the nodes from laurel.pipelines.evaluate_impacts.nodes into a single Pipeline object.
For full documentation of each node’s inputs, outputs, and algorithm,
see laurel.pipelines.evaluate_impacts.nodes.
Module contents
This is a boilerplate pipeline ‘evaluate_impacts’ generated using Kedro 0.19.1