Skip to content
Open
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,7 @@
- Updates to all GH Actions, pre-commit, isort, and ruff versioning. [PR 904](https://github.com/NatLabRockies/H2Integrate/pull/904)
- Update API resource models to be able to be able to handle nonannual simulations. [PR 897](https://github.com/NatLabRockies/H2Integrate/pull/897)
- Move PySAM model instantiation for wind and solar performance models to the `compute()` method, and validate recalculated wind power curves. [PR 909](https://github.com/NatLabRockies/H2Integrate/pull/909)
- Add data resampling capabilities via pandas methods. [PR 893](https://github.com/NatLabRockies/H2Integrate/pull/893)

## 0.9 [August 10, 2026]

Expand Down
4 changes: 2 additions & 2 deletions docs/_static/class_hierarchy.html

Large diffs are not rendered by default.

6 changes: 6 additions & 0 deletions docs/resource/resource_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ If running a simulation that is longer than 1 year, then there are a few options
2. Use resource data from a list of filenames. This option is enabled when `resource_filename` is a list. The list must be the same length as the number of years needed for the simulation. The filenames provided as `resource_filename` are **not used if the site changes from the site defined in the configuration file**. This can be used with `resource_year_order`, where `resource_year_order` is used if the site changes. If `resource_year_order` is not specified, the year order is inferred from the data in the provided files under the 'year' column (unless using a TMY dataset, in which case the year order is inferred from the filename). Supplying `resource_year_order` alongside `resource_filename` avoids inferring the years and makes the replacement years explicit. The inferred or provided `resource_year_order` is used instead of `resource_year`.
3. Use a specified order of resource years. This option is enabled when `resource_year_order` is a list of resource years and `resource_filename` is an empty string (or not provided). The list must be the same length as the number of years needed for the simulation. The `resource_year_order` is used instead of `resource_year`. As mentioned above, this can be used with the `resource_filename` option.

## Resampling resource data

API resource data may have a different native timestep from the simulation, but H2I does not resample it by default. You may explicitly choose a pandas interpolation method for upsampling or a pandas aggregation method for downsampling in the resource's `resource_parameters`.

Set `upsample_method` or `downsample_method` in `resource_parameters` to explicitly select the pandas operation. If the data timestep and simulation timestep already match, neither setting is needed. When resampling is requested, H2I emits a warning naming the source timestep, target timestep, direction, and method. If a mismatch requires resampling but no corresponding method is set, H2I raises an error. See {py:func}`h2integrate.resource.utilities.time_tools.resample_resource_data_to_dt` for method details.



## Setting resource data for a technology
Expand Down
2 changes: 1 addition & 1 deletion h2integrate/core/inputs/validation.py
Original file line number Diff line number Diff line change
Expand Up @@ -155,7 +155,7 @@ def load_plant_yaml(finput):
n_timesteps = plant_config["plant"]["simulation"]["n_timesteps"]
dt = plant_config["plant"]["simulation"]["dt"]

if int(n_timesteps) * int(dt) != 31536000: # seconds in simulation must be seconds/year
if int(n_timesteps) * int(dt) != 31_536_000:
msg = (
"H2Integrate does not currently support simulations that are less than or "
"greater than 1-year. Please ensure that "
Expand Down
23 changes: 20 additions & 3 deletions h2integrate/resource/resource_baseclass.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
check_data_length,
contains_leap_day,
add_resource_start_end_times,
resample_resource_data_to_dt,
get_number_of_resource_years_needed,
)
from h2integrate.resource.utilities.download_tools import download_from_api
Expand All @@ -40,6 +41,8 @@ class ResourceBaseAPIConfig(BaseConfig):
load resource files from. Defaults to "".
- **resource_filename** (*str*, optional): Filename to save resource data to or load
resource data from. Defaults to None.
- **resource_year_setting** (*str*, optional): How resource years are selected for
API datasets. Options include ``start_year``, ``year_order``, and ``filenames``.
- **valid_intervals** (*list[int]*): time interval(s) in minutes that resource data can be
downloaded in.

Expand All @@ -65,6 +68,12 @@ class ResourceBaseAPIConfig(BaseConfig):
resource_year_order (list, optional): Only used running a simulation requiring multiple
resource years. List of resource years in-order, such as [2012, 2011, 2013].
Defaults to None.
upsample_method (str | None, optional): Pandas interpolation method to use when the
simulation timestep is finer than the resource data timestep. Required only when
upsampling is needed; otherwise defaults to None and no automatic resampling occurs.
downsample_method (str | None, optional): Pandas resampling aggregation to use when the
simulation timestep is coarser than the resource data timestep. Required only when
downsampling is needed; otherwise defaults to None and no automatic resampling occurs.

Attributes:
dataset_desc (str): description of the dataset, used in file naming.
Expand All @@ -85,6 +94,9 @@ class ResourceBaseAPIConfig(BaseConfig):
resource_dir: Path | str | None = field(default=None)
include_leap_day: bool = field(default=False)
resource_year_order: list | None = field(default=None)
# Resampling methods used when the resource data's native timestep differs from the sim dt.
upsample_method: str | None = field(default=None)
downsample_method: str | None = field(default=None)


class ResourceBaseAPIModel(om.ExplicitComponent):
Expand Down Expand Up @@ -559,7 +571,6 @@ def get_data_for_year(
# filepath, and a new download isn't forced, then load data using `load_data()`
if filepath.is_file() and not forced_download:
data = self.load_data(filepath)
# NOTE: this where we could up/downsample
return data

# If the filepath (resource_dir/filename) does not exist, download data
Expand All @@ -571,7 +582,6 @@ def get_data_for_year(
if success:
# 6) Load data from the file created in Step 5 using `load_data()`
data = self.load_data(filepath)
# NOTE: this where we could up/downsample
return data

else:
Expand All @@ -591,8 +601,15 @@ def process_final_resource_data(self, resource_data):
Returns:
dict: resource_data after final processing and checks
"""
# NOTE: we could up/downsample here also
# Remove leap day (if needed) on the native data, resample to the simulation
# timestep, then clip to the number of timesteps requested by the simulation.
resource_data = process_leap_day(resource_data, self.config.include_leap_day)
resource_data = resample_resource_data_to_dt(
resource_data,
self.dt,
self.config.upsample_method,
self.config.downsample_method,
)
resource_data = clip_data_to_n_timesteps(resource_data, n_timesteps=self.n_timesteps)
resource_data = add_resource_start_end_times(resource_data)
check_data_length(resource_data, self.n_timesteps)
Expand Down
2 changes: 2 additions & 0 deletions h2integrate/resource/test/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,8 @@ def site_config(which, lat, lon, model, resource_year, model_name):
"resource_model": model,
"resource_parameters": {
"resource_year": resource_year,
"upsample_method": "time",
"downsample_method": "mean",
},
}
},
Expand Down
5 changes: 4 additions & 1 deletion h2integrate/resource/test/test_resource_api_baseclass.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,8 @@ def input_config(n_timesteps, resource_year, include_leap, resource_fname, yr_or
"include_leap_day": include_leap,
"resource_year_order": yr_order,
"resource_filename": resource_fname,
"upsample_method": "time",
"downsample_method": "mean",
"timezone": 0,
}

Expand Down Expand Up @@ -86,7 +88,7 @@ def get_data_for_year(
"month": dates.month.to_numpy().astype(float),
"day": dates.day.to_numpy().astype(float),
"hour": dates.hour.to_numpy().astype(float),
"minute": dates.hour.to_numpy().astype(float),
"minute": dates.minute.to_numpy().astype(float),
"ws": np.arange(len(dates), dtype=float),
"latitude": latitude,
"longitude": longitude,
Expand Down Expand Up @@ -304,6 +306,7 @@ def test_get_data_tmy_filenames_require_years_for_site_change(input_config):
[(2013, 2, False, "", None)],
)
def test_process_final_resource_data(subtests, input_config):
input_config["plant"]["simulation"]["dt"] = 172800
_, comp = _setup_resource_component(FakeResource, input_config)
dates = pd.to_datetime(["2012-02-28", "2012-02-29", "2012-03-01"])
data = {
Expand Down
230 changes: 230 additions & 0 deletions h2integrate/resource/utilities/test/test_time_tools.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
process_leap_day,
check_data_length,
contains_leap_day,
resample_resource_data_to_dt,
get_n_timesteps_from_year_list,
get_number_of_resource_years_needed,
)
Expand Down Expand Up @@ -117,6 +118,235 @@ def test_process_leap_day_without_february():
np.testing.assert_array_equal(result["day"], [1, 2])


@pytest.mark.unit
def _make_timeseries(n, freq_seconds, start="2012-01-01 00:00", values=None):
"""Build a resource data dict with time columns spaced at ``freq_seconds``."""

idx = pd.date_range(start=start, periods=n, freq=pd.Timedelta(seconds=freq_seconds))
if values is None:
values = np.arange(n, dtype=float)
return {
"wind_speed_100m": np.asarray(values, dtype=float),
"year": idx.year.to_numpy().astype(float),
"month": idx.month.to_numpy().astype(float),
"day": idx.day.to_numpy().astype(float),
"hour": idx.hour.to_numpy().astype(float),
"minute": idx.minute.to_numpy().astype(float),
"units": {"wind_speed_100m": "m/s"},
"site_lat": 1.0,
}


@pytest.mark.unit
def test_resample_no_op_when_dt_matches():
data = _make_timeseries(10, 3600)
result = resample_resource_data_to_dt(data, 3600)
assert result is data


@pytest.mark.unit
def test_resample_no_time_columns_raises():
data = {"wind_speed_100m": np.arange(10, dtype=float), "site_lat": 1.0}
with pytest.raises(ValueError, match="no time columns"):
resample_resource_data_to_dt(data, 1800)


@pytest.mark.unit
def test_resample_target_dt_larger_than_span_raises():
# 5 hourly samples span only a few hours; a 10-day timestep yields no full step
data = _make_timeseries(5, 3600)
with pytest.raises(ValueError, match="larger than the total time span"):
resample_resource_data_to_dt(data, 10 * 86400)


@pytest.mark.unit
def test_resample_non_increasing_timestamps_raises():
# The first two timestamps are identical, so the native timestep is non-positive
data = {
"wind_speed_100m": np.arange(3, dtype=float),
"year": np.array([2012, 2012, 2012], dtype=float),
"month": np.array([1, 1, 1], dtype=float),
"day": np.array([1, 1, 1], dtype=float),
"hour": np.array([0, 0, 1], dtype=float),
"minute": np.array([0, 0, 0], dtype=float),
}
with pytest.raises(ValueError, match="non-positive"):
resample_resource_data_to_dt(data, 1800)


@pytest.mark.unit
def test_upsample_interpolation(subtests):
# hourly data upsampled to 30-minute resolution
data = _make_timeseries(5, 3600, values=[0, 1, 2, 3, 4])
result = resample_resource_data_to_dt(data, 1800, upsample_method="time")

with subtests.test("upsampled length doubles"):
assert len(result["wind_speed_100m"]) == 10
expected = np.array([0, 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4], dtype=float)
with subtests.test("interpolated values match expectation"):
np.testing.assert_allclose(result["wind_speed_100m"], expected)
with subtests.test("minutes alternate on 30 minute grid"):
np.testing.assert_array_equal(result["minute"][:4], [0, 30, 0, 30])


@pytest.mark.unit
def test_downsample_average(subtests):
# 30-minute data downsampled to hourly resolution via pandas mean aggregation
data = _make_timeseries(10, 1800, values=[0, 1, 2, 3, 4, 5, 6, 7, 8, 9])
result = resample_resource_data_to_dt(data, 3600, downsample_method="mean")

with subtests.test("downsampled length halves"):
assert len(result["wind_speed_100m"]) == 5
expected = np.array([0.5, 2.5, 4.5, 6.5, 8.5], dtype=float)
with subtests.test("hourly means match expectation"):
np.testing.assert_allclose(result["wind_speed_100m"], expected)
with subtests.test("minutes stay on hour"):
np.testing.assert_array_equal(result["minute"], np.zeros(5))


@pytest.mark.unit
def test_resample_keeps_leap_day_excluded_at_new_timestep(subtests):
# Hourly data for a leap year with the leap day removed, resampled to 2-hour.
# February 29 must stay excluded in the regenerated calendar at the new timestep.

idx = pd.date_range("2012-01-01 00:30", "2012-12-31 23:30", freq="1h")
idx = idx[~((idx.month == 2) & (idx.day == 29))] # 8760 hourly, leap day removed
data = {
"wind_speed_100m": np.arange(len(idx), dtype=float),
"year": idx.year.to_numpy().astype(float),
"month": idx.month.to_numpy().astype(float),
"day": idx.day.to_numpy().astype(float),
"hour": idx.hour.to_numpy().astype(float),
"minute": idx.minute.to_numpy().astype(float),
}

result = resample_resource_data_to_dt(data, 7200, downsample_method="mean") # 2-hour timestep

with subtests.test("result length matches excluded leap year at 2 hour dt"):
assert len(result["wind_speed_100m"]) == 4380
with subtests.test("february 29 remains excluded"):
assert not ((result["month"] == 2) & (result["day"] == 29)).any()
with subtests.test("starts on january 1"):
assert int(result["month"][0]) == 1 and int(result["day"][0]) == 1
with subtests.test("ends on december 31"):
assert int(result["month"][-1]) == 12 and int(result["day"][-1]) == 31


@pytest.mark.unit
def test_resample_keeps_leap_day_when_present_at_new_timestep(subtests):
# Leap year with the leap day retained, resampled to 2-hour: Feb 29 must remain.

idx = pd.date_range("2012-01-01 00:30", "2012-12-31 23:30", freq="1h") # 8784, leap kept
data = {
"wind_speed_100m": np.arange(len(idx), dtype=float),
"year": idx.year.to_numpy().astype(float),
"month": idx.month.to_numpy().astype(float),
"day": idx.day.to_numpy().astype(float),
"hour": idx.hour.to_numpy().astype(float),
"minute": idx.minute.to_numpy().astype(float),
}

result = resample_resource_data_to_dt(data, 7200, downsample_method="mean") # 2-hour timestep

with subtests.test("result length matches leap year at 2 hour dt"):
assert len(result["wind_speed_100m"]) == 4392
with subtests.test("leap day remains present"):
assert ((result["month"] == 2) & (result["day"] == 29)).any()


@pytest.mark.unit
def test_downsample_preserves_mean():
rng = np.random.default_rng(0)
values = rng.random(24)
data = _make_timeseries(24, 900, values=values) # 15-min data
result = resample_resource_data_to_dt(data, 3600, downsample_method="mean") # hourly

assert len(result["wind_speed_100m"]) == 6
# overall mean is conserved by the averaging
np.testing.assert_allclose(result["wind_speed_100m"].mean(), values.mean())


@pytest.mark.unit
def test_downsample_with_calendar_gap_preserves_mean(subtests):
# Hourly data whose timestamps have a one-day gap in the middle (as when a leap
# day is removed). Resampling must treat the samples as an evenly spaced sequence,
# so the downsampled mean matches the exact pairwise mean.

part1 = pd.date_range("2012-02-28 00:30", periods=4, freq="1h")
part2 = pd.date_range("2012-03-01 00:30", periods=4, freq="1h")
idx = part1.append(part2)
values = np.arange(8, dtype=float)
data = {
"wind_speed_100m": values.copy(),
"year": idx.year.to_numpy().astype(float),
"month": idx.month.to_numpy().astype(float),
"day": idx.day.to_numpy().astype(float),
"hour": idx.hour.to_numpy().astype(float),
"minute": idx.minute.to_numpy().astype(float),
}
result = resample_resource_data_to_dt(data, 7200, downsample_method="mean") # to 2-hour

with subtests.test("downsampled length matches expected bins"):
assert len(result["wind_speed_100m"]) == 4
with subtests.test("pairwise averages preserved across gap"):
np.testing.assert_allclose(result["wind_speed_100m"], [0.5, 2.5, 4.5, 6.5])
with subtests.test("mean preserved despite gap"):
np.testing.assert_allclose(result["wind_speed_100m"].mean(), values.mean())


@pytest.mark.unit
def test_resample_scalar_metadata_preserved(subtests):
data = _make_timeseries(5, 3600, values=[0, 1, 2, 3, 4])
result = resample_resource_data_to_dt(data, 1800, upsample_method="time")
with subtests.test("units preserved"):
assert result["units"] == {"wind_speed_100m": "m/s"}
with subtests.test("site latitude preserved"):
assert result["site_lat"] == 1.0


@pytest.mark.unit
def test_resample_unknown_upsample_method_raises():
data = _make_timeseries(5, 3600)
# An invalid pandas interpolation method raises. The exact message differs across
# pandas versions, but the offending method name is always reported.
with pytest.raises(ValueError, match="not_a_method"):
resample_resource_data_to_dt(data, 1800, upsample_method="not_a_method")


@pytest.mark.unit
def test_resample_unknown_downsample_method_raises():
data = _make_timeseries(10, 1800)
# An invalid pandas aggregation raises. Different pandas versions raise different
# exception types (AttributeError vs ValueError) with different messages, so match
# on the offending method name that is common to all versions.
with pytest.raises((AttributeError, ValueError), match="not_a_method"):
resample_resource_data_to_dt(data, 3600, downsample_method="not_a_method")


@pytest.mark.unit
def test_resample_upsampling_without_method_raises():
data = _make_timeseries(5, 3600)
with pytest.raises(ValueError, match="requires upsampling"):
resample_resource_data_to_dt(data, 1800)


@pytest.mark.unit
def test_resample_downsampling_without_method_raises():
data = _make_timeseries(10, 1800)
with pytest.raises(ValueError, match="requires downsampling"):
resample_resource_data_to_dt(data, 3600)


@pytest.mark.unit
def test_resample_warns_when_resampling(subtests):
with subtests.test("upsampling warns"):
with pytest.warns(UserWarning, match="upsampling"):
resample_resource_data_to_dt(_make_timeseries(5, 3600), 1800, upsample_method="time")
with subtests.test("downsampling warns"):
with pytest.warns(UserWarning, match="downsampling"):
resample_resource_data_to_dt(_make_timeseries(10, 1800), 3600, downsample_method="mean")


@pytest.mark.unit
@pytest.mark.parametrize(
"n_timesteps,include_leap,expected_years",
Expand Down
Loading
Loading