Skip to content

Nonannual resource models - #897

Merged
johnjasa merged 56 commits into
NatLabRockies:developfrom
elenya-grant:resource_api/multi_year
Oct 7, 2026
Merged

johnjasa merged 56 commits into
NatLabRockies:developfrom
elenya-grant:resource_api/multi_year

Conversation

@elenya-grant

@elenya-grant elenya-grant commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Multi-year resource

Updated the API resource models to be able to accommodate multiple years and partial years (with somewhat exhausting handling of leap years even though the solar resource models won't even use it...).

There are 3 different ways that a user can provide inputs if running multiple resource years. The attributes resource_year_order and resource_filename are what indicates what option the user is using. The options are:

  1. resource_year is start year: user provides resource_year as the first year to get resource data for, the following resource years are auto-calculated based on the number of timesteps and dt of the simulation. If running a simulation <= 1 year, then resource_filename can be a string (or not input, same logic as before) and resource_year_order must be None. Consecutive years are auto-calculated based on the number of timesteps in the simulation and the years available for the resource model. An error is raised if not enough future years are available.
resource_year: 2012
  1. list of filenames: users provides a list of filenames in the resource_filename attribute to use to get the resource data. The number of filenames must match the number of years needed to get the proper number of timesteps. resource_year_order can be provided too. If resource_year_order is not provided, then the resource years are estimated from the resource files on the first call (from filenames if using a TMY dataset, from the timeseries data if using other datasets). resource_year_order is only used if the site changes.
resource_year: 2012 # only used for config validation 
resource_filename: ['site0_resource_data_2022.csv', 'site0_resource_data_2023.csv']
resource_year_order: [2022, 2023] # this is optional, but used for back-up if the site changes
  1. resource year order: user provides a list of resource years in the resource_year_order attribute to use to get resource data, the number of resource years provided must match the number of years needed to get the proper number of timesteps.
resource_year: 2012 # only used for config validation, unused otherwise
resource_filenames: "" # not necessary
resource_year_order: [2012, 2014, 2015]

Basically, for running more than 1 year, the valid options are:

  • provide resource_year only (do not provide resource_year_order or resource_filename)
  • provide resource_filename as a list and resource_year_order as None (years are inferred)
  • provide resource_filename as a list and resource_year_order as a list
  • provide resource_year_order as a list and do not provide resource_filename

Other notable changes:

  • updated the solar TMY models to throw an error if trying to include leap-day
  • updated all resource model configs to have an include_leap_day attribute which is used for API calls and resource modifications (as necessary)
  • updated openmeteo models so that their load_data() method doesnt use self.config.resource_year and instead generalized that logic to be in the resource baseclass.

Most new-functionality is tested (right now) in h2integrate/resource/solar/test/test_multiyear_solar_resource.py

Work planned for Follow-On PR(s):

Section 1: Type of Contribution

  • Feature Enhancement
    • Framework
    • New Model
    • Updated Model
    • Tools/Utilities
    • Other (please describe):
  • Bug Fix
  • Documentation Update
  • CI Changes
  • Other (please describe):

Section 2: Draft PR Checklist

  • Open draft PR
  • Describe the feature that will be added
  • Fill out TODO list steps
  • Describe requested feedback from reviewers on draft PR
  • Complete Section 8: New Model Checklist (if applicable)

TODO:

  • add in last test that I'm thinking about
  • make issues for follow-on PRs
  • general clean-up (removing commented-out code, adding doc-strings and inline comments)
  • Work with @jaredthomas68 to merge related work
  • tests for individual functions in time_tools.py and data_tools.py (in-progress)
  • tests for methods in resource_baseclass.py (in-progress)
    • add tests for raised warnings and raised errors (few-edge cases are hard-to-test, noted in section 6 or 7)
    • add tests for resource-year estimation from TMY dataset logic with 'filenames' setting (plus error)
    • add tests for specific methods in the class, (process_final_resource_data, get_resource_years_from_start_year, check_resource_year)
  • finish adding checks (warnings and error-raising) for extraneous inputs in the config class
  • update documentation page
  • see if theres a way to reduce the number of new resource files added while still having comprehensive testing. Or, remove any old resource files that are unused.

Type of Reviewer Feedback Requested (on Draft PR)

Structural feedback:

  • done should the _check_resource_year method in the resource_baseclass be turned into a function? not turned into a function
  • done should the get_resource_years_from_start_year() method in the resource_baseclass be turned into a function? not turned into a function
  • done should there be a method that checks or redownloads data if all 3 things are true:
    1.include_leap_day is True
    2. the resource year is a leap year
    3. the resource year does not contain leap-day-data
  • donefor the filenames setting, what should happen if the site changes? Should we try to infer the resource year order based on the resource data that was loaded before the site changed? Should we throw a user-warning? Right now, it will just use the resource_year input in the config and repeat it for however many timesteps are needed.
  • outdated (I've moved this to the openmeteo models and out of the baseclass): the usage of clip_data_to_resource_year() in the resource baseclass is purely because of the OpenMeteo resource models. Should the usage of this function (and it's sister-function estimate_resource_year_from_data()), be moved to the OpenMeteo resource models load_data() method or continue to stay in the resource baseclass? (I think they should be moved to the OpenMeteo resource models).

Implementation feedback:

  • outdated thoughts on the resource_setting options?

Other feedback:

  • done anyone have suggestions on testing individual methods in the resource baseclass?

Section 3: General PR Checklist

  • PR description thoroughly describes the new feature, bug fix, etc.
  • Added tests for new functionality or bug fixes
  • Tests pass (If not, and this is expected, please elaborate in the Section 6: Test Results)
  • Documentation
    • Docstrings are up-to-date
    • Related docs/ files are up-to-date, or added when necessary
    • Documentation has been rebuilt successfully
    • Examples have been updated (if applicable)
  • CHANGELOG.md
    • At least one complete sentence has been provided to describe the changes made in this PR
    • After the above, a hyperlink has been provided to the PR using the following format:
      "A complete thought. [PR XYZ]((https://github.com/NatLabRockies/H2Integrate/pull/XYZ)", where
      XYZ should be replaced with the actual number.

Section 4: Related Issues

Section 5: Impacted Areas of the Software

The main changes are in the following files:

  • h2integrate/resource/resource_baseclass.py: changes to config and baseclass. A lot of the logic that was in get_data() is now in the method get_data_for_year().
  • h2integrate/resource/utilities/data_tools.py: new file with new tools
  • h2integrate/resource/utilities/time_tools.py: added and modified some existing tool functions
  • h2integrate/resource/solar/test/test_multiyear_solar_resource.py: new test file to test multi and partial year logic

Section 5.1: New Files

  • h2integrate/resource/utilities/data_tools.py: new file containing function/tools for processing/modifying/interpretting resource data
    • separate_timeseries_and_meta_data(): separate the resource data dictionary into two separate dictionaries, one with meta-data and one with timeseries data
    • append_timeseries_data(): append timeseries data from one dictionary to another
    • clip_data_to_n_timesteps(): clip timeseries data to a specified number of timesteps
    • clip_data_to_resource_year(): clip timeseries data to a single resource year. Only needed because of OpenMeteo resource models.
    • estimate_resource_year_from_data(): clip timeseries data to a single resource year. Only needed because of OpenMeteo resource models.

Section 5.2: Modified Files (of interest)

  • h2integrate/resource/resource_baseclass.py
    • ResourceBaseAPIConfig: added new attributes or types, noted in section 1. Also added an __attrs_post_init__() method to check for conflicting inputs based on the resource_year_setting.
    • ResourceBaseAPIModel
      • setup(): added some checks if running more than 1 resource year, basically additional checks based on simulation length and resource_year_setting
      • get_data(): this contains the first two "steps" in the previous get_data() method but now also loops through resource years.
      • get_data_for_year(): this contains the remainder of the logic that previously existed in get_data().
      • process_final_resource_data(): new method that does some final clean-up and testing of resource data after all the resource data for all the necessary resource years has been loaded
      • _check_resource_year(): new method that checks if a resource year is valid (could be simplified, noted with a Note in the code)
      • get_resource_years_from_start_year(): new method. This gets the valid future resource years if using resource_year_setting as 'start_year'
      • create_url(): added resource_year as an input
      • create_filename(): added resource_year as an input
  • h2integrate/resource/utilities/time_tools.py
    • is_leap_year: new function that returns boolean indicating whether the input year is a leap year or not
    • process_leap_day(): modified, the 'checking' part of this function was moved to a new function check_data_length()
    • check_data_length(): new function that contains the previous 'checking' logic that previously existed in process_leap_day().
    • get_number_of_resource_years_needed(): new function, determines the number of resource years needed to get enough data for the simulation length

Section 5.3: Deleted Files

  • resource_files/wind/*.srw
  • examples/25_sizing_modes/tech_inputs/hopp_config_tx.yaml
  • resource_files/solar/27.18624_-96.9516_psmv3_60_2013.csv
  • resource_files/solar/34.22_-102.75_psmv3_60_2013.csv
  • resource_files/solar/34.865371_-116.783023_psmv3_60_tmy.csv
  • resource_files/solar/39.7555_-105.2211_psmv3_60_2012.csv
  • resource_files/solar/47.5233_-92.5366_psmv3_60_2013.csv

Section 6: Additional Supporting Information

The resource models walk a fine-line between allowing opportunities for possible needed workarounds for single site simulations while also ensuring that when the site has changed, that resource data is updated in cases that it can be.
If a user provides resource data in the config as a dictionary, nothing about this resource data is checked for accuracy or alignment (the resource model basically acts as a pass-through, regardless of if the site changes or anything).

For example, if a resource filename is provided (and that file exists) and the site hasn't changed since setup() has been called, then there is nothing preventing a user from providing a resource file that:

  • doesn't match that specific dataset (which is sometimes nice for testing, for example, a file downloaded for the Himawari dataset can be used with a GOES resource model)
  • doesn't match the timezone specified in the plant config (could provide a file with data in UTC although the plant config specifies local timezone)
  • doesnt match the resource year specified in the resource config
  • doesnt match the site specified in the resource config
  • doesnt match the data interval that would be used based on the simulation timestep

However, if the site changes, then user-provided resource filenames are no longer used and resource data is ensured to be accurate for that site, resource year, timezone, timestep, etc.

This PR aims to maintain that fine-line and allow for the same flexibility as long as the site doesn't change.

Section 7: Test Results, if applicable

Things that are hard to test

  • testing warnings/errors for redownloading resource data if missing leap-day
    • test warning is raised ("Resource data is missing leap-day, attempting a forced redownload")
    • test weird edge-case if forced redownload doesnt work and error is raised ("Leap day data may not be available for this dataset.")

@elenya-grant elenya-grant changed the title Multi-year resource WIP: Multi-year resource Sep 24, 2026

@jaredthomas68 jaredthomas68 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trying again

Returns:
bool: True if the year is a leap year
"""
is_leap = (year % 100 == 0 and year % 400 == 0 and year % 4 == 0) or (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is hard to read. Is there a problem with using the pandas check?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is different than proces_leap_day(). This just checks to see if a year (integer) is a resource year - not if resource data contains Feb 29th.

)

# If only running 1 year, then only valid option is 'start_year'
if n_data_years == 1 and self.config.resource_year_setting != "start_year":

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this still allow for one year with a file?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes - if self.config.resource_year_setting it 'start_year' then it runs as it always used to.

@jaredthomas68 jaredthomas68 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lots of what I'm seeing looks very similar to what I have, but with an improved interface and excluding the sampling

@elenya-grant elenya-grant changed the title WIP: Multi-year resource Nonannual resource models Sep 29, 2026
@jaredthomas68 jaredthomas68 mentioned this pull request Oct 5, 2026
20 of 50 tasks

@genevievestarke genevievestarke left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall, I think this looks great! I love the functionality here, @elenya-grant!!

I don't have a ton of comments beyond what others have said. I like the ability to have non-consecutive years, and acknowledge that that's responsibility that we put on the user. If I understand it correctly, you can set resource_year, and the number of time steps, and that will give you a consecutive year simulation, right?

I think it makes sense to have resource_year_setting be inferred internally, though I would have to put more time in to see what all the different combinations of specifying data and simulation years that are available.

Finally, I think this is an important feature to have an example on, as well as some nice docs pages to cover all the input options!

1) Check if resource data was input. If not, continue to Step 2.
2) Determine the resource years and resource filenames to loop through based on
``config.resource_year_setting``
3) Loop through the resource years and resource filenames, calling ``get_data()``

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
3) Loop through the resource years and resource filenames, calling ``get_data()``
3) Loop through the resource years and resource filenames, calling ``get_data_for_year()``

This should be updated in the doc string, right?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks! updated!

Comment on lines +641 to +642
if not infer_years_from_files:
resource_years = self.inferred_resource_years

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a little confused what's happening in this line?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to add inline comments - let me know if its been cleared up?

# Get the number of hours in the simulation
hours_simulated = (self.dt / 3600) * self.n_timesteps

if future_hours_available < hours_simulated:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Because this is in hours, it feels like there could be issues with this logic when we move to smaller time steps (if we are using 10 minute time steps, the hours_simulated number could be 8760.5 and future_hours_available could be 8760 since we're going off of rounded hours, even if the data includes the data for 8760.5 times in 10-minutes time steps).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see what you're saying - but I think that resource data is always downloaded for one year. If we use 10 min intervals, then future_hours_available would still be 8760 hours (but 6x the number of timesteps). Maybe I don't understand your comment?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, working through it again, I think it should still be fine! Sorry, it didn't make sense to me earlier in the day :)

@elenya-grant
elenya-grant requested a review from johnjasa October 7, 2026 16:44
@elenya-grant
elenya-grant marked this pull request as ready for review October 7, 2026 16:52
@kbrunik
kbrunik self-requested a review October 7, 2026 17:26

@jaredthomas68 jaredthomas68 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for this @elenya-grant !!!!

if len(most_often_yr) == 1:
return most_often_yr[0]

if len(most_often_yr) > 1:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@johnjasa

johnjasa commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Thanks for your edits, Elenya! I agree with Kaitlin that if we're going to allow non-consecutive years, we should raise a warning to make it clear to users that's happening. Could you please add that in, or is there a good reason to not do that?

I've made new issues #907 and #908 to help track follow-on work based on comments. Please edit those directly if you want!

@johnjasa
johnjasa enabled auto-merge October 7, 2026 20:42
@johnjasa
johnjasa merged commit 42a0238 into NatLabRockies:develop Oct 7, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready for review This PR is ready for input from folks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants