Skip to content

Update to be able to experience the latest version even in this branch - #1088

Closed
danielelotito wants to merge 41 commits into
maj_cplex_mipfrom
main
Closed

danielelotito wants to merge 41 commits into
maj_cplex_mipfrom
main

Conversation

@danielelotito

Copy link
Copy Markdown
Collaborator

Some commits would help here such as 6ec4c89 , I ask this merge to let me continue using this and main interchangeably for all except the solve_mip experiment

michael-johnston and others added 30 commits June 5, 2026 10:44
* feat(operators): Add script operator type

This allows dynamic operations to be recorded on discoveryspaces i.e. without having to install a separate operator plugin

* refactor(operators): use semantic types

* chore(orchestrate): address comments from code review
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
…1019)

* test: add failing test for avoid_oom_recommender

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

* fix(autoconf): avoid_oom_recommender takes into account original number_gpus

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

* test(autoconf): remove breaking tests that were mocking functionality

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

* refactor(autoconf): replace recommend_min_gpu() with get_model_prediction_and_metadata()

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

* fix(autoconf): do not set workers and gpu when not recommending

This is for the avoid_oom_recommender custom_experiment

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

---------

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
)

* refactor(cli): rename show entities to show measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update show entities references to show measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(examples): update show entities references to show measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(plugins): update show entities references to show measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(skills): update show entities references to show measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…al GPUs (#1025)

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
* refactor(core): add instantiation methods in sample store class

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): update call sites for sample store initialisation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(tests): revert incorrect change

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): raise error when sample store resource doesn't exist

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(cli): update parameter names

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): rename spec to specification

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Hey @AlessandroPomponio 👋

I ran your skills through `tessl skill review` at work and found some targeted improvements. Here's the full before/after:

| Skill | Before | After | Change |
|-------|--------|-------|--------|
| using-ado-cli | 73% | 94% | +21% |
| examining-ado-operations | 94% | — | — |
| resource-yaml-creation | 94% | — | — |
| examining-discovery-spaces | 94% | — | — |
| formulate-discovery-problem | 88% | — | — |
| remote-execution | 86% | — | — |
| query-ado-data | 86% | — | — |
| conduct-empirical-study | 85% | — | — |
| examining-ado-project | 84% | — | — |

Only `using-ado-cli` was changed — it had the most headroom and is the foundational CLI reference that most of your other skills point to.

<details>
<summary>Changes made</summary>

**Description improvements (biggest impact):**
- Rewrote the frontmatter description from a vague 197-char label into a concrete 441-char reference listing specific subcommands (`get`, `create`, `edit`, `show`, `describe`), key flags (`-o`, `--output-file`, `--use-latest`, `--set`, `--with`, `-l`), and `run_experiment` — all as natural trigger terms
- Added explicit "Use when..." clause covering four trigger scenarios: writing/verifying ado CLI commands, looking up syntax/flags, debugging CLI output, explaining command patterns

**Content improvements:**
- Consolidated two redundant `--output-file` sections into a single "Output Format and File Handling" section — same info, fewer tokens, clearer guidance on when to use `--output-file` vs shell redirects
- Added error recovery guidance to the iterative development workflow (`--dry-run` failure handling with common issues: missing experiment references, invalid parameter types, unresolvable `--use-latest`)
- Streamlined the Terminology section into a focused "show Commands Quick Reference" table — kept the critical `show entities` vs `show results` distinction, trimmed redundant entity explanation
- Condensed Documentation Best Practices by removing the verbose example documentation pattern (markdown-inside-markdown block) while retaining the actionable guidelines

**Unchanged skills:**
- examining-ado-operations (94%), resource-yaml-creation (94%), examining-discovery-spaces (94%), formulate-discovery-problem (88%), remote-execution (86%), query-ado-data (86%), conduct-empirical-study (85%), examining-ado-project (84%) — all left untouched

</details>

I also stress-tested your `remote-execution` skill against a few real-world task evals and it held up really well on dispatching operations to remote Ray clusters with local plugin shipping via `additionalFiles`. Kudos for that.

Honest disclosure — I work at @tesslio where we build tooling around skills like these. Not a pitch — just saw room for improvement and wanted to contribute.

Want to self-improve your skills? Just point your agent (Claude Code, Codex, etc.) at [this Tessl guide](https://docs.tessl.io/evaluate/optimize-a-skill-using-best-practices) and ask it to optimize your skill. Ping me — [@yogesh-tessl](https://github.com/yogesh-tessl) — if you hit any snags.

Thanks in advance 🙏

Signed-off-by: Yogesh Rao <yogesh-tessl@users.noreply.github.com>
Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Co-authored-by: Yogesh Rao <yogesh-tessl@users.noreply.github.com>
Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
…on (#1002)

* feat(core): Enable recording package provenance for plugins

* feat(core): Provide provenance for actuators and operators

* feat(core): Add provenance to created resources

* refactor(core): provenance implementation

* docs(core): fix typing

* refactor(operators): Use PackageProvenance in OperatorMetadata

* refactor(core): consolidate package version lookup code in PackageProveance

* feat: define pep440 version type with validator
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…mote runtime setup options (#998)

* feat(remote_execution): Allow setting params related to job env creation

If you are expecting many jobs to start simultaneously with autoscaling it can be useful to changes these from the ray defaults.

* feat(remote_execution): Allow setting params related to job env creation

If you are expecting many jobs to start simultaneously with autoscaling it can be useful to changes these from the ray defaults.

* test(remote_execution): Additional tests.

* feat(actuators): launch supervisor

The most common reason for operation hangs with distributed ray is ray tasks never starting. Current actuators "fire-and-forget" so there is no code for checking if a task ever started.

LaunchSupervisor provides a generic way to monitor if launched ray tasks start and handle if they don't.

* feat(actuators): integration tests

* test(actuators): launch supervisor

* feat(actuators): notify launch supervisor when executor queues a result

Notify the supervisor immediately after queue.put and ensure only timeout tasks that never reached RUNNING.

* docs(remote): doc additional options

* chore(refactor): tidying code

* chore(refactor): update filenames

* feat(actuator): handle pending tasks separately

* test(actuators): fixes for supervisor tests

* fix(actuators): pass actor handle not self

* fix(operators): stringify exceptions in discovery space manager

When manager catches an exception when monitoring measurement queue it informs subscribers.

Prior to this change it passed the exception - however this may contain non-ray serializable components causing a second exception crashing the manager.

This change fixes the problem

* feat(core): Harden supervisor polling against transient lookup failures and duplicate failures.

* fix(tests): Usding x

With pytest-xdist the pytest workers for ray supervision tests each call ray.init() in a session fixture, so parallel workers each start a separate local Ray cluster on the same machine.

The tests then query Ray’s State API to observe task scheduling state; under parallel cluster startup that API is often slow or temporarily unavailable (ConnectionError, empty results).
The supervisor under test then sees OTHER instead of real states, causing false timeouts and missed pending-resource failures. This then causes tests to fail as the required behaviour is not observed.

To avoid this this commit creates a xdist_group to keep these tests on one worker (one Ray cluster, serial execution).

tox must then use --dist loadgroup, because worksteal ignores xdist_group and the tests remain parallel.

* fix(actuators): stabilize executor supervisor against Ray API issues

The following problems were observed in tests in CI

* Ray State API becomes unreliable after many tests on the same xdist worker — list_tasks returns empty results or ServerUnavailable, so tasks look like OTHER instead of their real state.
* list_tasks retries were too short (3 quick attempts) for CI-level API lag, so lookups failed before the State API recovered.
* FAILED state fell through to launch timeout when grace hadn’t elapsed yet, causing false “did not start within Xs” failures on healthy or completed tasks.
* Resource timeout was blocked by seen_running — a brief false RUNNING report could prevent pending-resource failures from ever firing.
* Launch timeout didn’t re-check mark_completed in _check_pending, leaving a race where a duplicate failure could still be emitted.

This commit solves them as follows:

* Ray State API becomes unreliable and list_tasks retries too short — fixed by more list_tasks retries (8), longer backoff, and extra delay when error is ServerUnavailable.
* FAILED fell through to launch timeout — fixed by always returning after FAILED (only emit failure once grace has elapsed).
* Resource timeout blocked by seen_running — fixed by removing not pending.seen_running from resource-wait and OTHER resource-timeout checks.
* Launch timeout race with mark_completed — fixed by re-checking completed_request_ids in _check_pending before emitting launch timeout.

* fix(actuators): harden executor supervisor tests

Issues:

* Ray State API becomes unreliable after many tests on the same xdist worker — list_tasks returns empty results or ServerUnavailable, so tasks look like OTHER instead of their real state.
* State lookup integration test was too brittle — hard-failed after 5s if PENDING_NODE_ASSIGNMENT never appeared, even when the API was just temporarily unavailable under parallel load.

Fixes:

* Ray State API becomes unreliable (state-lookup test) — fixed by polling up to 15s, only stopping on PENDING_NODE_ASSIGNMENT, and pytest.skip when the API stays unavailable under parallel load.
* State lookup integration test too brittle — same changes as above (longer window + skip instead of hard fail).

* fix(ci): reduce ray noise

* fix(logging): make configure_logging() idempotent within a process

Modules that call configure_logging() at module level can be lazily imported
inside a typer CliRunner.invoke() call, at which point sys.stderr is the
runner's capture buffer. The resulting root logger handler then points to a
closed stream after the invoke returns, causing "ValueError: I/O operation on
closed file" to appear in subsequent invocations' result.output.

A _logging_configured flag ensures the full handler setup runs only once;
subsequent calls merely update the log level. Ray workers are unaffected as
each starts with a fresh interpreter.

Exposed by the --dist worksteal → --dist loadgroup change in tox.ini.

* refactor(actuators): naming in executor supervisor

* docs(core): update docstring

* docs(core): fix typing

* chore(refactor): comments from code review

* chore(refactor): comments from code review

* fix(core): readd validator

0 is disallowed as a value by ray so can't use ge=-1 and keep ray semantics.

* Apply suggestions from code review

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>

---------

Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
…n resources (#1028)

* feat(core): no validation on get

* fix(tests): manual downcast of parameter models

* refactor(validation): use pydantic validation contexts

* chore(ruff): lint

* chore(lint): black

---------

Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
* feat(core): add filtering for measurements in sample store

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* perf(core): optimise db metadata reflection

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): rename files

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): remove useless comments

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): avoid accessing maps twice

Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

* refactor(core): simplify table existence check

Reflect will raise an error if at least one doesn't exist, we don't need to check explicitly

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): remove stray :

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(core): add comment clarifying behaviour

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
* build(deps): update dependencies

Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>

* fix(test): avoid duplicate parameterization error

Pytest 9.1.0 would raise duplicate parametrization of 'valid_ado_project_context'

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Co-authored-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
#1040)

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(sfttrainer_download_hf_weights): remove runtime_env

With some version of ray the ray task raises exceptions that urllib3 is not found
even though the package is installed in the base virtual environment.

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

* refactor(sfttrainer_download_hf_weights): do not use

Since download_weights() no longer configures packages, there's no need to use ray at all.

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>

---------

Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
* feat(cli): add ado show trace command

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(cli): implement show trace operation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): add documentation for show trace

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update show trace documentation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update show trace description

Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

* docs(website): apply suggestions from code review

Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

* refactor(core): do not allow filtering requests on operation_id

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(cli): do not allow filtering show trace on results

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): add example of getting request yaml

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update documentation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(cli): rename --include-results to --unroll-entities

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(cli): handle potential errors

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(cli): add missing --no-trunc to help

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(cli): avoid errors on empty df

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(cli): add invalid reason column for invalid measurements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): clarify documentation

Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
* feat(core): create OperationMeasurementStatistics class

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(core): add operation_measurement_statistics

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* test: add tests for operation_measurement_statistics

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(cli): remove show requests and show results

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(skills): remove references to show requests/results

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): remove references to show requests/results

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): create ado migration guide

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs: update documentation

Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

* docs: formatting

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
docs(website): minor fixes
* fix: fix ci #938

This is the main fix, even if batchsize > 1 is supposed to be stable after a fix by @michaelj which I don't recall exactly

* fix: minor fix ci #938

This could make the process more stable

* fix(test): exclude Catboost and set outout dirs

cherrypicked from 450d4ee
There is also a small refactoring

* fix: typo in model name

ref: https://auto.gluon.ai/stable/api/autogluon.tabular.models.html
* feat(core): add get_related_resources_by_relationship

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* test: add tests for get_related_resources_by_relationship

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): incorrect relationship semantic for actuatorconfigurations

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(core): support both directions

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor: updates

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): use max-hops instead of stop_resource_kind

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(core): support fetching start resources as well

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): improvements

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): updates

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): use set instead of tuple

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): reorder code

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(core): clarify event should never happen

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): simplify code using sets instead of lists

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(core): simplify comment

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): minor fixes

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): centralise query creation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): remove stray comma

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
…r_operation (#1062)

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…esources with get_resources_by_relationship (#1059)

* refactor(core): replace getRelatedResourceIdentifiers and getRelatedResources with get_resources_by_relationship

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(cli): support configuring --max-hops in show related

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* test: add tests for show related

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update show related documentation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Users should now use ado show trace <resource> --filter

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
mgazz and others added 11 commits June 23, 2026 08:06
…1039)

* feat(vllm): add version checking utility for vLLM compatibility

- Add VLLMVersionChecker class to parse vLLM versions from image values
- Add supports_threadpool method to check version compatibility
- Minimum version for threadpool support is 0.20.0
- Include comprehensive tests for version parsing and threadpool support detection

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* chore(vllm_performance): remove comments

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* refactor(vllm): update supports_threadpool to accept only version strings

- Add extract_version_from_image() helper method to extract version from container image strings
- Update supports_threadpool() to only accept version strings (e.g., '0.20.1')
- Rewrite all tests to use the new helper method and pass only version strings
- Remove tests for non-existent parse_version method
- Add comprehensive tests for version extraction from various image formats
- All tests passing (13/13 in test_version_utils.py, 24/24 in vllm_performance plugin)
- Pre-commit hooks passing (ruff, black, codespell, copyright headers)

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Add vLLM threadpool version validation and enable list hashability in group samplers

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* feat(core): Enable hashing of lists

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* feat(vllm_performance): Add threadpool support

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* doc(vllm_performance): update docs

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* chore(vllm_performance): update property from threadpool to use_threadpool and set default quay image

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Include threadpool parameters in environment definition to prevent incorrect deployment reuse

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* chore(vllm_performance): update default image values

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Add threadpool support for vLLM geospatial models and fix list hashing in grouped sampling

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Update orchestrator/core/discoveryspace/group_samplers.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py

Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* refactor(vllm_performance): rename threadpool to use_threadpool and change type to bool

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* test(vllm_performance): remove non-meaningful exception tests

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/yamls/vllm_deployment_space.yaml

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* chore(vllm_performance): remove unnecessary conditional assignment

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* chore(vllm_performance): fix ruff error

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): update logic to check image version

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* test(vllm_performance): update tests description

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): add strict check on image value structure

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* feat(vllm_performance): constraint logic on use_threadpool

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): treat renderer_num_workers=0 as no threadpool to enable v0.18.0 testing, while positive values of renderer_num_workers enable threadpool functionalities

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): set renderer_num_workers default to 0 and add threadpool handling tests

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): raise InvalidImageStructureError for empty image values in vLLM performance actuator

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py

Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>

* fix(vllm_performance): split walrus operator

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): update function call from _determine_threadpool_usage to _is_threadpool_requested

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): fix pre-commit

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): fix image version error handling

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): remove use_threadpool occurrences

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): remove unnecessary test

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* fix(vllm_performance): remove use_threadpool references

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

* doc(vllm_performance): make it clear that the number of renderers need to be an integer

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>

---------

Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>
Signed-off-by: mgazz <michele.gazzetti1@ibm.com>
Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
* feat(cli): add ado show trace space/store

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(website): update section title

Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…tatistics (#1068)

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* build(hooks): add new pre-commit hooks

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix: autofix files with pre-commit hooks

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(docs): only one CRD per file

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(hooks): --unsafe flag is required for mkdocs-specific tags

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(autoconf): avoid reformatting autogluon models

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* feat(core): calculate statistics for spaces

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* fix(core): prevent NULLs in db function

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): update space stats

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* feat(cli): add -o stats option to ado get operation

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(skills): clarify get -o stats

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* docs(skills): simplify mentions of -o stats

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(core): use sample store identifier instead of uri in mapping

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* refactor(core): remove samplestore_statistics_for_stores

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

* feat(cli): support samplestores in ado get -o stats

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>

---------

Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
@danielelotito danielelotito added the ci Enables CI integration label Jun 25, 2026
@danielelotito

Copy link
Copy Markdown
Collaborator Author

@michael-johnston please have a quick look, if it is not straightforward I can easily cherrypick

@michael-johnston

Copy link
Copy Markdown
Member

I’ve already made the merge locally yesterday and forgot to push - I’ll push in next few mins

@AlessandroPomponio

Copy link
Copy Markdown
Collaborator

I would advise on restarting from a clean slate on main with the cplex branch, especially since the changes there seem to have already mostly been merged to main

@michael-johnston

Copy link
Copy Markdown
Member

I've already merged main into the branch. The only changes are to solve_mip experiment - #1089

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Enables CI integration

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants