Update to be able to experience the latest version even in this branch - #1088
Closed
danielelotito wants to merge 41 commits into
Closed
danielelotito wants to merge 41 commits into
danielelotito wants to merge 41 commits into
Conversation
* feat(operators): Add script operator type This allows dynamic operations to be recorded on discoveryspaces i.e. without having to install a separate operator plugin * refactor(operators): use semantic types * chore(orchestrate): address comments from code review
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com> Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
…1019) * test: add failing test for avoid_oom_recommender Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> * fix(autoconf): avoid_oom_recommender takes into account original number_gpus Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> * test(autoconf): remove breaking tests that were mocking functionality Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> * refactor(autoconf): replace recommend_min_gpu() with get_model_prediction_and_metadata() Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> * fix(autoconf): do not set workers and gpu when not recommending This is for the avoid_oom_recommender custom_experiment Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> --------- Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
) * refactor(cli): rename show entities to show measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update show entities references to show measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(examples): update show entities references to show measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(plugins): update show entities references to show measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(skills): update show entities references to show measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…al GPUs (#1025) Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
* refactor(core): add instantiation methods in sample store class Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): update call sites for sample store initialisation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(tests): revert incorrect change Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): raise error when sample store resource doesn't exist Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(cli): update parameter names Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): rename spec to specification Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Hey @AlessandroPomponio 👋 I ran your skills through `tessl skill review` at work and found some targeted improvements. Here's the full before/after: | Skill | Before | After | Change | |-------|--------|-------|--------| | using-ado-cli | 73% | 94% | +21% | | examining-ado-operations | 94% | — | — | | resource-yaml-creation | 94% | — | — | | examining-discovery-spaces | 94% | — | — | | formulate-discovery-problem | 88% | — | — | | remote-execution | 86% | — | — | | query-ado-data | 86% | — | — | | conduct-empirical-study | 85% | — | — | | examining-ado-project | 84% | — | — | Only `using-ado-cli` was changed — it had the most headroom and is the foundational CLI reference that most of your other skills point to. <details> <summary>Changes made</summary> **Description improvements (biggest impact):** - Rewrote the frontmatter description from a vague 197-char label into a concrete 441-char reference listing specific subcommands (`get`, `create`, `edit`, `show`, `describe`), key flags (`-o`, `--output-file`, `--use-latest`, `--set`, `--with`, `-l`), and `run_experiment` — all as natural trigger terms - Added explicit "Use when..." clause covering four trigger scenarios: writing/verifying ado CLI commands, looking up syntax/flags, debugging CLI output, explaining command patterns **Content improvements:** - Consolidated two redundant `--output-file` sections into a single "Output Format and File Handling" section — same info, fewer tokens, clearer guidance on when to use `--output-file` vs shell redirects - Added error recovery guidance to the iterative development workflow (`--dry-run` failure handling with common issues: missing experiment references, invalid parameter types, unresolvable `--use-latest`) - Streamlined the Terminology section into a focused "show Commands Quick Reference" table — kept the critical `show entities` vs `show results` distinction, trimmed redundant entity explanation - Condensed Documentation Best Practices by removing the verbose example documentation pattern (markdown-inside-markdown block) while retaining the actionable guidelines **Unchanged skills:** - examining-ado-operations (94%), resource-yaml-creation (94%), examining-discovery-spaces (94%), formulate-discovery-problem (88%), remote-execution (86%), query-ado-data (86%), conduct-empirical-study (85%), examining-ado-project (84%) — all left untouched </details> I also stress-tested your `remote-execution` skill against a few real-world task evals and it held up really well on dispatching operations to remote Ray clusters with local plugin shipping via `additionalFiles`. Kudos for that. Honest disclosure — I work at @tesslio where we build tooling around skills like these. Not a pitch — just saw room for improvement and wanted to contribute. Want to self-improve your skills? Just point your agent (Claude Code, Codex, etc.) at [this Tessl guide](https://docs.tessl.io/evaluate/optimize-a-skill-using-best-practices) and ask it to optimize your skill. Ping me — [@yogesh-tessl](https://github.com/yogesh-tessl) — if you hit any snags. Thanks in advance 🙏 Signed-off-by: Yogesh Rao <yogesh-tessl@users.noreply.github.com> Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Co-authored-by: Yogesh Rao <yogesh-tessl@users.noreply.github.com> Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
…on (#1002) * feat(core): Enable recording package provenance for plugins * feat(core): Provide provenance for actuators and operators * feat(core): Add provenance to created resources * refactor(core): provenance implementation * docs(core): fix typing * refactor(operators): Use PackageProvenance in OperatorMetadata * refactor(core): consolidate package version lookup code in PackageProveance * feat: define pep440 version type with validator
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…mote runtime setup options (#998) * feat(remote_execution): Allow setting params related to job env creation If you are expecting many jobs to start simultaneously with autoscaling it can be useful to changes these from the ray defaults. * feat(remote_execution): Allow setting params related to job env creation If you are expecting many jobs to start simultaneously with autoscaling it can be useful to changes these from the ray defaults. * test(remote_execution): Additional tests. * feat(actuators): launch supervisor The most common reason for operation hangs with distributed ray is ray tasks never starting. Current actuators "fire-and-forget" so there is no code for checking if a task ever started. LaunchSupervisor provides a generic way to monitor if launched ray tasks start and handle if they don't. * feat(actuators): integration tests * test(actuators): launch supervisor * feat(actuators): notify launch supervisor when executor queues a result Notify the supervisor immediately after queue.put and ensure only timeout tasks that never reached RUNNING. * docs(remote): doc additional options * chore(refactor): tidying code * chore(refactor): update filenames * feat(actuator): handle pending tasks separately * test(actuators): fixes for supervisor tests * fix(actuators): pass actor handle not self * fix(operators): stringify exceptions in discovery space manager When manager catches an exception when monitoring measurement queue it informs subscribers. Prior to this change it passed the exception - however this may contain non-ray serializable components causing a second exception crashing the manager. This change fixes the problem * feat(core): Harden supervisor polling against transient lookup failures and duplicate failures. * fix(tests): Usding x With pytest-xdist the pytest workers for ray supervision tests each call ray.init() in a session fixture, so parallel workers each start a separate local Ray cluster on the same machine. The tests then query Ray’s State API to observe task scheduling state; under parallel cluster startup that API is often slow or temporarily unavailable (ConnectionError, empty results). The supervisor under test then sees OTHER instead of real states, causing false timeouts and missed pending-resource failures. This then causes tests to fail as the required behaviour is not observed. To avoid this this commit creates a xdist_group to keep these tests on one worker (one Ray cluster, serial execution). tox must then use --dist loadgroup, because worksteal ignores xdist_group and the tests remain parallel. * fix(actuators): stabilize executor supervisor against Ray API issues The following problems were observed in tests in CI * Ray State API becomes unreliable after many tests on the same xdist worker — list_tasks returns empty results or ServerUnavailable, so tasks look like OTHER instead of their real state. * list_tasks retries were too short (3 quick attempts) for CI-level API lag, so lookups failed before the State API recovered. * FAILED state fell through to launch timeout when grace hadn’t elapsed yet, causing false “did not start within Xs” failures on healthy or completed tasks. * Resource timeout was blocked by seen_running — a brief false RUNNING report could prevent pending-resource failures from ever firing. * Launch timeout didn’t re-check mark_completed in _check_pending, leaving a race where a duplicate failure could still be emitted. This commit solves them as follows: * Ray State API becomes unreliable and list_tasks retries too short — fixed by more list_tasks retries (8), longer backoff, and extra delay when error is ServerUnavailable. * FAILED fell through to launch timeout — fixed by always returning after FAILED (only emit failure once grace has elapsed). * Resource timeout blocked by seen_running — fixed by removing not pending.seen_running from resource-wait and OTHER resource-timeout checks. * Launch timeout race with mark_completed — fixed by re-checking completed_request_ids in _check_pending before emitting launch timeout. * fix(actuators): harden executor supervisor tests Issues: * Ray State API becomes unreliable after many tests on the same xdist worker — list_tasks returns empty results or ServerUnavailable, so tasks look like OTHER instead of their real state. * State lookup integration test was too brittle — hard-failed after 5s if PENDING_NODE_ASSIGNMENT never appeared, even when the API was just temporarily unavailable under parallel load. Fixes: * Ray State API becomes unreliable (state-lookup test) — fixed by polling up to 15s, only stopping on PENDING_NODE_ASSIGNMENT, and pytest.skip when the API stays unavailable under parallel load. * State lookup integration test too brittle — same changes as above (longer window + skip instead of hard fail). * fix(ci): reduce ray noise * fix(logging): make configure_logging() idempotent within a process Modules that call configure_logging() at module level can be lazily imported inside a typer CliRunner.invoke() call, at which point sys.stderr is the runner's capture buffer. The resulting root logger handler then points to a closed stream after the invoke returns, causing "ValueError: I/O operation on closed file" to appear in subsequent invocations' result.output. A _logging_configured flag ensures the full handler setup runs only once; subsequent calls merely update the log level. Ray workers are unaffected as each starts with a fresh interpreter. Exposed by the --dist worksteal → --dist loadgroup change in tox.ini. * refactor(actuators): naming in executor supervisor * docs(core): update docstring * docs(core): fix typing * chore(refactor): comments from code review * chore(refactor): comments from code review * fix(core): readd validator 0 is disallowed as a value by ray so can't use ge=-1 and keep ray semantics. * Apply suggestions from code review Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> --------- Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com>
…n resources (#1028) * feat(core): no validation on get * fix(tests): manual downcast of parameter models * refactor(validation): use pydantic validation contexts * chore(ruff): lint * chore(lint): black --------- Signed-off-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
* feat(core): add filtering for measurements in sample store Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * perf(core): optimise db metadata reflection Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): rename files Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): remove useless comments Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): avoid accessing maps twice Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> * refactor(core): simplify table existence check Reflect will raise an error if at least one doesn't exist, we don't need to check explicitly Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): remove stray : Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(core): add comment clarifying behaviour Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
* build(deps): update dependencies Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com> * fix(test): avoid duplicate parameterization error Pytest 9.1.0 would raise duplicate parametrization of 'valid_ado_project_context' Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> Co-authored-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
#1040) Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(sfttrainer_download_hf_weights): remove runtime_env With some version of ray the ray task raises exceptions that urllib3 is not found even though the package is installed in the base virtual environment. Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> * refactor(sfttrainer_download_hf_weights): do not use Since download_weights() no longer configures packages, there's no need to use ray at all. Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com> --------- Signed-off-by: Vassilis Vassiliadis <vassilis.vassiliadis@ibm.com>
* feat(cli): add ado show trace command Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(cli): implement show trace operation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): add documentation for show trace Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update show trace documentation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update show trace description Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> * docs(website): apply suggestions from code review Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> * refactor(core): do not allow filtering requests on operation_id Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(cli): do not allow filtering show trace on results Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): add example of getting request yaml Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update documentation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(cli): rename --include-results to --unroll-entities Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(cli): handle potential errors Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(cli): add missing --no-trunc to help Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(cli): avoid errors on empty df Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(cli): add invalid reason column for invalid measurements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): clarify documentation Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
* feat(core): create OperationMeasurementStatistics class Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(core): add operation_measurement_statistics Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * test: add tests for operation_measurement_statistics Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(cli): remove show requests and show results Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(skills): remove references to show requests/results Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): remove references to show requests/results Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): create ado migration guide Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs: update documentation Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> * docs: formatting Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
docs(website): minor fixes
* fix: fix ci #938 This is the main fix, even if batchsize > 1 is supposed to be stable after a fix by @michaelj which I don't recall exactly * fix: minor fix ci #938 This could make the process more stable * fix(test): exclude Catboost and set outout dirs cherrypicked from 450d4ee There is also a small refactoring * fix: typo in model name ref: https://auto.gluon.ai/stable/api/autogluon.tabular.models.html
* feat(core): add get_related_resources_by_relationship Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * test: add tests for get_related_resources_by_relationship Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): incorrect relationship semantic for actuatorconfigurations Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(core): support both directions Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor: updates Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): use max-hops instead of stop_resource_kind Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(core): support fetching start resources as well Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): improvements Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): updates Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): use set instead of tuple Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): reorder code Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(core): clarify event should never happen Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): simplify code using sets instead of lists Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(core): simplify comment Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): minor fixes Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): centralise query creation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): remove stray comma Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
Signed-off-by: DRL NextGen <220003231+DRL-NextGen@users.noreply.github.com>
…r_operation (#1062) Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…esources with get_resources_by_relationship (#1059) * refactor(core): replace getRelatedResourceIdentifiers and getRelatedResources with get_resources_by_relationship Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(cli): support configuring --max-hops in show related Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * test: add tests for show related Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update show related documentation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Users should now use ado show trace <resource> --filter Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…1039) * feat(vllm): add version checking utility for vLLM compatibility - Add VLLMVersionChecker class to parse vLLM versions from image values - Add supports_threadpool method to check version compatibility - Minimum version for threadpool support is 0.20.0 - Include comprehensive tests for version parsing and threadpool support detection Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * chore(vllm_performance): remove comments Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * refactor(vllm): update supports_threadpool to accept only version strings - Add extract_version_from_image() helper method to extract version from container image strings - Update supports_threadpool() to only accept version strings (e.g., '0.20.1') - Rewrite all tests to use the new helper method and pass only version strings - Remove tests for non-existent parse_version method - Add comprehensive tests for version extraction from various image formats - All tests passing (13/13 in test_version_utils.py, 24/24 in vllm_performance plugin) - Pre-commit hooks passing (ruff, black, codespell, copyright headers) Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Add vLLM threadpool version validation and enable list hashability in group samplers Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * feat(core): Enable hashing of lists Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * feat(vllm_performance): Add threadpool support Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * doc(vllm_performance): update docs Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * chore(vllm_performance): update property from threadpool to use_threadpool and set default quay image Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Include threadpool parameters in environment definition to prevent incorrect deployment reuse Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * chore(vllm_performance): update default image values Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Add threadpool support for vLLM geospatial models and fix list hashing in grouped sampling Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Update orchestrator/core/discoveryspace/group_samplers.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * refactor(vllm_performance): rename threadpool to use_threadpool and change type to bool Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * test(vllm_performance): remove non-meaningful exception tests Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/yamls/vllm_deployment_space.yaml Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * chore(vllm_performance): remove unnecessary conditional assignment Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * chore(vllm_performance): fix ruff error Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): update logic to check image version Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/version_utils.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * test(vllm_performance): update tests description Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): add strict check on image value structure Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * feat(vllm_performance): constraint logic on use_threadpool Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): treat renderer_num_workers=0 as no threadpool to enable v0.18.0 testing, while positive values of renderer_num_workers enable threadpool functionalities Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): set renderer_num_workers default to 0 and add threadpool handling tests Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): raise InvalidImageStructureError for empty image values in vLLM performance actuator Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * Update plugins/actuators/vllm_performance/ado_actuators/vllm_performance/experiment_executor.py Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> * fix(vllm_performance): split walrus operator Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): update function call from _determine_threadpool_usage to _is_threadpool_requested Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): fix pre-commit Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): fix image version error handling Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): remove use_threadpool occurrences Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): remove unnecessary test Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * fix(vllm_performance): remove use_threadpool references Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> * doc(vllm_performance): make it clear that the number of renderers need to be an integer Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> --------- Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com> Signed-off-by: mgazz <michele.gazzetti1@ibm.com> Co-authored-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Co-authored-by: Christian Pinto <55737893+christian-pinto@users.noreply.github.com>
* feat(cli): add ado show trace space/store Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(website): update section title Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> Signed-off-by: Alessandro Pomponio <10339005+AlessandroPomponio@users.noreply.github.com> Co-authored-by: Michael Johnston <66301584+michael-johnston@users.noreply.github.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
…tatistics (#1068) Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* build(hooks): add new pre-commit hooks Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix: autofix files with pre-commit hooks Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(docs): only one CRD per file Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(hooks): --unsafe flag is required for mkdocs-specific tags Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(autoconf): avoid reformatting autogluon models Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* feat(core): calculate statistics for spaces Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * fix(core): prevent NULLs in db function Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): update space stats Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* feat(cli): add -o stats option to ado get operation Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(skills): clarify get -o stats Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * docs(skills): simplify mentions of -o stats Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
* refactor(core): use sample store identifier instead of uri in mapping Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * refactor(core): remove samplestore_statistics_for_stores Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> * feat(cli): support samplestores in ado get -o stats Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com> --------- Signed-off-by: Alessandro Pomponio <alessandro.pomponio1@ibm.com>
Collaborator
Author
|
@michael-johnston please have a quick look, if it is not straightforward I can easily cherrypick |
Member
|
I’ve already made the merge locally yesterday and forgot to push - I’ll push in next few mins |
Collaborator
|
I would advise on restarting from a clean slate on main with the cplex branch, especially since the changes there seem to have already mostly been merged to main |
Member
|
I've already merged main into the branch. The only changes are to solve_mip experiment - #1089 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Some commits would help here such as 6ec4c89 , I ask this merge to let me continue using this and main interchangeably for all except the solve_mip experiment