-
Notifications
You must be signed in to change notification settings - Fork 56
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[BUG] EvaluationReport.to_file() writes RFC-8259-invalid JSON (literal NaN), has no explicit encoding, and can leave a truncated file
area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesCore eval framework: Case, Experiment, task handler, evaluation data storesarea-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilitybugSomething isn't workingSomething isn't workingStatus: Open.#384 In strands-agents/evals;[BUG] TrajectoryEvaluator silently drops user-supplied tools on serialization and exposes no public
toolsattributearea-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsbugSomething isn't workingSomething isn't workingStatus: Open.#381 In strands-agents/evals;[BUG] Experiment.to_file() truncates the destination before serializing — a serialization error destroys the existing file
area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesCore eval framework: Case, Experiment, task handler, evaluation data storesarea-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilitybugSomething isn't workingSomething isn't workingStatus: Open.#382 In strands-agents/evals;[BUG] parse_timestamp: millisecond epochs silently misparsed, plus two robustness gaps
area-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityarea-tracingTrace/session ingestion: providers, session mappers, extractors, telemetry/OTELTrace/session ingestion: providers, session mappers, extractors, telemetry/OTELbugSomething isn't workingSomething isn't workingStatus: Open.#378 In strands-agents/evals;Interop: Case[InputT, OutputT] <-> EvalPort testcase.json
area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesCore eval framework: Case, Experiment, task handler, evaluation data storesdesignProposal or discussion about API shape, architecture, or UX before implementationProposal or discussion about API shape, architecture, or UX before implementationStatus: Open.#376 In strands-agents/evals;[FEATURE] Add metadata() to remaining evaluators
area-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#363 In strands-agents/evals;[BUG] CorrectnessEvaluator reference mode renders a blank USER QUERY when a tool call precedes the final answer
area-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsbugSomething isn't workingSomething isn't workingPythonLanguage: PythonStatus: Open.#355 In strands-agents/evals;[FEATURE] Failure cohort analysis for evaluation results
area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesCore eval framework: Case, Experiment, task handler, evaluation data storesarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#348 In strands-agents/evals;[FEATURE] Evaluator metadata taxonomy (tier, method, description)
area-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#350 In strands-agents/evals;[FEATURE] Risk classification utility for multi-run results (bug vs flaky vs pass)
area-coreCore eval framework: Case, Experiment, task handler, evaluation data storesCore eval framework: Case, Experiment, task handler, evaluation data storesarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#349 In strands-agents/evals;[FEATURE] Tool efficiency evaluator (LLM-as-judge, trajectory-based)
area-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#345 In strands-agents/evals;[FEATURE] Could-Not-Evaluate verdict for evaluators that cannot score a case
area-devxDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityDeveloper experience: papercuts, confusing public APIs, error messages, ergonomics, usabilityarea-evaluatorsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsEvaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metricsenhancementNew feature or requestNew feature or requestStatus: Open.#346 In strands-agents/evals;