Conversation
jimador
added this pull request to stack #2081
September 27, 2026 22:11
jimador
marked this pull request as ready for review
September 27, 2026 22:11
This was referenced Sep 27, 2026
Contributor
|
@jimador - please consider breaking PR into 2 PRs, thank you |
jimador
force-pushed
the
feature/decision-execution
branch
4 times, most recently
from
September 28, 2026 04:26
6055231 to
80d32fb
Compare
jimador
force-pushed
the
feature/decision-execution
branch
from
September 28, 2026 05:19
86e5181 to
b3c6d84
Compare
|
DecisionService.ask checks the whole request against the service's capabilities and hooks, then answers it through askNative or one assess, choose or rate call per question. DecisionResponseAssembler applies one set of answer rules for native providers. ObservedDecisionService observes each ask and its provider calls. NoOpDecisionService and StubDecisionService serve disabled configuration and tests. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
TypeSafeDecisionService answers a whole spec in one Jev systemOne call and maps each Jev answer to its question. LlmDecisionService prompts the chat model with the whole question set and reads a JSON array of answers. Both implement the one-question hooks for direct calls. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Covers asking, routing by hook, preflight, decorators, answer rules, the service families, observations, diagnostics and content capture. The TypeSafe page describes how a TypeSafe service answers a spec. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
…y value Specs no longer carry ids, so execution and telemetry compare the response with the request's spec through requireMatches. - A native response that does not match the request's spec is still an IllegalStateException naming the service. Its message now says what differs. - Answer events are recorded only for a response that matches the request's spec. - Tests cover a native answer whose options differ from the question's. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
…lpers Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
A choice question is a classification request built from the question, and every decision service can classify. Per-question execution now answers a choice through classify, so ChoiceAssessment and every choose implementation are gone. - Default capabilities are PROPOSITION and CHOICE; RATING still needs RatingAssessment. Preflight only asks for a hook on rating questions. - A per-question choice records embabel.ai.classification with operation=classify. - The stub's classify returns the choice scripted for the question name, then the outcome scripted for the category ids. - Tests cover a ClassificationRequest asked through ask on native and per-question services. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Update routing, capability derivation, preflight and telemetry tables for choice questions answered through classify. Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
jimador
force-pushed
the
feature/decision-execution
branch
from
September 28, 2026 16:07
b3c6d84 to
4bee754
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



This PR runs decision specs against a decision service.
DecisionService.asktakes an input and a whole spec, or aDecisionRequest, and returns one answer per question in spec order. The service answers every question in one provider operation when it has a native hook, and one question per call otherwise. Four service families implement it: TypeSafe Jev, prompted LLM, no-op and stub. Every ask on an observed service produces one observation with bounded tags, one event per answer, and log lines that name the cause and the fix. Jev and prompted services are observed as built.Routing
A service that implements
NativeQuestionSetExecutionanswers the whole spec in one call. Any other service answers one question per call in spec order: propositions throughPropositionAssessment(elseassess), choices throughclassifyand ratings throughRatingAssessment. A choice question becomesClassificationRequest.of(input, ClassificationSpec.of(question)), so the provider sees the question's own instructions and categories. AClassificationRequestis also aDecisionRequest, soask(request)accepts one andresponse.answer(spec.getQuestion())returns theClassificationResult.A stub answers the whole spec in one native call by default. After
perQuestion()it answers one question per call:A question kind the service's capabilities leave out throws
UnsupportedDecisionExceptionbefore any provider call. The message says whether the service lacks the hook or only leaves the kind out ofcapabilities(). Only rating questions need a hook. Here the service reports propositions only:capabilities()reports what a service accepts:Diagnosing
A failed provider call logs one WARN line on the provider's class logger, and the request it belongs to logs one on
DecisionExecution. Neither holds input, instructions, option text or provider text:A partial response on
DecisionExecution, a native provider's unusable answers onDecisionResponseAssembler, and an answer outside its question when questions go one per call:DEBUG on
com.embabel.common.ai.decision.spi.DecisionExecutionlogs the start of each ask and each question's outcome:TRACE logs request and response content only when
DecisionContentCapture.enable()has been called. Captured lines hold the input and provider output.Observations:
embabel.ai.askaskon an observed serviceoperation=ask,outcome,service,provider,question_countanswer.<kind>.<outcome>per questionembabel.ai.decisionaskNative,assessorratecalloperation,outcomeembabel.ai.classificationclassifycall, including a choice question asked on its ownoperation=classify,outcomeoperationaskonembabel.ai.ask.ask_native,assess,rateorclassifyon provider calls. A per-question choice recordsclassify.outcomeonembabel.ai.askcomplete,partial,request_failure,unsupported,exception,interrupted,cancelledquestion_count1,2-4,5-16,17+service,providerjev-latestandTypeSafe<kind>in an answer eventproposition,choice,rating<outcome>in an answer eventanswered,selected,no_match,inconclusive,failureProvider-call spans are children of the ask span. A meter handler counts each answer event as
embabel.ai.ask.answer.<kind>.<outcome>with the ask's tags. A tracing handler records it as a span event named<question name> <kind> <outcome>. Question names, roles, input and model output are absent from tag values. Service names come from configuration and code, which keepsservicebounded.Service families
TypeSafeDecisionServicepTruefor propositions. Category and confidence for choices. Score as expected level index and a level distribution for ratings, with no selected level.LlmDecisionServiceNoOpDecisionServiceUNAVAILABLE, so every question isFailure(UNAVAILABLE). Providernone, no provider callStubDecisionServiceperQuestion(), a choice throughclassifyIllegalStateExceptionfor a scripted outcome outside its question's options or levels.All four support proposition, choice and rating questions and report no limits. A prompted service asks the whole spec in one chat model call and reads one JSON object whose
answersarray holds one answer per question. Both providers send questions under the keysq1toqNin spec order, so question names stay with the caller. A choice question asked on its own goes throughclassify, which Jev receives as one call under the keyclassificationwith the question's instructions and categories.Execution rules
maxQuestionsandmaxInputCharacters, a hook behind every question, and a decorator's forwarded hooks. A kind or limit miss throwsUnsupportedDecisionExceptionnaming the service, the questions with their kinds, the capabilities and a remedy. When the service lacksRatingAssessment, a rating kind miss names that hook, the only one a question can need. Otherwise a kind miss names the method that backs the question:assess,PropositionAssessment,classifyorRatingAssessment. The last two checks throwIllegalStateException.classify.RatingAssessmentadds rating questions. Capabilities that claim rating questions on a service with neitherRatingAssessmentnorNativeQuestionSetExecutionthrowIllegalStateExceptionbefore any provider call.classify, with a classification request built from the question's instructions and options. A proposition goes throughPropositionAssessment, which receives the whole question, or throughassesson a service without it. A legacy service'sclassifyandassessstay directly usable.IllegalArgumentExceptionfromclassifyorrate, or from validating their answer, becomes that question'sINVALID_RESPONSE. Any other exception stops the request and propagates unchanged. Execution checks the interrupt flag before each question and throwsInterruptedExceptionwith the flag still set.INVALID_RESPONSE, as does an envelope that cannot be matched to questions. An answer with an unknown name is ignored and counted, because it cannot change which answer a spec question receives. A missing, wrong-kind, unreadable, out-of-domain or out-of-range answer, or a distribution that does not cover the options or levels or sum to 1, fails only its question. The response follows spec order.requireMatcheschecks its answers one for one, in order, by name, kind and options or levels. A response that doesn't match throwsIllegalStateExceptionnaming the service. Questions are compared by value, so a response to reworded questions with the same names, kinds and options is accepted.ObservedDecisionServiceforwardscapabilities(),classifyand every hook, and adds no kind. A choice question needs no hook on the decorator, because execution calls the decorator's ownclassify. It implementsDelegatingDecisionService, whosehookSourcenames the innermost service, and routing reads the hooks of that service. A decorator that lacks a hook the request is routed through fails preflight withIllegalStateExceptionnaming the hook.ObservedDecisionService.askruns the shared preflight and execution with those hooks, so a delegate's ownaskoverride is not called through it. A service customizes execution throughcapabilities()and the hooks.Execution
flowchart TD Caller["service.ask(input, spec) / ask(request)"] --> Svc{observed service?} Svc -->|"yes: embabel.ai.ask"| Pre{preflight} Svc -->|no| Pre Pre -->|kind or limit missing| UDE[UnsupportedDecisionException] Pre -->|kind without a hook, or decorator missing a hook| ISE[IllegalStateException] Pre -->|NativeQuestionSetExecution| Native["askNative (ask_native)"] Pre -->|no native hook| Split["assess / classify / rate per question"] Native --> Asm[DecisionResponseAssembler] Asm --> Resp[DecisionResponse in spec order] Split --> RespCompatibility
Every member added to
DecisionServicehas a default body, and no abstract member is added to an existing interface. The hook interfaces declare no default methods, so a class can implement all three. ExistingDecisionServiceandClassificationServiceimplementors compile and link unchanged. Throughask, a legacy decision service answers propositions with oneassesscall each and choices with oneclassifycall each.ClassificationServiceis unchanged.A failed prompted
classifyorassessnow logs one WARN line in place of the old DEBUG line.The prompted path calls the chat model with no LLM request event, so the run, agent and action tags reach its model call through the parent spans.
Experimental status
Every type added here carries
@ApiStatus.Experimental. The owner is James Dunnam (@jimador). Promotion to stable needs use ofDecisionService.askby a consumer application, the four families working throughDecisionService, a compatibility review ofDecisionService, privacy checks on logs, tags and exception messages, a runnable consumer proof, and a recorded run against the hosted Jev service. Promotion is revisited at the next release review after the consumer proof.Docs: a new reference page,
decision-execution, covers asking, capabilities, routing, preflight, answer rules, service families, observations and diagnostics. The TypeSafe page covers running a decision spec.Stacked on #2075. #2076 adds selecting, binding and registering services on top.