Skip to content

Feature: run decision specs against a service with Jev and LLM providers - #2080

Open
jimador wants to merge 10 commits into
feature/decision-spec-contractsfrom
feature/decision-execution
Open

jimador wants to merge 10 commits into
feature/decision-spec-contractsfrom
feature/decision-execution

Conversation

@jimador

@jimador jimador commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

This PR runs decision specs against a decision service. DecisionService.ask takes an input and a whole spec, or a DecisionRequest, and returns one answer per question in spec order. The service answers every question in one provider operation when it has a native hook, and one question per call otherwise. Four service families implement it: TypeSafe Jev, prompted LLM, no-op and stub. Every ask on an observed service produces one observation with bounded tags, one event per answer, and log lines that name the cause and the fix. Jev and prompted services are observed as built.

Routing

A service that implements NativeQuestionSetExecution answers the whole spec in one call. Any other service answers one question per call in spec order: propositions through PropositionAssessment (else assess), choices through classify and ratings through RatingAssessment. A choice question becomes ClassificationRequest.of(input, ClassificationSpec.of(question)), so the provider sees the question's own instructions and categories. A ClassificationRequest is also a DecisionRequest, so ask(request) accepts one and response.answer(spec.getQuestion()) returns the ClassificationResult.

A stub answers the whole spec in one native call by default. After perQuestion() it answers one question per call:

var spec = DecisionSpec.builder()
        .proposition("is_urgent", question -> question.asking("Does this convey urgency?"))
        .proposition("is_refund", question -> question.asking("Is this a refund request?"))
        .build();

var stub = StubDecisionService.builder("triage-stub")
        .proposition("is_urgent", new PropositionResult.Answered(true, PROVENANCE))
        .proposition("is_refund", new PropositionResult.Answered(false, PROVENANCE))
        .perQuestion()
        .build();

var response = stub.ask("The export API returns 500 since this morning.", spec);

var refund = (PropositionQuestionSpec) spec.question("is_refund");
assertEquals(new PropositionResult.Answered(false, PROVENANCE), response.answer(refund));
assertEquals(List.of("assess", "assess"), stub.calls());

A question kind the service's capabilities leave out throws UnsupportedDecisionException before any provider call. The message says whether the service lacks the hook or only leaves the kind out of capabilities(). Only rating questions need a hook. Here the service reports propositions only:

Decision service 'legacy-stub' cannot run this request: the service's capabilities leave out CHOICE questions: 'department' (CHOICE), which the service backs with classify. Questions: 'department' (CHOICE). Service capabilities: kinds [PROPOSITION], maxQuestions not reported, maxInputCharacters not reported. Report CHOICE in the service's capabilities(), use a service whose capabilities include CHOICE, or remove these questions.

capabilities() reports what a service accepts:

DecisionCapabilities capabilities = service.capabilities();
Set<QuestionKind> kinds = capabilities.getQuestionKinds();
Integer maxQuestions = capabilities.getMaxQuestions();   // null when the service reports no limit

Diagnosing

A failed provider call logs one WARN line on the provider's class logger, and the request it belongs to logs one on DecisionExecution. Neither holds input, instructions, option text or provider text:

TypeSafe question set failed with reason UNAVAILABLE: service=jev-latest, provider=TypeSafe, operation=ask_native, cause=http_5xx, status=5xx, attempts=1, elapsedMs=7, exception=SafeHttpFailure, requestId=none
Decision call failed: service=gpt-4.1-mini, provider=OpenAI, operation=ask, reason=UNAVAILABLE, cause=timeout, httpStatus=none, attempts=3, elapsedMs=12, exception=SocketTimeoutException. Check the model's availability, credentials and retry settings under embabel.agent.platform.decisions.llm.services.llm-review.
Decision ask failed: service=native-service, provider=native-provider, requestFailure=UNAVAILABLE, elapsedMs=2. The provider's own log lines give the cause.

A partial response on DecisionExecution, a native provider's unusable answers on DecisionResponseAssembler, and an answer outside its question when questions go one per call:

Decision ask partially failed: service=log-service, provider=log-provider, failedQuestions=['urgent' UNAVAILABLE], elapsedMs=3. The other answers are usable. The provider's own log lines give the cause.
Decision service 'jev-latest' returned answers it could not use: is_urgent=WRONG_KIND, department=OUT_OF_DOMAIN, frustration=OUT_OF_RANGE, channel=MISSING; unexpected answers: 1
Decision answer anomaly: service=log-service, question='team', anomaly=OUT_OF_DOMAIN. The answer does not fit the question's options or levels and is recorded as INVALID_RESPONSE. Check the service's mapping of provider output to the question.

DEBUG on com.embabel.common.ai.decision.spi.DecisionExecution logs the start of each ask and each question's outcome:

Decision ask started: service=log-service, provider=log-provider, questions=3
Decision ask completed: service=log-service, provider=log-provider, answers=['urgent' PROPOSITION answered, 'team' CHOICE selected, 'anger' RATING answered], elapsedMs=4

TRACE logs request and response content only when DecisionContentCapture.enable() has been called. Captured lines hold the input and provider output.

Observations:

Name Recorded Tags Events
embabel.ai.ask Once per ask on an observed service operation=ask, outcome, service, provider, question_count One answer.<kind>.<outcome> per question
embabel.ai.decision Once per askNative, assess or rate call operation, outcome None
embabel.ai.classification Once per classify call, including a choice question asked on its own operation=classify, outcome None
Tag Values
operation ask on embabel.ai.ask. ask_native, assess, rate or classify on provider calls. A per-question choice records classify.
outcome on embabel.ai.ask complete, partial, request_failure, unsupported, exception, interrupted, cancelled
question_count 1, 2-4, 5-16, 17+
service, provider The answering service's own name and provider, such as jev-latest and TypeSafe
<kind> in an answer event proposition, choice, rating
<outcome> in an answer event answered, selected, no_match, inconclusive, failure

Provider-call spans are children of the ask span. A meter handler counts each answer event as embabel.ai.ask.answer.<kind>.<outcome> with the ask's tags. A tracing handler records it as a span event named <question name> <kind> <outcome>. Question names, roles, input and model output are absent from tag values. Service names come from configuration and code, which keeps service bounded.

Service families

Family Service Answers Evidence Confidence
Jev TypeSafeDecisionService The whole spec in one call pTrue for propositions. Category and confidence for choices. Score as expected level index and a level distribution for ratings, with no selected level. Reported for choices and ratings
Prompted LLM LlmDecisionService The whole spec in one call A verdict, a category id or a level id None
No-op NoOpDecisionService The whole spec in one call A request failure UNAVAILABLE, so every question is Failure(UNAVAILABLE). Provider none, no provider call None
Stub StubDecisionService The whole spec in one call, or one question per call with perQuestion(), a choice through classify The scripted outcomes. A native ask throws IllegalStateException for a scripted outcome outside its question's options or levels. As scripted

All four support proposition, choice and rating questions and report no limits. A prompted service asks the whole spec in one chat model call and reads one JSON object whose answers array holds one answer per question. Both providers send questions under the keys q1 to qN in spec order, so question names stay with the caller. A choice question asked on its own goes through classify, which Jev receives as one call under the key classification with the question's instructions and categories.

Execution rules

  • Preflight checks the whole request before any provider call, in this order: question kinds, maxQuestions and maxInputCharacters, a hook behind every question, and a decorator's forwarded hooks. A kind or limit miss throws UnsupportedDecisionException naming the service, the questions with their kinds, the capabilities and a remedy. When the service lacks RatingAssessment, a rating kind miss names that hook, the only one a question can need. Otherwise a kind miss names the method that backs the question: assess, PropositionAssessment, classify or RatingAssessment. The last two checks throw IllegalStateException.
  • Default capabilities derive from the hooks a service implements. Every service accepts proposition and choice questions, since every decision service can classify. RatingAssessment adds rating questions. Capabilities that claim rating questions on a service with neither RatingAssessment nor NativeQuestionSetExecution throw IllegalStateException before any provider call.
  • A choice question asked on its own goes through classify, with a classification request built from the question's instructions and options. A proposition goes through PropositionAssessment, which receives the whole question, or through assess on a service without it. A legacy service's classify and assess stay directly usable.
  • When questions go one per call, a typed failure is recorded and the next question is asked. An IllegalArgumentException from classify or rate, or from validating their answer, becomes that question's INVALID_RESPONSE. Any other exception stops the request and propagates unchanged. Execution checks the interrupt flag before each question and throws InterruptedException with the flag still set.
  • Wire answers follow one set of rules for every native provider. A name repeated in any form fails the whole request with INVALID_RESPONSE, as does an envelope that cannot be matched to questions. An answer with an unknown name is ignored and counted, because it cannot change which answer a spec question receives. A missing, wrong-kind, unreadable, out-of-domain or out-of-range answer, or a distribution that does not cover the options or levels or sum to 1, fails only its question. The response follows spec order.
  • A native response must answer the request's spec: requireMatches checks its answers one for one, in order, by name, kind and options or levels. A response that doesn't match throws IllegalStateException naming the service. Questions are compared by value, so a response to reworded questions with the same names, kinds and options is accepted.
  • ObservedDecisionService forwards capabilities(), classify and every hook, and adds no kind. A choice question needs no hook on the decorator, because execution calls the decorator's own classify. It implements DelegatingDecisionService, whose hookSource names the innermost service, and routing reads the hooks of that service. A decorator that lacks a hook the request is routed through fails preflight with IllegalStateException naming the hook. ObservedDecisionService.ask runs the shared preflight and execution with those hooks, so a delegate's own ask override is not called through it. A service customizes execution through capabilities() and the hooks.

Execution

flowchart TD
    Caller["service.ask(input, spec) / ask(request)"] --> Svc{observed service?}
    Svc -->|"yes: embabel.ai.ask"| Pre{preflight}
    Svc -->|no| Pre
    Pre -->|kind or limit missing| UDE[UnsupportedDecisionException]
    Pre -->|kind without a hook, or decorator missing a hook| ISE[IllegalStateException]
    Pre -->|NativeQuestionSetExecution| Native["askNative (ask_native)"]
    Pre -->|no native hook| Split["assess / classify / rate per question"]
    Native --> Asm[DecisionResponseAssembler]
    Asm --> Resp[DecisionResponse in spec order]
    Split --> Resp
Loading

Compatibility

Every member added to DecisionService has a default body, and no abstract member is added to an existing interface. The hook interfaces declare no default methods, so a class can implement all three. Existing DecisionService and ClassificationService implementors compile and link unchanged. Through ask, a legacy decision service answers propositions with one assess call each and choices with one classify call each. ClassificationService is unchanged.

A failed prompted classify or assess now logs one WARN line in place of the old DEBUG line.

The prompted path calls the chat model with no LLM request event, so the run, agent and action tags reach its model call through the parent spans.

Experimental status

Every type added here carries @ApiStatus.Experimental. The owner is James Dunnam (@jimador). Promotion to stable needs use of DecisionService.ask by a consumer application, the four families working through DecisionService, a compatibility review of DecisionService, privacy checks on logs, tags and exception messages, a runnable consumer proof, and a recorded run against the hosted Jev service. Promotion is revisited at the next release review after the consumer proof.

Docs: a new reference page, decision-execution, covers asking, capabilities, routing, preflight, answer rules, service families, observations and diagnostics. The TypeSafe page covers running a decision spec.

Stacked on #2075. #2076 adds selecting, binding and registering services on top.

@jimador
jimador added this pull request to stack #2081 September 27, 2026 22:11
@jimador jimador changed the title feature/decision execution Feature: run decision specs against a service with Jev and LLM providers Sep 27, 2026
@jimador
jimador marked this pull request as ready for review September 27, 2026 22:11
@igordayen

Copy link
Copy Markdown
Contributor

@jimador - please consider breaking PR into 2 PRs, thank you

@jimador
jimador force-pushed the feature/decision-execution branch 4 times, most recently from 6055231 to 80d32fb Compare September 28, 2026 04:26
@jimador
jimador force-pushed the feature/decision-execution branch from 86e5181 to b3c6d84 Compare September 28, 2026 05:19
@sonarqubecloud

Copy link
Copy Markdown

@jimador
jimador requested a review from igordayen September 28, 2026 06:14
DecisionService.ask checks the whole request against the service's
capabilities and hooks, then answers it through askNative or one
assess, choose or rate call per question. DecisionResponseAssembler
applies one set of answer rules for native providers.
ObservedDecisionService observes each ask and its provider calls.
NoOpDecisionService and StubDecisionService serve disabled
configuration and tests.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
TypeSafeDecisionService answers a whole spec in one Jev systemOne call
and maps each Jev answer to its question. LlmDecisionService prompts
the chat model with the whole question set and reads a JSON array of
answers. Both implement the one-question hooks for direct calls.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Covers asking, routing by hook, preflight, decorators, answer rules,
the service families, observations, diagnostics and content capture.
The TypeSafe page describes how a TypeSafe service answers a spec.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
…y value

Specs no longer carry ids, so execution and telemetry compare the
response with the request's spec through requireMatches.

- A native response that does not match the request's spec is still an
  IllegalStateException naming the service. Its message now says what
  differs.
- Answer events are recorded only for a response that matches the
  request's spec.
- Tests cover a native answer whose options differ from the question's.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
…lpers

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
A choice question is a classification request built from the question,
and every decision service can classify. Per-question execution now
answers a choice through classify, so ChoiceAssessment and every choose
implementation are gone.

- Default capabilities are PROPOSITION and CHOICE; RATING still needs
  RatingAssessment. Preflight only asks for a hook on rating questions.
- A per-question choice records embabel.ai.classification with
  operation=classify.
- The stub's classify returns the choice scripted for the question name,
  then the outcome scripted for the category ids.
- Tests cover a ClassificationRequest asked through ask on native and
  per-question services.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Update routing, capability derivation, preflight and telemetry tables
for choice questions answered through classify.

Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
Signed-off-by: James Dunnam <7660553+jimador@users.noreply.github.com>
@jimador
jimador force-pushed the feature/decision-execution branch from b3c6d84 to 4bee754 Compare September 28, 2026 16:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants