Skip to content

Improve tutor syllabus quality: single-pass generation, deterministic references - #194

Merged
noor-lpi merged 3 commits into
mainfrom
fix/tutor-syllabus-quality
Sep 30, 2026
Merged

noor-lpi merged 3 commits into
mainfrom
fix/tutor-syllabus-quality

Conversation

@noor-lpi

@noor-lpi noor-lpi commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Improves the current tutor (POST /tutor/syllabus and /tutor/syllabus/feedback). The newer tutor_test version is untouched.

  • References bug fixed: extract_doc_info read the Qdrant payload (a dict) with getattr, so the LLM got every WeLearn resource with an empty title, URL and content. This is why it made up links.
  • One structured LLM call instead of the 3-agent chain. The LLM returns a SyllabusDraft (Pydantic), which is rendered to markdown in tutor/service/syllabus.py:
    • fixed fr/en section headings (FR: Résultats d'apprentissage, RA1…);
    • the numbers of objectives, outcomes and competencies are set from the course duration (table based on teaching-centre and ECTS guidance);
    • GreenComp: at least one per syllabus, each linked to outcomes, with its official name (French names from the JRC's French edition); unlinked or duplicate codes are dropped;
    • class plans are student-centred and active-learning.
  • References are built in code from the documents in the request (the selected resources, or all retrieved ones), each listed once. The LLM never writes them.
  • Feedback:
    • the syllabus and template are now passed as text, not Python object dumps;
    • references are split off before the LLM and re-attached unchanged;
    • LLM comments before the first heading and after the schedule are trimmed;
    • the syllabus is resized when the feedback changes the duration;
    • calls are now traced in LangSmith as Tutor (FeedbackAgent).
  • Other fixes:
    • data collection stored the first draft instead of the final syllabus;
    • stray "weather…" text in the GreenComp prompt;
    • the disciplinary skills were joined incorrectly in the prompt;
    • invalid _target=blank in the template;
    • the no-op @with_backoff on the two syllabus endpoints is removed.
  • Generation takes about 18–25 s instead of over 60 s.

No frontend change needed: the response keeps source="PedagogicalEngineerAgent".

Revert switch

TUTOR_SINGLE_PASS (default true). Set it to false to run the old 3-agent chain. It still gets the deterministic references and the bug fixes.

Test plan

  • pytest src/app/tests: 260 pass (new: tests/services/tutor/test_syllabus.py; test_extract_doc_info fixed to use dict payloads)
  • Live Mistral runs (fr): sections, counts, GreenComp names, references
  • Manual GUI tests on the old tutor page: fr/en, with and without selected resources, feedback rounds
  • LangSmith: check that Tutor (SinglePassSyllabus) and Tutor (FeedbackAgent) runs appear on staging
  • Compare a few syllabi, old chain vs single pass, before prod

🤖 Generated with Claude Code

Noor A and others added 2 commits September 30, 2026 14:46
… references

- Fix extract_doc_info reading Qdrant payload dicts with getattr, which sent
  every WeLearn resource to the LLM with empty title/url/content
- Replace the 3-agent chain with one structured LLM call (SyllabusDraft)
  rendered to markdown in code: fixed fr/en headings, counts sized by course
  duration, GreenComp competencies with official names linked to outcomes,
  student-centred class plans; old chain kept behind TUTOR_SINGLE_PASS=false
- Build the references section from the selected documents, never the LLM
- Feedback: pass syllabus/template as text, keep references verbatim, trim
  LLM chatter, rescale on duration change, trace in LangSmith (FeedbackAgent)
- Store the final syllabus (not the first draft) in data collection
- Remove stray GreenComp prompt text, fix disciplinary skills join and
  template link attribute

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@noor-lpi noor-lpi self-assigned this Sep 30, 2026
Comment thread src/app/core/config.py
LLM_MODEL_NAME: str
LLM_TEMPERATURE: float
# old tutor: one structured LLM call instead of the 3-agent chain (False = revert)
TUTOR_SINGLE_PASS: bool = True

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this needs to be added to the values.yaml in the k8s folder

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

"Syllabus draft has no GreenComp competency linked to an outcome"
)
greencomp = greencomp[: limits.competencies]
competencies = greencomp + others[: limits.competencies - len(greencomp)]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

to be clear, we make the choice to keep greencomp and ditch others if the nb of greencomp is equal limits.competencies

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed, and added a code comment. Done in 8600ba8

Comment thread src/app/tutor/service/syllabus.py Outdated
Comment thread src/app/tutor/service/syllabus.py Outdated
Comment on lines +204 to +260
def render_markdown(draft: SyllabusDraft, lang: str, limits: Limits) -> str:
h = HEADINGS.get(lang, HEADINGS["en"])
objectives = draft.objectives[: limits.objectives]
outcomes = draft.outcomes[: limits.outcomes]
n_obj, n_lo, lo = len(objectives), len(outcomes), h["lo"]

greencomp, others, seen_codes = [], [], set()
for comp in draft.competencies:
linked = _refs(lo, comp.outcome_numbers, n_lo)
suffix = f" *({linked})*" if linked else ""
code = (comp.greencomp_code or "").strip()
if code not in GREENCOMP_NAMES:
others.append(f"- {comp.text}{suffix}")
elif linked and code not in seen_codes:
# GreenComp only when it supports a listed outcome, each code once
seen_codes.add(code)
name = GREENCOMP_NAMES[code].get(lang, GREENCOMP_NAMES[code]["en"])
greencomp.append(f"- **GreenComp {code} – {name}** : {comp.text}{suffix}")
if not greencomp:
logger.warning(
"Syllabus draft has no GreenComp competency linked to an outcome"
)
greencomp = greencomp[: limits.competencies]
competencies = greencomp + others[: limits.competencies - len(greencomp)]

sections = [f"# {draft.course_title}", f"## {h['description']}", draft.description]
sections += [
f"## {h['objectives']}",
"\n".join(f"{i}. {o}" for i, o in enumerate(objectives, 1)),
]
sections += [
f"## {h['outcomes']}",
"\n".join(
f"- **{lo}{i}** {o.text}"
+ (
f" *({h['objectives_ref']} {refs})*"
if (refs := _refs("", o.objective_numbers, n_obj))
else ""
)
for i, o in enumerate(outcomes, 1)
),
]
sections += [f"## {h['competencies']}", "\n".join(competencies)]
sections += [
f"## {h['assessment']}",
"\n".join(
f"- **{a.method}** ({a.weight})"
+ (f" — {refs}" if (refs := _refs(lo, a.outcome_numbers, n_lo)) else "")
for a in draft.assessment
),
]
table = [h["table"], "|---|---|---|---|"] + [
f"| {i} | {_cell(s.topics)} | {_refs(lo, s.outcome_numbers, n_lo)} | {_cell(s.class_plan)} |"
for i, s in enumerate(draft.schedule, 1)
]
sections += [f"## {h['schedule']}", "\n".join(table)]
return "\n\n".join(sections)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Si ça permet de faire la conversion entre objet sérialisé et markdown il me semble judicieux de déléguer ça à une lib standard comme pandoc peut être ?

https://stackoverflow.com/questions/44768989/how-to-convert-json-object-to-markdown-using-pypandoc-without-writing-to-file

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pandoc sert à convertir un document d'un format à un autre (par ex. Word → markdown), mais il ne sait pas construire notre syllabus à partir des données. On devrait quand même écrire tout le formatting nous-mêmes, avec en plus un outil de plus à installer et à maintenir. Vu que le rendu actuel fait ~50 lignes, je propose de le garder tel quel.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Selon moi, le parseur proposé est impossible à maintenir sur le temps long, il faudra trouver une solution plus stable, écrire un template jinja par exemple ou se tourner vers les lib qui permettent d'écrire du markdown directement, j'ai proposé pandoc car c'est ce qui m'est passé par la tête rapidement, mais cette fonction risque à mon sens de devenir une énorme dette technique par la suite et mérite qu'on se penche sur la question.

Il n'en reste pas moins que selon moi c'est pas bloquant immédiatement, mais faudra repasser rapidement dessus

Comment thread src/app/tutor/service/syllabus.py Outdated
Comment thread src/app/tutor/service/syllabus.py Outdated


async def generate_syllabus(
message: Any,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Le "Any" me semble servir à rien, il est plus utile de trouver le vrai type pris ici

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

Comment thread src/app/tutor/service/syllabus.py Outdated
Comment thread src/app/tutor/service/syllabus.py
Comment thread src/app/tutor/service/syllabus.py
Comment thread src/app/tutor/service/models.py Outdated
)


class SyllabusDraft(BaseModel):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💅 Choisir si "Draft" est d'abord ou non pour maintenir une consistence

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

Comment thread src/app/tutor/api/router.py Outdated
Comment thread src/app/tutor/service/syllabus.py Outdated
Comment thread src/app/tutor/service/syllabus.py Outdated
{"user_prompt": build_user_prompt(message, limits, disciplinary_skills)},
config=config,
)
draft = SyllabusDraft.model_validate(result, from_attributes=True)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this could be wrapped in a try catch so if the validation fails the error is handled

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

Comment thread src/app/tutor/service/syllabus.py Outdated
Comment on lines +186 to +192
url, title = res.get("url") or "", res.get("title") or res.get("url") or ""
key = url or title
if not key or key in seen:
continue
seen.add(key)
link = f' <a href="{url}" target="_blank">[{h["link"]}]</a>' if url else ""
items.append(f"- {title}{link}")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice, there is a small issue that I see if we don't have a url the link wo=ill not work and if we don't have a title the item will look odd. we should make sure that both url and title are present

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

Comment thread src/app/tutor/service/syllabus.py Outdated
h = HEADINGS.get(lang, HEADINGS["en"])
seen, items = set(), []
for res in resources:
url, title = res.get("url") or "", res.get("title") or res.get("url") or ""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is the double check for url wanted ? title could either be a title or an url

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8600ba8

- Retry once on a malformed LLM draft; /syllabus returns 502
  SYLLABUS_GENERATION_FAILED instead of a raw 500
- References require both title and url, deduplicated by url
- Type message as MessageWithResources; move Limits to models;
  rename SyllabusDraft to DraftSyllabus
- Add TUTOR_SINGLE_PASS to k8s values.yaml
- Clarify final syllabus selection, GreenComp priority and draft
  normalisation with comments
@noor-lpi
noor-lpi merged commit 1cdcd32 into main Sep 30, 2026
3 checks passed
@noor-lpi
noor-lpi deleted the fix/tutor-syllabus-quality branch September 30, 2026 16:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants