-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathtutorial.html
More file actions
232 lines (227 loc) · 12.7 KB
/
Copy pathtutorial.html
File metadata and controls
232 lines (227 loc) · 12.7 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
<h1 id="a-socratic-ai-tutor-for-programming-classes">A Socratic AI tutor
for programming classes</h1>
<h2 id="the-problem">The problem</h2>
<p>Teaching programming to a roomful of students is, before anything
else, an attention-allocation problem. Twenty students get stuck on
twenty different things at the same moment, and the instructor can only
be in one place at once. The pressure to relieve that by handing out the
worked solutions is constant — but doing so quietly undermines the thing
you were trying to teach. Reading the solution feels like progress; it
tells the student <em>what</em> the answer is without forcing them
through the discomfort that builds the skill. Most students, given the
choice between struggling for twenty minutes and copying a working
snippet, copy. They are not wrong to: the immediate reward is real. But
the productive struggle is the part that sticks.</p>
<p>The other tools available to a student in 2026 push in the same
direction. A general-purpose chat assistant, asked "how do I solve
this?", will gladly produce a clean, runnable solution — often within
seconds. From the student's point of view, that is indistinguishable
from getting the answer from the instructor. The fact that an AI happens
to be the one handing it over does not change what was lost.</p>
<p>What would help, instead, is a tutor that has seen the reference
solution but is constrained to never give it away. A tutor that asks
where you are stuck, points at the missing concept, and nudges you one
step further, but stops short of completing the task. That is what this
project is.</p>
<h2 id="the-idea">The idea</h2>
<p>The tutor is a <strong>skill</strong> in Posit Assistant, backed by a
small <strong>MCP server</strong> you run.</p>
<p>When a student selects <em>Tutor</em> from the agent dropdown and
writes "I'm stuck":</p>
<ol type="1">
<li>The skill reads a task identifier the student has marked in their
notebook (e.g. <code>#| task: r-lab-1</code>).</li>
<li>It calls <code>get_task_context</code> on your server, which looks
up the reference solution and returns it as private model context —
along with an escalation level.</li>
<li>The model replies with a diagnosis, a question, or a small scaffold,
calibrated to that level.</li>
<li>The request is logged, so you can see where the class is stuck.</li>
</ol>
<p>The student never sees the reference solution. The model uses it only
as ground truth for diagnosing what is missing and deciding how strong a
hint to give.</p>
<h2 id="what-it-looks-like">What it looks like</h2>
<blockquote>
<p><strong>Student:</strong> <em>my model won't converge — what am I
doing wrong?</em></p>
<p><strong>Tutor:</strong> <em>Before we look at the model, what
assumption is your formula making about which variables are part of the
network and which are node attributes? Try running
<code>summary()</code> on your network object first and tell me what you
see.</em></p>
<p><strong>Student:</strong> <em>(pastes summary output)</em></p>
<p><strong>Tutor:</strong> <em>Good — notice that your tie variable is
being read as a node attribute, not an edge. Your formula is then asking
the estimator to predict something that isn't in the data. Look at the
line where you build the network object; one argument is in the wrong
position.</em></p>
</blockquote>
<h2 id="the-escalation-ladder-and-why-it-needs-a-server">The escalation
ladder, and why it needs a server</h2>
<p>The tutor's responses get more concrete the longer a student is
stuck: a guiding question first, then a scaffold, then a one-or-two-line
snippet. The obvious way to implement this is to count turns in the
conversation — which is what v0.1 did, and it has a hole you can drive a
truck through. Open a new chat, and you're back at rung one.</p>
<p>Because the level is computed in Postgres from the request history,
it survives fresh chat windows, restarts, and reinstalls. A student on
their fourth ask about <code>r-lab-3</code> gets a fourth-ask answer
even if this particular chat is thirty seconds old.</p>
<p>That is the one thing no purely client-side design can do, and it is
most of the reason the server exists. The other reason is the
dashboard.</p>
<h2 id="seeing-where-the-class-is-stuck">Seeing where the class is
stuck</h2>
<p>Every request is logged: student, task, level, question text,
timestamp. A Quarto report renders that into the things you actually
want to know.</p>
<p>The most useful panel is not "requests per task" but <strong>the
share of students who reached level 3 or higher</strong>. A task with
many requests but mostly level 1 is producing quick clarifications —
probably a wording problem in the prompt. A task where half the class
reaches level 3 is a task where people are genuinely stuck, and that is
the one to rewrite.</p>
<p>The second most useful panel is the raw text of what students typed.
Counts tell you <em>where</em>; the wording tells you <em>why</em>.</p>
<p>Two honest caveats. It measures asking, not struggling — a task with
zero requests is ambiguous between "everyone got it" and "nobody tried
it", so pair it with submission data before you act. And with thirty
students across forty tasks the cells are thin; treat the top few tasks
as signal and the rest as noise.</p>
<h2 id="how-it-works-under-the-hood">How it works under the hood</h2>
<ul>
<li><strong>Task ID detection.</strong> The skill reads the
<code>#| task: <id></code> marker from the active editor, which
Posit Assistant attaches as context by default. In v0.1 this was a regex
scan in extension code; now it is a sentence of prose in a markdown
file.</li>
<li><strong>Solution storage.</strong> A loader script parses your
private solution notebooks and upserts them into Postgres. Solutions
never live in the public repo or in the deployment, and <strong>updating
them needs no redeploy</strong>.</li>
<li><strong>Solution extraction.</strong> Inside a notebook, the parser
finds the heading whose text contains <code>{r-lab-1}</code>, then scans
forward for the next Quarto callout titled <code>"Solution"</code> and
captures its body, respecting nested fenced divs. This code is moved
verbatim from v0.1 — it never depended on VS Code, so the notebook
format you have already authored against is unchanged.</li>
<li><strong>The skill.</strong> The pedagogy — diagnose before
answering, prefer questions to scaffolds and scaffolds to snippets,
escalate, never reproduce the solution — lives in <code>SKILL.md</code>
in the lab repo. It follows the Agent Skills spec, so the same file
works in Claude Code and other compliant assistants.</li>
<li><strong>Tool restriction.</strong> The <code>Tutor</code> agent's
<code>tools:</code> list omits editing and code execution. In v0.1,
"don't write the solution into their file" was a request to the model.
Now it is a capability the tutor doesn't have.</li>
</ul>
<h2 id="setting-it-up-for-your-course">Setting it up for your
course</h2>
<h3 id="step-1--prepare-your-solutions-repo">Step 1 — Prepare your
solutions repo</h3>
<p>A private repo with one Quarto file per lesson. The structure, which
<code>templates/solution-template.qmd</code> demonstrates:</p>
<pre class="markdown"><code># Sum the even numbers in a vector `{r-lab-1}`
**To-do:** Write a function `sum_even(x)` that returns the sum of all even numbers.
::: {.callout-caution collapse="true" title="Solution"}
```r
sum_even <- function(x) sum(x[x %% 2 == 0])
```
**Key points:**
- `x %% 2 == 0` produces a logical vector marking even entries.
:::</code></pre>
<p>Two requirements, one of which the v0.1 docs got wrong:</p>
<ul>
<li>The heading must contain the task ID inside <strong>bare
braces</strong>: <code>`{r-lab-1}`</code>. The parser matches the
literal <code>{r-lab-1}</code> including the braces. The old tutorial
claimed a Quarto anchor like <code>{#sec-r-lab-1 .task}</code> would
also work <em>because the anchor contains the bare token</em> — it does
not, and it never did. The shipped template always used the correct
form, so notebooks written from it are fine.</li>
<li>The solution must be wrapped in a callout titled
<code>"Solution"</code>. Any callout type works; only the title
matters.</li>
</ul>
<h3 id="step-2--deploy">Step 2 — Deploy</h3>
<p>Railway project → add Postgres → add a service from this repo with
<strong>Root Directory</strong> <code>server</code>. Set
<code>DATABASE_URL</code> to the Postgres reference. Then from
<code>server/</code>:</p>
<pre class="bash"><code>npm install
npm run migrate # create the tables
npm run load -- ../../my-solutions/notebook-solutions # load content</code></pre>
<p><code>curl https://<app>.up.railway.app/healthz</code> should
return <code>{"ok":true}</code>.</p>
<h3 id="step-3--wire-up-the-lab-repo">Step 3 — Wire up the lab repo</h3>
<p>Copy <code>templates/lab-repo/.posit/</code> into the repo your
students clone and put your Railway URL in <code>settings.json</code>.
Copy <code>templates/lab-repo/README.md</code> too, and fill in the
disclosure section.</p>
<h3 id="step-4--tell-students-two-things">Step 4 — Tell students two
things</h3>
<p>Set <code>TUTOR_TOKEN</code> and <code>TUTOR_STUDENT</code> in
<code>~/.Renviron</code>. Mark tasks with <code>#| task:</code>.</p>
<p>That is the whole setup. No extension to install, no GitHub token on
any student machine.</p>
<h2 id="customizing-the-tutor-for-your-domain">Customizing the tutor for
your domain</h2>
<p>The default skill is deliberately generic — it says "programming
exercises" and avoids naming a language. For your course you will want
to name the language, mention the libraries students should reach for
first, and swap the Socratic example questions for ones in your domain's
vocabulary.</p>
<p>The difference from v0.1 is that this is now a markdown file in the
lab repo. Edit it, commit, students pull. In v0.1 the same change meant
editing the prompt, running <code>vsce package</code>, uploading a
release, and asking twenty students to reinstall — which in practice
meant the prompt was written once and never tuned. Tuning it against
real student interactions is the whole game, so lowering that cost
matters more than it sounds.</p>
<h2 id="honest-caveats">Honest caveats</h2>
<ul>
<li><p><strong>Students can read the skill.</strong> It is a file in
their open project. They can read the escalation policy, and they can
edit or delete it. The <code>.vsix</code> was extractable too, but
"unzip and rebuild an extension" and "open a file that's already in your
editor" are different levels of friction. There is no technical fix.
Some instructors will prefer to show it to students deliberately and
talk about why the constraint is there — in a course where students will
use AI professionally, that is a lesson rather than a leak.</p></li>
<li><p><strong>It doesn't stop the plain Agent.</strong> The dropdown
that offers <em>Tutor</em> also offers <em>Agent</em>, which will
happily write the function. The tutor is now an opt-in mode, not a
gatekeeper. A project-level <code>permission</code> block denying
<code>edit</code> in the lab repo makes the tutor the path of least
resistance, but it is a speed bump — the setting is a text file students
can change. Past that, this is a course-design question about what you
grade, not a tooling one.</p></li>
<li><p><strong>The class token is shared, and the payload contains the
solution.</strong> Together those mean a determined student can extract
every solution via the API. Rate limiting, rotating the token each
semester, and the dashboard's visibility all raise the cost, but none of
them close it. If that trade is wrong for you, strip fenced code from
the payload in the loader and keep the prose — the
<code>**Key points:**</code> sections exist for exactly that, and it is
about ten lines.</p></li>
<li><p><strong>Tool results may be visible.</strong> Depending on your
Positron version, a student may be able to expand a tool call in the
transcript and read what came back. Check this. If they can, the
reference solution is one click away — weaker than v0.1, where it lived
in an invisible prompt.</p></li>
<li><p><strong>Identity is self-reported.</strong>
<code>TUTOR_STUDENT</code> is a label, not authentication. It is fine
for "which tasks is the class finding hard" and unfit for anything
touching a grade.</p></li>
<li><p><strong>This is not a replacement for office hours.</strong> It
is a partial substitute for "I can't be everywhere at once" — not for
the deeper guidance that comes from a human reading what a student has
been struggling with for a week.</p></li>
</ul>
<h2 id="try-it">Try it</h2>
<p>The code is at <a
href="https://github.com/benrosche/socratic-tutor">github.com/benrosche/socratic-tutor</a>.
If you adopt it for your course, I'd be interested to hear how it goes —
both the wins and the places where the model's hint quality breaks
down.</p>