From 467d9d1d7c6c0dd7cd0b26af0c22994916277c8b Mon Sep 17 00:00:00 2001 From: MechaCat02 Date: Wed, 16 Sep 2026 22:50:33 +0200 Subject: [PATCH] =?UTF-8?q?fix(tutor):=20what=20live=20gpt-oss-20b=20lesso?= =?UTF-8?q?ns=20showed=20=E2=80=94=20marks=20lost,=20feedback=20swallowed,?= =?UTF-8?q?=20answers=20given=20away?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three unit-1.1 lessons with gpt-oss-20b through LM Studio and the server — the first real model on the reworked turn. The letter-level check went out right, and the +25 clamp held: a reported 80 on the first answer was stored as 25. What failed was how the model wrote its blocks, a different way each session. All three transcripts are in test/fixtures/, verbatim, and each failure below is a test against them. Marks lost. The prompt asks for `여덟 | wrong | 여덜`. The first session wrote `we | wrong | 우라 → 우리`, English prompt first; the second wrote no ::result at all and marked only in prose, `✗ 나 | I (humble) → 저`. evidence.ts keys on the first field of a ::result row, so nothing was ever recorded — no evidence, no schedule, no confusions, and a 다지기 review that could never close. The artifact would have lost them the same way. domain/marking.ts attaches each mark to its word only where that is unambiguous: one Korean word first, or through a prompt of the exercise he answered, read via that exercise's ::words as the letter check reads it. With no ::result block the ✓/✗ lines are read on the same terms, so a mark can never name a word the exercise did not ask for; a mark on a whole sentence is still dropped. What he mistook a word for is taken from what he actually wrote whenever the mark itself gives no other word — the third session put the right answer there. Its third session, marked through all of this: 20 evidence rows, 20 cards. Progress on requests. The prompt allows marks, ::confirmed and ::progress only in reply to an answer. The model wrote ::progress on every message, and three requests for a new exercise took the unit from 50% to 80% with nothing answered. A reply to anything but an answer now changes none of them. Feedback swallowed. The model closed no blocks, so lib read what followed each one as rows: "your score is about 5%" became a result row the student never saw, and a "---" became a recall item he was asked to write in 한글. Another session fenced every block in ```. gloss.ts now decides every block's extent from the raw text — at "::", the next block, a rule or fence line, a blank line with no row after it, or for the piped blocks the first line without a "|" — and hands lib the blocks properly closed. The gate audit, now run through the parser the lesson uses, still flags 7 and 2. Answers given away. Translate rows came with their meanings ("나 | I") and recall hints were the answers ("two | 이"). A translate row keeps only its Korean line, and a recall hint that is the expected word, or any word the message declares, is dropped. Also: the spelling a recall prompt expects now keeps its qualifiers. With 나 "I, me (casual)" and 저 "I, me (humble)" in one list, "I (humble)" matched 나: no letter check was sent, and the mark for 저 was filed under 나. Co-Authored-By: Claude Opus 5 (1M context) --- app/src/domain/gloss.ts | 219 ++++++++++++---- app/src/domain/letters.ts | 48 +++- app/src/domain/marking.ts | 170 +++++++++++++ app/src/domain/turn.ts | 35 ++- app/src/ui/tutor/TutorTab.tsx | 14 +- test/domain/block-leak.test.ts | 58 ++++- test/domain/gate-audit.test.ts | 4 +- test/domain/learner.test.ts | 9 + test/domain/marking.test.ts | 249 +++++++++++++++++++ test/fixtures/gpt-oss-20b-fenced-blocks.json | 37 +++ test/fixtures/gpt-oss-20b-prose-marks.json | 69 +++++ test/fixtures/gpt-oss-20b-recall-round.json | 8 + 12 files changed, 850 insertions(+), 70 deletions(-) create mode 100644 app/src/domain/marking.ts create mode 100644 test/domain/marking.test.ts create mode 100644 test/fixtures/gpt-oss-20b-fenced-blocks.json create mode 100644 test/fixtures/gpt-oss-20b-prose-marks.json create mode 100644 test/fixtures/gpt-oss-20b-recall-round.json diff --git a/app/src/domain/gloss.ts b/app/src/domain/gloss.ts index cc6afac..e0d8b77 100644 --- a/app/src/domain/gloss.ts +++ b/app/src/domain/gloss.ts @@ -3,8 +3,8 @@ The 16 Sep lib fixed the two defects this module used to work around: a gloss block now closes at each "=" line, and every block is collected - rather than only the first. What remains is how a block ends, and both - remaining defects reach the student. + rather than only the first. What remains is how a block ends, and every + one of these reaches the student. ── 1. Removing one block takes the next block's colons ──────────────── @@ -25,19 +25,49 @@ ── 2. A block runs to the next "::", or to the end ──────────────────── - Every local model tested forgets the closing "::" on a gloss block. lib - then reads the paragraph that follows as gloss PARTS, so two sentences - of English render as Korean example text inside the card. And two gloss - blocks written back to back share one terminator, so lib never sees the - second. + The prompt closes every block with "::". Every local model tested leaves + it off, and lib then reads whatever follows as rows of the block: - So gloss blocks are found in the raw text too. A gloss block ends where - its rows end — a row is `한글 | role | gloss`, a translation starts with - "=" — and lib parses each block on its own, so the row format is still - entirely lib's. */ + · after a gloss block, the paragraph that follows becomes gloss PARTS, + so two sentences of English render as Korean example text; + · after ::result, the tutor's own feedback — "you wrote an extra line", + "your score for this unit is about 5%" — became result rows, so the + student never saw it; + · a markdown rule, "---", became an exercise item, and he was asked to + write the 한글 for "---". + + So every block's extent is decided here, from the raw text, and lib is + handed the blocks properly closed — the row formats stay entirely lib's. + A block ends where its rows end: at "::", at the next block, at a rule + line, at a blank line that nothing row-like follows, and — for the + blocks whose rows always carry a "|" — at the first line without one. A + gloss row is `한글 | role | gloss` and its translation starts with "=". + + ── 3. Rows that give the exercise away ──────────────────────────────── + + Measured on gpt-oss-20b: translate rows arrived with their meaning + appended ("나 | I"), and recall hints were the answer itself ("two | 이"). + A translate row is one Korean line, so a "|" tail is not part of it; a + recall hint that is the expected word, or any word the message declares, + is dropped. Neither loosens anything — the exercise is only ever made + harder to cheat. + + And in one session every block came wrapped in a markdown code fence, + which left empty "```" lines in the message. A fence line is dropped. */ import { parse } from "@lib/blocks.js"; -import type { GlossBlock, ParsedMessage } from "@lib/blocks.js"; +import type { ParsedMessage, Task, WordEntry } from "@lib/blocks.js"; +import { expectedFor } from "./letters.js"; + +/** A markdown rule: never a row, and never the start of one. */ +const RULE = /^\s*(?:-{3,}|\*{3,}|_{3,})\s*$/; + +/** A code fence. The prompt forbids them; one session wrapped every block + in one. Not prose, not a row: dropped, and it ends a block. */ +const FENCE = /^\s*```/; + +/** Blocks whose every row carries a "|". */ +const PIPED = new Set(["words", "result", "task match", "task build", "task choice"]); /** * Is this a line of ::gloss content? @@ -58,79 +88,158 @@ export function glossContent(body: string): string { return (end === -1 ? lines : lines.slice(0, end)).join("\n"); } +interface Block { + /** The opening line, as written: "::task recall", "::progress 40 | note". */ + head: string; + /** "gloss", "words", "task recall", "progress" … */ + kind: string; + lines: string[]; +} + interface Scanned { prose: string[]; - /** The content of each gloss block, in order, trimmed to its rows. */ - glosses: string[]; + blocks: Block[]; +} + +function kindOf(head: string): string { + const [name = "", sub = ""] = head.slice(2).trim().split(/\s+/); + return name === "task" ? `task ${sub}` : name; } /** * One pass over the raw text. * - * The grammar is small: a line of exactly "::" closes a block, any other - * line beginning "::" opens one (::gloss, ::task translate, ::words, - * ::progress 40 | note), and a block runs until it is closed, until the - * next block opens, or to the end — except a gloss block, which also ends - * at its first line that is not a gloss row. Trailing whitespace is common - * (one model ends every line with markdown's two spaces), so nothing here - * compares exact strings. + * A line of exactly "::" closes a block; any other line beginning "::" + * opens one (::gloss, ::task translate, ::words, ::progress 40 | note); a + * ::progress line is a whole block by itself. The rest of the rules are at + * the top of this file. Trailing whitespace is common — one model ends every + * line with markdown's two spaces — so nothing here compares exact strings. */ function scan(raw: string): Scanned { const prose: string[] = []; - const glosses: string[] = []; - let block: string | null = null; - let gloss: string[] = []; + const blocks: Block[] = []; + const lines = raw.split("\n"); + let block: Block | null = null; - const enter = (next: string | null) => { - if (block === "gloss") glosses.push(gloss.join("\n")); - block = next; - gloss = []; + const close = () => { + if (block) blocks.push(block); + block = null; }; - for (const line of raw.split("\n")) { + for (let i = 0; i < lines.length; i++) { + const line = lines[i]!; const t = line.trim(); + + if (FENCE.test(line)) { + close(); + continue; + } if (t === "::") { - enter(null); + close(); continue; } if (t.startsWith("::")) { - enter(t.slice(2).split(/\s/)[0] ?? null); + close(); + const opened: Block = { head: t, kind: kindOf(t), lines: [] }; + if (opened.kind === "progress") blocks.push(opened); + else block = opened; continue; } - // An unclosed gloss ends where its rows do; what follows is prose again. - if (block === "gloss" && !isGlossContent(line)) enter(null); - if (block === "gloss") gloss.push(line); - else if (!block) prose.push(line); + const open = block as Block | null; + if (open) { + if (open.kind === "gloss") { + // An unclosed gloss ends where its rows do; what follows is prose. + if (!isGlossContent(line)) close(); + } else if (RULE.test(line)) { + close(); + } else if (t === "") { + // A blank line inside a block is only a gap if a row comes next. + const next = lines.slice(i + 1).find((l) => l.trim() !== "")?.trim() ?? ""; + if (!next.startsWith("::") && !(next.includes("|") && !RULE.test(next))) close(); + } else if (PIPED.has(open.kind) && !t.includes("|")) { + close(); + } + } + + const current = block as Block | null; + if (current) { + if (t !== "") current.lines.push(line); + } else { + prose.push(line); + } } - enter(null); + close(); - return { prose, glosses }; + return { prose, blocks }; } /** The prose of a message: every line that is not inside a block. */ export function proseOf(raw: string): string { - return scan(raw).prose.join("\n").trim(); + return tidyProse(scan(raw).prose); } /** - * parse(), with the body and the gloss blocks read from the raw text. - * Use this everywhere instead of calling parse() directly. + * The prose, without the separators a closed block leaves behind: runs of + * blank lines, and "---" rules with nothing left to separate at either end. */ -export function parseMessage(text: string): ParsedMessage { - const parsed = parse(text); - const { prose, glosses } = scan(text); - const body = prose.join("\n").trim(); - - if (!glosses.length) return { ...parsed, body, gloss: null }; - - const blocks: GlossBlock[] = []; - for (const content of glosses) { - const one = parse(`::gloss\n${content}\n::`); - if (one.gloss) blocks.push(...one.gloss); - } - - return { ...parsed, body, gloss: blocks.length ? blocks : null }; +function tidyProse(lines: string[]): string { + const out = [...lines]; + const edge = (l: string | undefined) => l !== undefined && (l.trim() === "" || RULE.test(l)); + while (edge(out[0])) out.shift(); + while (edge(out[out.length - 1])) out.pop(); + return out.join("\n").replace(/\n{3,}/g, "\n\n"); } -export type { GlossBlock, ParsedMessage }; +/** Has at least one letter — Hangul or Latin — and so can be an item. */ +const hasLetter = (s: string) => /\p{L}/u.test(s); + +/** Section 3 of the header: rows that give the exercise away, and non-rows. */ +function tidyTask(task: Task | null, words: WordEntry[] | null): Task | null { + if (!task) return task; + const rows = task.rows?.filter(hasLetter); + switch (task.type) { + case "translate": + return { + ...task, + rows, + items: task.items.map((it) => ({ q: it.q.split("|")[0]!.trim() })).filter((it) => hasLetter(it.q)), + }; + case "recall": + return { + ...task, + rows, + items: task.items + .filter((it) => hasLetter(it.q)) + .map((it) => { + // A hint holding the answer, or any word the message declares, + // is the answer given away. Real hints are about the word: + // "double batchim". + const answer = expectedFor(it.q, words); + const hintWords = it.hint.match(/[가-힣]+/g) ?? []; + const gives = + (answer && it.hint.includes(answer)) || hintWords.some((h) => words?.some((w) => w.ko === h)); + return gives ? { ...it, hint: "" } : it; + }), + }; + default: + return { ...task, rows }; + } +} + +/** + * parse(), with every block's extent and the prose decided from the raw + * text. Use this everywhere instead of calling parse() directly. + */ +export function parseMessage(text: string): ParsedMessage { + const { prose, blocks } = scan(text); + const closed = blocks.map((b) => [b.head, ...b.lines, "::"].join("\n")).join("\n"); + const parsed = parse(closed); + return { + ...parsed, + body: tidyProse(prose), + task: tidyTask(parsed.task, parsed.words), + }; +} + +export type { ParsedMessage }; diff --git a/app/src/domain/letters.ts b/app/src/domain/letters.ts index 712a16f..f10d818 100644 --- a/app/src/domain/letters.ts +++ b/app/src/domain/letters.ts @@ -22,15 +22,55 @@ const normal = (s: string): string => .replace(/\s+/g, " ") .trim(); +/** Case and spacing only — "I (humble)" stays distinct from "I". */ +const verbatim = (s: string): string => s.toLowerCase().replace(/\s+/g, " ").trim(); + +const FILLER = new Set(["a", "an", "the", "to", "be"]); + +/** The words of a meaning, qualifiers included: "I, me (humble)" → i, me, humble. */ +const wordsOf = (s: string): Set => + new Set( + s + .toLowerCase() + .split(/[^a-z]+/) + .filter((w) => w && !FILLER.has(w)), + ); + /** - * The Korean a recall prompt is asking for: the ::words entry whose meaning - * matches the prompt exactly, else one whose meaning contains it or is - * contained by it. "" when nothing matches — and then no check is sent, which - * the prompt tells the tutor how to handle. + * The Korean a recall prompt is asking for, from the ::words entries: + * + * 1. the meaning that is the prompt word for word; + * 2. the closest meaning holding every word of the prompt, qualifiers + * included — "I (humble)" is in "I, me (humble)" and not in + * "I, me (casual)"; + * 3. the meaning that matches once qualifiers are dropped, else one that + * contains the prompt or is contained by it. + * + * "" when nothing matches — and then no check is sent, which the prompt + * tells the tutor how to handle. + * + * Measured on gpt-oss-20b: with 나 "I, me (casual)" and 저 "I, me (humble)" + * in one list, dropping the qualifiers first matched "I (humble)" to 나. The + * letter check was never sent, and the tutor's mark for 저 was filed under 나. */ export function expectedFor(prompt: string, words: WordEntry[] | null): string { const q = normal(prompt); if (!q || !words?.length) return ""; + const literal = words.find((w) => verbatim(w.gloss) === verbatim(prompt)); + if (literal) return literal.ko; + + const want = wordsOf(prompt); + let closest: WordEntry | null = null; + let extra = Infinity; + for (const w of words) { + const have = wordsOf(w.gloss); + if (want.size && [...want].every((x) => have.has(x)) && have.size - want.size < extra) { + closest = w; + extra = have.size - want.size; + } + } + if (closest) return closest.ko; + const exact = words.find((w) => normal(w.gloss) === q); if (exact) return exact.ko; const near = words.find((w) => { diff --git a/app/src/domain/marking.ts b/app/src/domain/marking.ts new file mode 100644 index 0000000..8dda9f5 --- /dev/null +++ b/app/src/domain/marking.ts @@ -0,0 +1,170 @@ +/* Reading the marks — which word each one is for. + + The prompt asks for the Korean first: `여덟 | wrong | 여덜`. The evidence, + the schedule and the confusion list all hang off that first field, and a + row naming no studiable word is dropped — the artifact's rule, and still + this app's (evidence.ts). + + Measured on gpt-oss-20b, a recall round came back marked the other way + round, English prompt first and the Korean after it: + + I | ok | 나 + we | wrong | 우라 → 우리 + + Every row was dropped, the round counter moved, and nothing was recorded. + So a row is attached to its word here, before evidence.ts sees it — but + only where that is unambiguous: + + · a first field that is one Korean word is that word, punctuation aside; + · an English first field that is one of the answered recall exercise's + prompts is the word that prompt asked for — found the way the letter + check finds it, from that exercise's ::words — or, failing that, the + Korean the row itself gives: the part after an arrow, or on a right + answer the lone word; + · anything else is left as it is. A mark on a whole sentence, "머리 아파 + | wrong", does not say which word was wrong, and is still dropped. + + In a second session the same model wrote no ::result block at all. It + marked in prose, as the prompt also asks — ✓ or ✗, the Korean, his + reading: + + ✓ 나 | I + ✗ 나 | I (humble) → 저 + + With no block, those lines are read instead, on the same terms: a line + counts only through a prompt or item of the exercise he answered, so it + can never mark a word that exercise did not ask for. + + What he mistook a word for is the third field — but in a third session + the model wrote the RIGHT answer there, `We | wrong | 우리`, and the slip + was lost. For a recall exercise the app has something better than the + model's guess: what he actually wrote, in his own answer message. It is + used whenever the mark gives no other Korean word. */ + +import type { ParsedMessage, ResultRow, Task, WordEntry } from "@lib/blocks.js"; +import { expectedFor } from "./letters.js"; + +const ARROW = /\s*(?:→|->|=>)\s*/; + +/** The one Korean word a field is — "“나”" is 나 — or "" if it is not one. */ +function soleWord(field: string): string { + const runs = field.match(/[가-힣]+/g) ?? []; + return runs.length === 1 ? runs[0]! : ""; +} + +/** "We → 우라" lines of his answer message: what he wrote, by prompt. */ +function writtenAnswers(answer: string): Map { + const out = new Map(); + for (const line of answer.split("\n")) { + const at = line.lastIndexOf(" → "); + if (at > 0) out.set(line.slice(0, at).trim().toLowerCase(), line.slice(at + 3).trim()); + } + return out; +} + +/** What he mistook `word` for: the mark's own Korean, else what he wrote. */ +function mistook(word: string, fromMark: string, prompt: string | undefined, written: Map): string { + if (fromMark && fromMark !== word) return fromMark; + const wrote = prompt ? soleWord(written.get(prompt.trim().toLowerCase()) ?? "") : ""; + return wrote && wrote !== word ? wrote : ""; +} + +/** + * `results` with each row's item made the word it marks, where that can be + * told. `answered` is the message holding the exercise he was answering, + * and `answer` his reply to it. + */ +export function readResults( + results: ResultRow[] | null, + answered: ParsedMessage | null, + answer = "", +): ResultRow[] | null { + if (!results) return results; + const prompts = answered?.task?.type === "recall" ? answered.task.items : []; + const written = writtenAnswers(answer); + + return results.map((r) => { + const [before = "", after] = r.mistakenFor.split(ARROW); + const arrowed = after !== undefined; + + if (/[가-힣]/.test(r.item)) { + // As the prompt asks. "우라 → 우리" in the third field still means he wrote 우라. + const word = soleWord(r.item); + return { + item: word || r.item, + ok: r.ok, + mistakenFor: arrowed ? soleWord(before) : r.mistakenFor, + }; + } + + const prompt = prompts.find((p) => p.q.trim().toLowerCase() === r.item.trim().toLowerCase()); + const asked = prompt ? expectedFor(prompt.q, answered?.words ?? null) : ""; + const word = asked || (arrowed ? soleWord(after) : r.ok ? soleWord(before) : ""); + if (!word) return r; + + return { item: word, ok: r.ok, mistakenFor: r.ok ? "" : mistook(word, soleWord(before), prompt?.q, written) }; + }); +} + +const MARK = /^\s*(✓|✔|✅|✗|✘|❌)\s*(.+)$/; +const RIGHT = new Set(["✓", "✔", "✅"]); + +const escape = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); + +/** `phrase` in `text` as a whole phrase, any case: "I" is not in "It". */ +const hasPhrase = (text: string, phrase: string) => + new RegExp(`(? { + const w = soleWord(s); + return w && w === s.trim() && korean.has(w) ? w : ""; + }; + switch (task.type) { + case "recall": { + // Longest prompt first: a line about "I (humble)" is not about "I". + const prompt = [...task.items] + .sort((a, b) => b.q.length - a.q.length) + .find((it) => hasPhrase(text, it.q)); + return prompt ? expectedFor(prompt.q, words) : ""; + } + case "translate": + return task.items.map((it) => single(it.q)).find(Boolean) ?? ""; + case "match": + return task.pairs.map((pair) => single(pair.ko)).find(Boolean) ?? ""; + default: + return ""; + } +} + +/** ✓/✗ lines in the prose, read as marks — only when there is no ::result. */ +function proseMarks(body: string, answered: ParsedMessage | null, answer: string): ResultRow[] | null { + const task = answered?.task; + if (!task) return null; + const written = writtenAnswers(answer); + const rows: ResultRow[] = []; + for (const line of body.split("\n")) { + const m = MARK.exec(line); + if (!m) continue; + const word = markedWord(m[2]!, task, answered.words); + if (!word) continue; + const ok = RIGHT.has(m[1]!); + const other = (m[2]!.match(/[가-힣]+/g) ?? []).find((w) => w !== word) ?? ""; + const prompt = task.type === "recall" ? task.items.find((it) => hasPhrase(m[2]!, it.q))?.q : undefined; + rows.push({ item: word, ok, mistakenFor: ok ? "" : mistook(word, other, prompt, written) }); + } + return rows.length ? rows : null; +} + +/** + * Every mark a reply carries, each attached to its word: the ::result + * block if it has one, else its ✓/✗ lines. `answered` is the message + * holding the exercise he answered, and `answer` his reply to it. + */ +export function readMarks(reply: ParsedMessage, answered: ParsedMessage | null, answer = ""): ResultRow[] | null { + return reply.results + ? readResults(reply.results, answered, answer) + : proseMarks(reply.body, answered, answer); +} diff --git a/app/src/domain/turn.ts b/app/src/domain/turn.ts index 736ffd5..e3c8ca1 100644 --- a/app/src/domain/turn.ts +++ b/app/src/domain/turn.ts @@ -21,6 +21,7 @@ import { GATE_TRIES, proseIsKorean, scanTask, taskMaterial } from "@lib/gate.js" import { editMeta } from "../db/writes.js"; import { allowedSet, assemblePrompt, courseScaffold, unitOf, type FocusMode } from "./gate.js"; import { parseMessage } from "./gloss.js"; +import { readMarks } from "./marking.js"; import { headsFor } from "./resolver.js"; import { renderTail, type Retry } from "./prompt-tail.js"; import { practicePool, practiceSet } from "./practice.js"; @@ -171,14 +172,36 @@ export interface Applied { export async function applyReply( db: Db, parsed: ParsedMessage, - { lookups, today }: { lookups: Iterable; today: number }, + { + lookups, + today, + answered = null, + answer = "", + afterAnswer = true, + }: { + lookups: Iterable; + today: number; + /** The message holding the exercise he answered — see marking.ts. */ + answered?: ParsedMessage | null; + /** His message this replies to — what he actually wrote. */ + answer?: string; + /** + * Whether the message this replies to was an answer (isExerciseAnswer). + * The prompt allows marks, confirmations and a progress line only then. + * gpt-oss-20b wrote ::progress on every message, and three requests for + * a new exercise took a unit from 50% to 80% with nothing answered — + * so a reply to anything else changes none of them. + */ + afterAnswer?: boolean; + }, ): Promise { const progress = await readProgress(db); - const report = parsed.progress - ? await applyProgressReport(db, progress, parsed.progress.score, parsed.progress.note) - : null; - const results = await applyResults(db, parsed.results, lookups, today); - const confirmation = await applyConfirmed(db, currentUnit(progress), parsed.confirmed); + const report = + afterAnswer && parsed.progress + ? await applyProgressReport(db, progress, parsed.progress.score, parsed.progress.note) + : null; + const results = afterAnswer ? await applyResults(db, readMarks(parsed, answered, answer), lookups, today) : null; + const confirmation = await applyConfirmed(db, currentUnit(progress), afterAnswer ? parsed.confirmed : null); let recent = await readRecent(db); if (parsed.task) { diff --git a/app/src/ui/tutor/TutorTab.tsx b/app/src/ui/tutor/TutorTab.tsx index 21dbcae..78dd3f7 100644 --- a/app/src/ui/tutor/TutorTab.tsx +++ b/app/src/ui/tutor/TutorTab.tsx @@ -41,7 +41,7 @@ import { type FlatUnit, } from "../../domain/gate.js"; import type { FocusMode } from "../../domain/gate.js"; -import { currentUnit, noteAnswer } from "../../domain/progress.js"; +import { currentUnit, isExerciseAnswer, noteAnswer } from "../../domain/progress.js"; import { makeStubTutor, SampleError, type Sample, type StubWord } from "../../domain/stub-tutor.js"; import { makeRemoteTutor } from "../../domain/tutor-client.js"; import { koreanTokens, lookupMany } from "../../domain/lexicon.js"; @@ -377,6 +377,10 @@ export function TutorTab() { // What he looked up belongs to THIS answer; a typed message has none. lastLookups.current = lookups; + // The exercise he is answering, for reading the marks that come back. + const answeredTurn = [...turns].reverse().find((t) => t.role === "assistant"); + const answered = answeredTurn ? parseMessage(answeredTurn.body) : null; + // Answering scrolled the log up to the fields; sending means he wants // the reply, so follow it again. stick.current = true; @@ -442,7 +446,13 @@ export function TutorTab() { setStrays(flagged); if (replyId && flagged.length) setFlags((f) => ({ ...f, [replyId]: flagged })); - const applied = await applyReply(db, result.parsed, { lookups: lastLookups.current, today }); + const applied = await applyReply(db, result.parsed, { + lookups: lastLookups.current, + today, + answered, + answer: body, + afterAnswer: isExerciseAnswer(body), + }); // A new exercise resets the per-exercise lookup set — that set is // what the answer reports back, so it must not carry over — and diff --git a/test/domain/block-leak.test.ts b/test/domain/block-leak.test.ts index 58ab658..4bf58b8 100644 --- a/test/domain/block-leak.test.ts +++ b/test/domain/block-leak.test.ts @@ -87,7 +87,7 @@ describe("proseOf", () => { }); it("drops every block, not merely from the first one on", () => { - expect(proseOf("keep\n::task translate\ndrop\n::words\ndrop")).toBe("keep"); + expect(proseOf("keep\n::task translate\ndrop\n::words\ndrop | x")).toBe("keep"); }); it("resumes the prose after a block closes — gloss blocks sit inline", () => { @@ -221,3 +221,59 @@ describe("an unclosed ::gloss block", () => { expect(parseMessage(closed).body).toBe("After."); }); }); + +/* A recall round with gpt-oss-20b, captured verbatim through the server: + the exercise, his answer, and the marking. The model closed none of its + blocks, and each unclosed block swallowed what followed it. */ +describe("gpt-oss-20b, a recall round — blocks it never closed", () => { + const round = JSON.parse( + readFileSync(new URL("../fixtures/gpt-oss-20b-recall-round.json", import.meta.url), "utf8"), + ) as { answered: string; marking: string; reply: string }; + + const recall = (task: ParsedMessage["task"]) => { + if (!task || task.type !== "recall") throw new Error(`expected a recall task, got ${task?.type}`); + return task.items; + }; + + it("shows the feedback written after ::result, instead of reading it as marks", () => { + const m = parseMessage(round.marking); + expect(m.body).toContain("You wrote an extra line"); + expect(m.body).toContain("your score for this unit is about"); + expect(m.body).not.toMatch(/\n{3,}|---$/); + expect(m.results?.map((r) => r.item)).toEqual(["I", "you", "we", "there", "one"]); + }); + + it("never makes a markdown rule an exercise item", () => { + expect(recall(parseMessage(round.answered).task).map((i) => i.q)).toEqual([ + "I", + "you", + "we", + "there", + "one", + ]); + expect(parse(round.answered).task?.rows).toContain("---"); // lib, unaided + }); + + it("drops a recall hint that is the answer itself", () => { + const items = recall(parseMessage(round.marking).task); + expect(items.map((i) => [i.q, i.hint])).toEqual([ + ["two", ""], + ["that", ""], + ["I (humble)", ""], + ]); + }); + + it("keeps only the Korean line of a translate row", () => { + expect(translateItems(parseMessage(round.reply).task).map((i) => i.q)).toEqual([ + "나", + "너", + "우리", + "거기", + "하나", + ]); + }); + + it("keeps the closing remark after ::words as prose", () => { + expect(parseMessage(round.answered).body).toContain("You have just started"); + }); +}); diff --git a/test/domain/gate-audit.test.ts b/test/domain/gate-audit.test.ts index 27e2623..57bab54 100644 --- a/test/domain/gate-audit.test.ts +++ b/test/domain/gate-audit.test.ts @@ -14,7 +14,7 @@ import { SqliteWasmDb } from "@app/db/sqlite-wasm-core.js"; import { migrate } from "@app/db/migrate.js"; import { editCard } from "@app/db/writes.js"; import type { Db } from "@app/db/types.js"; -import { parse } from "@lib/blocks.js"; +import { parseMessage } from "@app/domain/gloss.js"; import { proseIsKorean, scanTask, taskMaterial } from "@lib/gate.js"; import { markKnown } from "@lib/srs.js"; import { allowedSet, courseScaffold, gateFor, metWords, UNITS, unitOf } from "@app/domain/gate.js"; @@ -83,7 +83,7 @@ async function audit(file: string): Promise { const flagged: string[] = []; for (const [i, m] of messages.entries()) { - const parsed = parse(m); + const parsed = parseMessage(m); const tokens = taskMaterial(parsed.task).join(" ").match(/[가-힣]+/g) ?? []; const heads = await headsFor(db, tokens); const found = scanTask(parsed, { allowed, scaffold, heads, unitOf }).map((f) => f.word); diff --git a/test/domain/learner.test.ts b/test/domain/learner.test.ts index b8a8df6..c1bcef0 100644 --- a/test/domain/learner.test.ts +++ b/test/domain/learner.test.ts @@ -224,6 +224,15 @@ describe("the letter-level check", () => { expect(expectedFor("a mountain", words)).toBe(""); }); + it("tells a qualified meaning from the bare one", () => { + const pronouns = [ + { ko: "나", gloss: "I", note: "" }, + { ko: "저", gloss: "I (humble)", note: "" }, + ]; + expect(expectedFor("I (humble)", pronouns)).toBe("저"); + expect(expectedFor("I", pronouns)).toBe("나"); + }); + it("hands the tutor the jamo comparison for a wrong spelling, and nothing for a right one", () => { const task = { type: "recall" as const, items: [{ q: "to be short", hint: "" }, { q: "the sea", hint: "" }] }; const block = recallLetterBlock(task, ["빫다", "바다"], words); diff --git a/test/domain/marking.test.ts b/test/domain/marking.test.ts new file mode 100644 index 0000000..14ef0a4 --- /dev/null +++ b/test/domain/marking.test.ts @@ -0,0 +1,249 @@ +/* Which word each mark is for — against three real gpt-oss-20b sessions: + one that wrote ::result English prompt first, one that wrote no ::result + at all and marked in prose, and one that fenced every block in ``` and put + the right answer where the slip belongs. Captured verbatim through the + server. */ + +import { describe, it, expect, beforeAll, afterAll, vi } from "vitest"; +import { readFileSync } from "node:fs"; +import type { Db } from "@app/db/types.js"; +import { parseMessage } from "@app/domain/gloss.js"; +import { readMarks, readResults } from "@app/domain/marking.js"; +import { expectedFor } from "@app/domain/letters.js"; +import { applyReply } from "@app/domain/turn.js"; +import { readEvidence } from "@app/domain/evidence.js"; +import { isExerciseAnswer, noteAnswer, readProgress } from "@app/domain/progress.js"; +import { dictionaryDb } from "../helpers/dict-db.js"; + +const round = JSON.parse( + readFileSync(new URL("../fixtures/gpt-oss-20b-recall-round.json", import.meta.url), "utf8"), +) as { answered: string; answer: string; marking: string }; + +const answered = parseMessage(round.answered); +const marking = parseMessage(round.marking); + +describe("reading the marks", () => { + it("attaches the model's English-first rows to the words they mark", () => { + expect(readResults(marking.results, answered)).toEqual([ + { item: "나", ok: true, mistakenFor: "" }, + { item: "너", ok: true, mistakenFor: "" }, + { item: "우리", ok: false, mistakenFor: "우라" }, + { item: "거기", ok: true, mistakenFor: "" }, + { item: "하나", ok: true, mistakenFor: "" }, + ]); + }); + + it("leaves rows written as the prompt asks exactly as they are", () => { + const rows = [ + { item: "닭", ok: true, mistakenFor: "" }, + { item: "여덟", ok: false, mistakenFor: "여덜" }, + ]; + expect(readResults(rows, answered)).toEqual(rows); + }); + + it("takes the word out of quotation marks, and what he wrote out of an arrow", () => { + expect(readResults([{ item: "“우리”", ok: false, mistakenFor: "우라 → 우리" }], null)).toEqual([ + { item: "우리", ok: false, mistakenFor: "우라" }, + ]); + }); + + it("does not guess which word of a sentence a mark is for", () => { + const rows = [{ item: "머리 아파", ok: false, mistakenFor: "read 머리 as leg" }]; + expect(readResults(rows, answered)).toEqual(rows); + }); + + it("does not attach an English row that is not one of the exercise's prompts", () => { + const rows = [{ item: "the sea", ok: false, mistakenFor: "" }]; + expect(readResults(rows, answered)).toEqual(rows); + }); +}); + +describe("the round, recorded", () => { + let db: Db; + beforeAll(async () => { + db = await dictionaryDb(1); + }); + afterAll(async () => { + vi.unstubAllGlobals(); + await db.close(); + }); + + it("puts every mark from the live round into the evidence", async () => { + const applied = await applyReply(db, marking, { lookups: [], today: 20_000, answered }); + expect(applied.results?.recorded).toEqual(["나", "너", "우리", "거기", "하나"]); + expect(applied.results?.ignored).toEqual([]); + + const evidence = await readEvidence(db, ["우리", "나"]); + expect(evidence.get("우리")).toMatchObject({ ok: 0, wrong: 1 }); + expect(evidence.get("나")).toMatchObject({ ok: 1, wrong: 0 }); + expect(await db.get("SELECT mistook FROM confusion WHERE word = '우리'")).toEqual({ mistook: "우라" }); + }); +}); + +/* The second session: no ::result at all, marks in prose, and a ::progress + line on messages that answered nothing. The whole transcript, verbatim. */ +const session = JSON.parse( + readFileSync(new URL("../fixtures/gpt-oss-20b-prose-marks.json", import.meta.url), "utf8"), +) as { turns: { role: "user" | "assistant"; body: string }[] }; +const turn = (i: number) => parseMessage(session.turns[i]!.body); + +describe("marks written in prose", () => { + it("reads ✓ and ✗ lines through the prompts of the exercise he answered", () => { + expect(turn(3).results).toBeNull(); + expect(readMarks(turn(3), turn(1))).toEqual([ + { item: "나", ok: true, mistakenFor: "" }, + { item: "너", ok: true, mistakenFor: "" }, + { item: "이", ok: true, mistakenFor: "" }, + { item: "그", ok: true, mistakenFor: "" }, + { item: "저", ok: false, mistakenFor: "나" }, + ]); + }); + + it("tells the prompt a line is about from a shorter one inside it", () => { + // "I (humble)" contains "I"; the longer prompt is the one meant. + expect(readMarks(turn(7), turn(5))).toEqual([ + { item: "바다", ok: true, mistakenFor: "" }, + { item: "다리", ok: false, mistakenFor: "나" }, + { item: "저", ok: false, mistakenFor: "나" }, + { item: "너", ok: true, mistakenFor: "" }, + { item: "하나", ok: true, mistakenFor: "" }, + ]); + }); + + it("marks nothing the answered exercise did not ask for", () => { + const stray = parseMessage("✓ 바다 | the sea\n✗ 소리 | sound → 소라"); + expect(readMarks(stray, turn(1))).toBeNull(); + }); + + it("prefers the ::result block when there is one", () => { + expect(readMarks(marking, answered)?.map((r) => r.item)).toEqual(["나", "너", "우리", "거기", "하나"]); + }); +}); + +describe("the second session, replayed", () => { + let db: Db; + beforeAll(async () => { + db = await dictionaryDb(1); + }); + afterAll(async () => { + vi.unstubAllGlobals(); + await db.close(); + }); + + it("records every round, and moves progress only in reply to an answer", async () => { + const stored: number[] = []; + let exercise: ReturnType | null = null; + for (let i = 1; i < session.turns.length; i++) { + const { role, body } = session.turns[i]!; + if (role === "user") { + await noteAnswer(db, await readProgress(db), body); + continue; + } + const reply = parseMessage(body); + const asked = session.turns[i - 1]!; + await applyReply(db, reply, { + lookups: [], + today: 20_000, + answered: exercise, + afterAnswer: asked.role === "user" && isExerciseAnswer(asked.body), + }); + if (reply.task) exercise = reply; + stored.push((await readProgress(db)).confidence["1.1"] ?? 0); + } + // The model reported 80, 60, 65, 70, 75, 80. Clamped, and read only + // after an answer: 80 → 25, 60 → 50, 70 → 70; the requests change nothing. + expect(stored).toEqual([0, 25, 25, 50, 50, 70, 70, 70]); + + const evidence = await readEvidence(db, ["저", "거기", "나무"]); + expect(evidence.get("저")).toMatchObject({ ok: 0, wrong: 2 }); + expect(evidence.get("거기")).toMatchObject({ wrong: 1 }); + expect(evidence.get("나무")).toMatchObject({ ok: 1 }); + expect(await db.get("SELECT mistook FROM confusion WHERE word = '거기'")).toEqual({ mistook: "여기" }); + }); +}); + +/* The third session: every block fenced in ```, the right answer where the + slip belongs, and qualified prompts — "I (humble)" beside "I". */ +const fenced = JSON.parse( + readFileSync(new URL("../fixtures/gpt-oss-20b-fenced-blocks.json", import.meta.url), "utf8"), +) as { turns: { role: "user" | "assistant"; body: string }[] }; +const fencedTurn = (i: number) => parseMessage(fenced.turns[i]!.body); + +describe("fenced blocks and qualified prompts", () => { + it("reads the blocks inside the fences, and leaves no fence in the prose", () => { + const m = fencedTurn(3); + expect(m.results).toHaveLength(17); + expect(m.task?.type).toBe("recall"); + expect(m.body).not.toContain("```"); + }); + + it("tells I (humble) from I, and there from over there", () => { + const { words } = fencedTurn(3); + expect(expectedFor("I (humble)", words)).toBe("저"); + expect(expectedFor("I", words)).toBe("나"); + expect(expectedFor("There (there)", words)).toBe("거기"); + }); + + it("drops every hint that is one of the message's own words", () => { + const task = fencedTurn(3).task; + if (task?.type !== "recall") throw new Error("expected recall"); + expect(task.items.map((i) => i.hint)).toEqual(["", "", "", "", "", ""]); + }); + + it("takes what he mistook a word for from what he wrote, when the mark gives the right answer", () => { + expect(readMarks(fencedTurn(3), fencedTurn(1), fenced.turns[2]!.body)?.find((r) => r.item === "우리")).toEqual({ + item: "우리", + ok: false, + mistakenFor: "우라", + }); + }); +}); + +describe("the third session, replayed", () => { + let db: Db; + beforeAll(async () => { + db = await dictionaryDb(1); + }); + afterAll(async () => { + vi.unstubAllGlobals(); + await db.close(); + }); + + it("files every mark under the word it is for", async () => { + let exercise: ReturnType | null = null; + for (let i = 1; i < fenced.turns.length; i++) { + const { role, body } = fenced.turns[i]!; + if (role === "user") { + await noteAnswer(db, await readProgress(db), body); + continue; + } + const asked = fenced.turns[i - 1]!; + const reply = parseMessage(body); + await applyReply(db, reply, { + lookups: [], + today: 20_000, + answered: exercise, + answer: asked.body, + afterAnswer: asked.role === "user" && isExerciseAnswer(asked.body), + }); + if (reply.task) exercise = reply; + } + + const evidence = await readEvidence(db, ["저", "나", "거기", "우리"]); + // "I (humble) → 나" was marked wrong in round two and right in round + // three; it used to be filed under 나, and dropped as a repeat. + expect(evidence.get("저")).toMatchObject({ ok: 1, wrong: 1 }); + expect(evidence.get("나")).toMatchObject({ ok: 2, wrong: 0 }); + expect(evidence.get("거기")).toMatchObject({ ok: 1, wrong: 1 }); + expect(evidence.get("우리")).toMatchObject({ ok: 0, wrong: 1 }); + + const confusions = await db.all<{ word: string; mistook: string }>( + "SELECT word, mistook FROM confusion ORDER BY word", + ); + expect(confusions).toEqual([ + { word: "거기", mistook: "여기" }, + { word: "우리", mistook: "우라" }, + { word: "저", mistook: "나" }, + ]); + }); +}); diff --git a/test/fixtures/gpt-oss-20b-fenced-blocks.json b/test/fixtures/gpt-oss-20b-fenced-blocks.json new file mode 100644 index 0000000..b25ad1e --- /dev/null +++ b/test/fixtures/gpt-oss-20b-fenced-blocks.json @@ -0,0 +1,37 @@ +{ + "source": "gpt-oss-20b (LM Studio) through server/src/backends/openai.ts, unit 1.1, 16 Sep 2026, third session. The whole transcript, verbatim. Every block wrapped in a markdown code fence; ::result third fields hold the right answer rather than the slip; recall hints are the answers.", + "turns": [ + { + "role": "user", + "body": "Let's start unit 1.1 자음과 모음 (Consonants & vowels). Give me the full introduction, then a first exercise." + }, + { + "role": "assistant", + "body": "Welcome to Unit 1.1: **자음과 모음** – the consonants and vowels that make up Korean syllables. \nIn this unit you will learn the ten basic vowels, the plain consonants, and how a single consonant plus one vowel creates a block (a “syllable”). We’ll keep it simple: no final consonants, no compound vowels, no tense or aspirated consonants.\n\n**What you’re learning now**\n\n1. **Plain consonants** – ㄱ ㄴ ㄷ ㄹ ㅁ ㅂ ㅅ ㅇ ㅈ ㅎ \n These are the first letters you’ll see in any Korean word. \n2. **Ten basic vowels** – ㅏ ㅑ ㅓ ㅕ ㅗ ㅛ ㅜ ㅠ ㅡ ㅣ \n Each vowel is a block on its own or combined with a consonant to form a syllable. \n3. **Blocks (syllables)** – one initial consonant + one vowel, no final consonant. \n For example: “가” (ㄱ+ㅏ) and “나” (ㄴ+ㅏ). \n\n**Why this is important**\n\nKorean is written in blocks that are each a syllable. If you can read the shape of every block, you can read any word without having to know how it sounds. We’ll start with the simplest shapes so you don’t get lost.\n\n**Common mistakes beginners make**\n\n- Mixing up consonants that look similar: ㄱ and ㅋ, ㅂ and ㅍ, ㅅ and ㅎ. \n Remember we only use plain consonants here – no aspirated or tense versions. \n- Forgetting the vowel part of a block. \n Every block has one vowel; if you miss it, the whole syllable becomes unreadable.\n\n**Examples to keep in mind**\n\n* 가 (ㄱ+ㅏ) → “ga” \n* 나 (ㄴ+ㅏ) → “na” \n* 바 (ㅂ+ㅏ) → “ba” \n* 다 (ㄷ+ㅏ) → “da” \n\nYou’ll see many of these in the words we will practice. The new words you can start reading now are: **나**, **너**, **우리**, **이**, **그**, **저**, **여기**, **거기**, **저기**, **어디**, **누구**, **아이**, **어머니**, **아버지**, **나무**, **머리**, **소리**, **하나**, **바다**, **다리**. \nEach of these has only one consonant and one vowel, so they are perfect for this unit.\n\n---\n\n### First exercise: **Recall** (you write the Korean)\n\nYou’ll see an English word or phrase and you must type the correct Korean block(s) using the on‑screen keyboard.\n\n```\n::task recall\nI | 나\nYou (casual) | 너\nWe | 우리\nTwo (Sino) | 이\nThat (near) | 그\nThere (over there) | 저기\nWhere? | 어디\nWho? | 누구\nChild | 아이\nMother | 어머니\nFather | 아버지\nTree | 나무\nHead | 머리\nSound | 소리\nOne (Native) | 하나\nSea | 바다\nBridge | 다리\n```\n\n---\n\n**Words used in this lesson**\n\n```\n::words\n나 | I, me (casual)\n너 | you (casual)\n우리 | we, our\n이 | two (Sino)\n그 | that (near you or already mentioned)\n저기 | over there\n어디 | where?\n누구 | who?\n아이 | child\n어머니 | mother\n아버지 | father\n나무 | tree\n머리 | head\n소리 | sound\n하나 | one (Native)\n바다 | sea\n다리 | bridge\n```" + }, + { + "role": "user", + "body": "My written answers:\nI → 나\nYou (casual) → 너\nWe → 우라\nTwo (Sino) → 이\nThat (near) → 그\nThere (over there) → 저기\nWhere? → 어디\nWho? → 누구\nChild → 아이\nMother → 어머니\nFather → 아버지\nTree → 나무\nHead → 머리\nSound → 소리\nOne (Native) → 하나\nSea → 바다\nBridge → 다리\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• We\n wanted 우리 [우=ㅇ+ㅜ · 리=ㄹ+ㅣ]\n wrote 우라 [우=ㅇ+ㅜ · 라=ㄹ+ㅏ]\n syllable 2 (라 for 리) — WRONG: vowel: wrote ㅏ, should be ㅣ — CORRECT, do not call these mistakes: first consonant ㄹ, batchim none\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "Great work on the first round! \nYou almost nailed every word—just a single mistake in **We**. \n\n---\n\n### Your results\n\n```\n::result\nI | ok\nYou (casual) | ok\nWe | wrong | 우리\nTwo (Sino) | ok\nThat (near) | ok\nThere (over there) | ok\nWhere? | ok\nWho? | ok\nChild | ok\nMother | ok\nFather | ok\nTree | ok\nHead | ok\nSound | ok\nOne (Native) | ok\nSea | ok\nBridge | ok\n```\n\n### Progress\n\n```\n::progress 94 | You’re almost done with this unit—just one mistake to fix.\n```\n\n---\n\nNow let’s reinforce a few new words that we haven’t practiced yet. \n\n```\n::task recall\nI | 나\nTwo (Sino) | 이\nThat (near) | 그\nI (humble) | 저\nHere | 여기\nThere (there) | 거기\n```\n\n### Words used in this lesson\n\n```\n::words\n나 | I, me (casual)\n이 | two (Sino)\n그 | that (near you or already mentioned)\n저 | I, me (humble)\n여기 | here\n거기 | there\n```" + }, + { + "role": "user", + "body": "My written answers:\nI → 나\nTwo (Sino) → 이\nThat (near) → 그\nI (humble) → 나\nHere → 여기\nThere (there) → 여기\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• There (there)\n wanted 거기 [거=ㄱ+ㅓ · 기=ㄱ+ㅣ]\n wrote 여기 [여=ㅇ+ㅕ · 기=ㄱ+ㅣ]\n syllable 1 (여 for 거) — WRONG: first consonant: wrote ㅇ, should be ㄱ; vowel: wrote ㅕ, should be ㅓ — CORRECT, do not call these mistakes: batchim none\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "You’re getting the hang of it! \nTwo small mistakes—let’s fix them.\n\n---\n\n### Your results\n\n```\n::result\nI | ok\nTwo (Sino) | ok\nThat (near) | ok\nI (humble) | wrong | 저\nHere | ok\nThere (there) | wrong | 거기\n```\n\n### Progress\n\n```\n::progress 98 | You’re almost finished with this unit—just a couple of words to practice.\n```\n\n---\n\nNow let’s practice the two words you missed. We’ll do another recall exercise, and we’ll add one new word that hasn’t appeared yet: **너** (you, casual).\n\n```\n::task recall\nI (humble) | 저\nThere (there) | 거기\nYou (casual) | 너\n```\n\n### Words used in this lesson\n\n```\n::words\n저 | I, me (humble)\n거기 | there\n너 | you (casual)\n```" + }, + { + "role": "user", + "body": "My written answers:\nI (humble) → 저\nThere (there) → 거기\nYou (casual) → 너\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "Excellent! All three words were read correctly.\n\n---\n\n### Your results\n\n```\n::result\nI (humble) | ok\nThere (there) | ok\nYou (casual) | ok\n```\n\n### Progress\n\n```\n::progress 100 | You’ve mastered all the basic blocks in this unit. Ready to move on to the next one!\n```" + } + ] +} diff --git a/test/fixtures/gpt-oss-20b-prose-marks.json b/test/fixtures/gpt-oss-20b-prose-marks.json new file mode 100644 index 0000000..4b72472 --- /dev/null +++ b/test/fixtures/gpt-oss-20b-prose-marks.json @@ -0,0 +1,69 @@ +{ + "source": "gpt-oss-20b (LM Studio) through server/src/backends/openai.ts, unit 1.1, 16 Sep 2026, second session. The whole transcript, verbatim. The model marked in prose only, with no ::result block, and wrote ::progress on messages that answered nothing.", + "turns": [ + { + "role": "user", + "body": "Let's start unit 1.1 자음과 모음 (Consonants & vowels). Give me the full introduction, then a first exercise." + }, + { + "role": "assistant", + "body": "Welcome! I’m your reading tutor for the very first lesson in Korean, unit 1.1 자음과 모음 (Consonants & vowels). \nThis unit is all about **how a block is built**: one initial consonant + one vowel. The blocks you will read are the simplest syllables in Korean – no final consonants, no compound vowels, no tense or aspirated sounds. \n\n### Why this rule matters \nEvery Korean word is made of these blocks stacked together. If you can read a block, you already know how to decode every other block that follows it. The first step is therefore the most fundamental: recognizing and saying the consonant‑vowel pair. \nThe ten basic vowels are ㅏ ㅑ ㅓ ㅕ ㅗ ㅛ ㅜ ㅠ ㅡ ㅣ. \nThe plain consonants you will see in this unit are ㄱ ㄴ ㄷ ㄹ ㅁ ㅂ ㅅ ㅇ ㅈ ㅎ. \n\n### A quick example of a block \nTake **나**: it is made of the consonant ㄴ and the vowel ㅏ, written together as 하나의 블록. \nWhen you see this block, you read “na”. \nThe same idea applies to every other block: \n- **너** (ㄴ+ㅓ) → “neo” \n- **이** (ㅇ+ㅣ) → “i” (the ㅇ is silent at the start) \n- **그** (ㄱ+ㅡ) → “geu” \n- **저** (ㅈ+ㅓ) → “jeo”\n\n### Common mistake \nA new learner often mixes up ㄴ and ㄹ because they look similar. If you read a block that starts with ㄴ but sound it like ㄹ, the meaning will be wrong. Practice the shapes and sounds of each consonant separately.\n\n### New words for this unit \nThe words we’ll use in your exercises are all simple blocks that fit the rule: **나 너 이 그 저**. \n\nYou’re ready to start reading! Let’s practice by writing these words from English prompts.\n\n---\n\n::task recall\nI | 나\nyou (casual) | 너\ntwo (Sino) | 이\nthat (near you) | 그\nI (humble) | 저\n\n::words\n나 | I | ㄴ+ㅏ\n너 | you (casual) | ㄴ+ㅓ\n이 | two (Sino) | ㅇ+ㅣ\n그 | that (near you) | ㄱ+ㅡ\n저 | I (humble) | ㅈ+ㅓ" + }, + { + "role": "user", + "body": "My written answers:\nI → 나\nyou (casual) → 너\ntwo (Sino) → 이\nthat (near you) → 그\nI (humble) → 나\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• I (humble)\n wanted 저 [저=ㅈ+ㅓ]\n wrote 나 [나=ㄴ+ㅏ]\n syllable 1 (나 for 저) — WRONG: first consonant: wrote ㄴ, should be ㅈ; vowel: wrote ㅏ, should be ㅓ — CORRECT, do not call these mistakes: batchim none\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "✓ 나 | I \n✓ 너 | you (casual) \n✓ 이 | two (Sino) \n✓ 그 | that (near you) \n✗ 나 | I (humble) → 저\n\n::words\n나 | I | ㄴ+ㅏ\n너 | you (casual) | ㄴ+ㅓ\n이 | two (Sino) | ㅇ+ㅣ\n그 | that (near you) | ㄱ+ㅡ\n나 | I (humble) | ㄴ+ㅏ\n\n::progress 80 | You’re reading most of the simple blocks correctly! Keep practicing to get all five right." + }, + { + "role": "user", + "body": "새 문제 주세요. Give me a new exercise." + }, + { + "role": "assistant", + "body": "Here’s a new practice set. Write the Korean words from the English prompts.\n\n---\n\n::task recall\nthe sea | 바다 \nbridge | 다리 \nI (humble) | 저 \nyou (casual) | 너 \none | 하나\n\n::words\n나 | I | ㄴ+ㅏ\n너 | you (casual) | ㄴ+ㅓ\n이 | two (Sino) | ㅇ+ㅣ\n그 | that (near you) | ㄱ+ㅡ\n저 | I (humble) | ㅈ+ㅓ\n바다 | the sea | ㅂ+ㅏ + ㄷ+ㅏ \n다리 | bridge | ㄷ+ㅏ + ㄹ+ㅣ \n하나 | one | 하(ㅎ+ㅏ) + 나(ㄴ+ㅏ)" + }, + { + "role": "user", + "body": "My written answers:\nthe sea → 바다\nbridge → 나\nI (humble) → 나\nyou (casual) → 너\none → 하나\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• bridge\n wanted 다리 [다=ㄷ+ㅏ · 리=ㄹ+ㅣ]\n wrote 나 [나=ㄴ+ㅏ]\n length differs: 2 syllables expected, 1 written | syllable 1 (나 for 다) — WRONG: first consonant: wrote ㄴ, should be ㄷ — CORRECT, do not call these mistakes: vowel ㅏ, batchim none\n• I (humble)\n wanted 저 [저=ㅈ+ㅓ]\n wrote 나 [나=ㄴ+ㅏ]\n syllable 1 (나 for 저) — WRONG: first consonant: wrote ㄴ, should be ㅈ; vowel: wrote ㅏ, should be ㅓ — CORRECT, do not call these mistakes: batchim none\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "✓ 바다 | the sea \n✗ 나 | bridge → 다리 \n✗ 나 | I (humble) → 저 \n✓ 너 | you (casual) \n✓ 하나 | one \n\n::words\n바다 | the sea | ㅂ+ㅏ + ㄷ+ㅏ \n다리 | bridge | ㄷ+ㅏ + ㄹ+ㅣ \n하나 | one | 하(ㅎ+ㅏ) + 나(ㄴ+ㅏ) \n너 | you (casual) | ㄴ+ㅓ \n이 | two (Sino) | ㅇ+ㅣ \n그 | that (near you) | ㄱ+ㅡ \n저 | I (humble) | ㅈ+ㅓ \n나 | wrong | ㄴ+ㅏ \n\n::progress 60 | You’re reading most of the simple blocks correctly! Keep practicing to get all five right." + }, + { + "role": "user", + "body": "새 문제 주세요. Give me a new exercise." + }, + { + "role": "assistant", + "body": "Here’s a fresh practice set. Write the Korean word that matches each English prompt.\n\n---\n\n::task recall\nhere | 여기 \nthere | 거기 \nover there | 저기 \ntree | 나무 \nsound | 소리\n\n::words\n여기 | here | ㅇ+ㅕ \n거기 | there | ㄱ+ㅓ + 기(ㄱ+ㅣ) \n저기 | over there | ㅈ+ㅓ + 기(ㄱ+ㅣ) \n나무 | tree | ㄴ+ㅏ + 무(ㅁ+ㅜ) \n소리 | sound | ㅅ+ㅗ + 리(ㄹ+ㅣ) \n\n::progress 65 | You’re getting more comfortable reading simple blocks! Keep practicing to nail all five." + }, + { + "role": "user", + "body": "My written answers:\nhere → 여기\nthere → 여기\nover there → 여기\ntree → 나무\nsound → 소리\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• there\n wanted 거기 [거=ㄱ+ㅓ · 기=ㄱ+ㅣ]\n wrote 여기 [여=ㅇ+ㅕ · 기=ㄱ+ㅣ]\n syllable 1 (여 for 거) — WRONG: first consonant: wrote ㅇ, should be ㄱ; vowel: wrote ㅕ, should be ㅓ — CORRECT, do not call these mistakes: batchim none\n• over there\n wanted 저기 [저=ㅈ+ㅓ · 기=ㄱ+ㅣ]\n wrote 여기 [여=ㅇ+ㅕ · 기=ㄱ+ㅣ]\n syllable 1 (여 for 저) — WRONG: first consonant: wrote ㅇ, should be ㅈ; vowel: wrote ㅕ, should be ㅓ — CORRECT, do not call these mistakes: batchim none\n\n(No lookups.)" + }, + { + "role": "assistant", + "body": "✓ 여기 | here \n✗ 여기 | there → 거기 \n✗ 여기 | over there → 저기 \n✓ 나무 | tree \n✓ 소리 | sound \n\n::words\n여기 | here | ㅇ+ㅕ \n거기 | there | ㄱ+ㅓ + 기(ㄱ+ㅣ) \n저기 | over there | ㅈ+ㅓ + 기(ㄱ+ㅣ) \n나무 | tree | ㄴ+ㅏ + 무(ㅁ+ㅜ) \n소리 | sound | ㅅ+ㅗ + 리(ㄹ+ㅣ) \n\n::progress 70 | You’re reading most of the simple blocks correctly! Keep practicing to get all five right." + }, + { + "role": "user", + "body": "새 문제 주세요. Give me a new exercise." + }, + { + "role": "assistant", + "body": "You’re doing great! Let’s keep practicing by writing the Korean words that match each English prompt. \n\n---\n\n::task recall\nI (humble) | 저 \nyou (casual) | 너 \ntree | 나무 \nsound | 소리 \nthe sea | 바다\n\n::words\n저 | I (humble) | ㅈ+ㅓ \n너 | you (casual) | ㄴ+ㅓ \n나무 | tree | ㄴ+ㅏ + 무(ㅁ+ㅜ) \n소리 | sound | ㅅ+ㅗ + 리(ㄹ+ㅣ) \n바다 | the sea | ㅂ+ㅏ + 다(ㄷ+ㅏ) \n\n::progress 75 | Keep up! You’re reading most of these simple blocks correctly." + }, + { + "role": "user", + "body": "새 문제 주세요. Give me a new exercise." + }, + { + "role": "assistant", + "body": "Here’s a fresh practice set. Write the Korean word that matches each English prompt.\n\n---\n\n::task recall\nI (humble) | 저 \nyou (casual) | 너 \nchild | 아이 \nmother | 어머니 \nhead | 머리\n\n::words\n저 | I (humble) | ㅈ+ㅓ \n너 | you (casual) | ㄴ+ㅓ \n아이 | child | 아(ㅇ+ㅏ) + 이(ㅇ+ㅣ) \n어머니 | mother | 어(ㅇ+ㅓ) + 머(ㅁ+ㅔ?) + 니(ㄴ+ㅣ) \n머리 | head | 머(ㅁ+ㅔ?) + 리(ㄹ+ㅣ) \n\n::progress 80 | You’re reading most of these simple blocks correctly! Keep practicing to get all five right." + } + ] +} diff --git a/test/fixtures/gpt-oss-20b-recall-round.json b/test/fixtures/gpt-oss-20b-recall-round.json new file mode 100644 index 0000000..2b1ceda --- /dev/null +++ b/test/fixtures/gpt-oss-20b-recall-round.json @@ -0,0 +1,8 @@ +{ + "source": "gpt-oss-20b (LM Studio) through server/src/backends/openai.ts, unit 1.1, 16 Sep 2026. Verbatim.", + "answered": "**Unit 1.1 자음과 모음 – Consonants & vowels**\n\nThe first thing you’ll learn is how Korean letters are put together into *blocks* (or syllables). \nEvery block has two parts: an initial consonant (the *기초 자음*) and a vowel. \nWe will only use the plain, single‑letter consonants that can start a block:\n\nㄱ ㄴ ㄷ ㄹ ㅁ ㅂ ㅅ ㅇ ㅈㅎ\n\nand the ten basic vowels:\n\nㅏ ㅑ ㅓ ㅕ ㅗ ㅛ ㅜ ㅠ ㅡ ㅣ\n\nNo final consonant will appear yet, so each block looks like “ㄱ + ㅏ” → **가**, “ㅇ + ㅔ” → **여**.\n\nBecause the syllable is just a pair of letters, you can read it by looking at the shape: \nthe initial consonant on the left and the vowel on the right. \nIf you see ㄴ and ㅓ together, you know it’s **너**; if you see ㅇ and ㅜ together, it’s **우리**.\n\nThe goal of this unit is to recognize these blocks and write them out. \nYou’ll practice by reading Korean words that contain only one block (no final consonant) and writing the English meaning or the Korean word itself.\n\n---\n\n### First exercise – *Recall* \n\nWrite the Korean syllables for each English prompt. \n(Use the new words listed in this unit.)\n\n::task recall\nI | 나 \nyou | 너 \nwe | 우리 \nthere | 거기 \none | 하나\n\n---\n\n::words\n나 | I | (plain consonant ㄴ + vowel ㅏ) \n너 | you | (ㄴ + ㅓ) \n우리 | we | (ㅇ + ㅜ) \n거기 | there | (여(ㅇ+ㅕ), 기(ㄱ+ㅣ)) \n하나 | one | (ㅎ + ㅏ)\n\n---\n\n*You have just started. Keep practicing the shapes of these syllables, and you’ll be ready to read more Korean words soon.*", + "answer": "My written answers:\nI → 나\nyou → 너\nwe → 우라\nthere → 거기\none → 하나\n--- → 나\n\n════ LETTER-LEVEL CHECK — computed by the app ════\nThis is the actual jamo comparison. Use it exactly. Do NOT work out for yourself which letter was wrong, and never call a letter wrong that is listed as correct.\n• we\n wanted 우리 [우=ㅇ+ㅜ · 리=ㄹ+ㅣ]\n wrote 우라 [우=ㅇ+ㅜ · 라=ㄹ+ㅏ]\n syllable 2 (라 for 리) — WRONG: vowel: wrote ㅏ, should be ㅣ — CORRECT, do not call these mistakes: first consonant ㄹ, batchim none\n\n(No lookups.)", + "marking": "You did most of them right, just a couple of small mistakes:\n\n::result\nI | ok | 나 \nyou | ok | 너 \nwe | wrong | 우라 → 우리 \nthere | ok | 거기 \none | ok | 하나\n\nYou wrote an extra line \"--- → 나\" that didn’t match any prompt – you can ignore it.\n\nThe first time you’ve answered, your score for this unit is about **5 %**. Keep practicing the shapes of these syllables; we’ll add a few more words next.\n\n---\n\n### Next exercise – *Recall* \n\nWrite the Korean syllables for each English prompt. \n(Use the new words listed in this unit.)\n\n::task recall\ntwo | 이 \nthat | 그 \nI (humble) | 저\n\n---\n\n::words\n나 | I | (ㄴ + ㅏ) \n너 | you | (ㄴ + ㅓ) \n우리 | we | (ㅇ + ㅜ) \n거기 | there | (여(ㅇ+ㅕ), 기(ㄱ+ㅣ)) \n하나 | one | (ㅎ + ㅏ) \n우라 | wrong answer for 우리 | (ㅇ + ㅜ, 라(ㄹ+ㅏ)) \n이 | two | (ㅇ + ㅣ) \n그 | that | (ㄱ + ㅡ) \n저 | I (humble) | (ㅈ + ㅓ)\n\n::progress 5 %", + "asked": "Give me a translate exercise that uses the words 도서관 and 병원.", + "reply": "**Translate – read the Korean, type the English meaning**\n\nHere are simple words that contain only one block each (no final consonant). \n\n::task translate\n나 | I \n너 | you \n우리 | we \n거기 | there \n하나 | one\n\n---\n\n::words\n나 | I | (ㄴ + ㅏ) \n너 | you | (ㄴ + ㅓ) \n우리 | we | (ㅇ + ㅜ) \n거기 | there | (여(ㅇ+ㅕ), 기(ㄱ+ㅣ)) \n하나 | one | (ㅎ + ㅏ)\n\n---\n\n*Try to read each Korean word and write the English translation. Good luck!*" +}