On a reasoning model the reasoning is spent from the same completion budget
and it goes first, so a small ceiling truncates the actual reply away. It
fails silently and looks exactly like a model that cannot follow the system
prompt, which is the expensive way to debug it.
Measured on gpt-oss-20b against the real ~3.5k-token system prompt, same
prompt and same model, only the ceiling changing:
1400 an empty string, 0 bytes, after 27s
2048 a truncated half-Korean fragment, none of the required blocks
8000 99% English prose, no romanization, 0 batchim violations in the
task lines, all three blocks, 11s
The old comment justified 2048 by worrying about 4k-context local models.
That reasoning was wrong: the system prompt alone is ~3.5k tokens, so such
a model cannot run this app at all and there was nothing to protect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1.7 KiB
1.7 KiB