Never cut an emoji in half at the text cap - #115
Merged
Merged
Conversation
The compact view cuts a long block at TEXT_CAP UTF-16 units. An emoji, or any character past U+FFFF, is two of them, so when its first unit was the last one inside the cap the view printed half of it: a lone surrogate, which a terminal shows as the replacement character and which makes the printed line malformed UTF-16. A block with no sentence end inside the cap (a post, a title, a list of tags) takes this plain cut. The plain cut now steps back one unit when it would land between the two halves. The line and sentence cuts already end on a newline or on punctuation, so they never split one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
only-cli
pushed a commit
that referenced
this pull request
Sep 28, 2026
#115 stopped the compact view's cap from printing half of a character past U+FFFF. Two other cuts count UTF-16 units the same way and were left as they were: find's snippet window, which opens 60 units before the match and runs 200, and read's cut of a first block bigger than its whole budget, at budget * 4 units. Either edge can land between the two halves of an emoji, and then the line carries a lone surrogate, which a terminal prints as the replacement character. Each edge now steps back one unit when it would land there, the rule render.js already uses. A snippet's opening edge so takes the whole emoji, and its closing edge and read's cut leave it out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The compact view cuts a long block at
TEXT_CAP(200) UTF-16 units. An emoji, or any character past U+FFFF, takes two units, so when its first unit is the 200th the view printed half of it. That is a lone surrogate, which a terminal shows as�, and the printed line is no longer well-formed UTF-16. The plain cut is the one a block takes when no sentence ends inside the cap, which is common for posts, titles, and tag lists, exactly where emoji are.A block of 199
x, then 😀, then more text:The plain cut now steps back one unit when it would land between the two halves. The line cut ends on a newline and the sentence cut on punctuation, so neither can split a character, and both are unchanged.
oc readand--jsonprint whole blocks and were never affected.Test.
tests/distill.test.jsgains one test next to the sentence-cut test: the block above renders well-formed, with the 199xfollowed directly by the cut marker. Onmainit fails onisWellFormed. Full suite: 311 tests, 308 pass, 3 skipped, 0 fail.Left alone. The
+N charscount stays in UTF-16 units, as it was.🤖 Generated with Claude Code