The rules were a Markdown table, and a table cell cannot hold a bare `|`. It
has to be written `\|`, which renders as a pipe and copies as a backslash and
a pipe. Every one of these rules is an alternation full of pipes, so anyone
reading README.md rather than a rendered view -- which, for a self-hosted
thing, is most of the time -- got a regex whose alternation had quietly become
literal characters. It matches nothing, and nothing about it looks wrong.
That is not hypothetical: it is how this came up. The Reddit rule is the
longest row, and reading it out of the file gave something that plainly did
not work, so the backslashes came out. Which fixed the copy and broke the
table -- four pipes turned into column separators, the row became seven cells,
the separator row was widened to seven to match, and `(.*)` picked up an
escape on the way past.
Restoring the row would have left the trap exactly where it was, for the next
person or the same one. So the table is now a code block: each rule is a
comment naming the platform, then the find field, then the replace field, one
per line. Nothing is escaped, the file and the render agree, and each field is
a whole line to select.
Every rule is checked from both directions. Out of README.md byte for byte
with no unescaping step, which is the copy-from-the-file path; and out of the
rendered HTML with entities decoded, which is the copy-from-the-page path.
Both give all seven rules, both compile, and both rewrite all seventeen sample
URLs correctly -- mobile.x.com, threads.net, the TikTok vm. and vt. hosts,
five Reddit subdomains, a /s/ share link and a redd.it short code. The two
extractions are byte-identical to each other, which is the property that was
missing before.
Also reformats the rest of the file, and stops calling it a table on the way
past.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu