6 Commits
Author SHA1 Message Date
thatguygriff b9b56e2195 Merge pull request 'Show the pictures inside Reddit comments' (#5) from reddit-inline-images into main
CI / Typecheck, test, build (push) Successful in 17s
Publish / Build and push (push) Successful in 16s
Publish / Move the working version on (push) Failing after 2s
Reviewed-on: #5
2026-08-27 19:29:59 +00:00
thatguygriffandClaude Opus 5 031101c382 Show a picture whose address was typed without a scheme
CI / Typecheck, test, build (pull_request) Successful in 26s
`preview.redd.it/lz4drsqh0clh1.jpeg?width=1290&...` was rendering as plain
text -- not a picture, and not even a link. The bare-address rule has always
insisted on `https://`, so anything copied out of an address bar, where the
browser hides the scheme, fell through to nothing at all. That predates the
inline images from the last commit; it just did not matter until pictures
started being worth placing.

The rule is deliberately narrow: a host with a dot, a path, and an image
extension. Widening it to every schemeless address would be the obvious move
and is wrong, because a comment thread is full of dotted, slashed prose --
`src/render/post.ts` would become a link to a website in Tonga, and
`node_modules/foo/bar.js` a website in Jersey. Requiring the extension costs
nothing here, since the only thing worth guessing a scheme for is a picture.

https is assumed. Every host that serves these redirects to it anyway.

Checked both ways: the four address shapes that should become pictures do,
and nine pieces of ordinary prose that must not -- file paths, a Windows
drive letter, a relative path, an email address followed by a filename, a
version number -- still do not.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu
2026-08-27 16:19:25 -03:00
thatguygriffandClaude Opus 5 0db18547c8 Show the pictures inside Reddit comments
CI / Typecheck, test, build (pull_request) Successful in 11s
A comment that was a picture rendered as either a link or, for a Giphy, the
literal text `![gif](giphy|Ve7wX45gaOFmw8eeEM)`. On r/aww that is most of the
thread.

Reddit writes an inline image as a token rather than an address, in three
shapes: `![gif](giphy|ID)` for a Giphy, `![img](emote|t5_2th52|4358)` for a
subreddit emote, and `![img](jo8gf0ca92zd1)` for an image uploaded straight to
the comment. The useful part is that all three tokens are keys in that same
comment's own `media_metadata`, so this is one lookup and not three special
cases. Nothing in the adapter has to know what Giphy is.

The fourth shape is someone pasting the address of a picture, which on Reddit
is how most images in comments actually arrive -- 117 of them against 21
Giphys in the sample I scanned. Those are shown as pictures too, decided by
the file extension. A link that is not to an image stays a link.

Animated ones take `s.gif` over `s.mp4` even though the MP4 is several times
smaller: a GIF moves on its own in an `<img>`, and an MP4 would need a player
element with autoplay, loop and muted set, for something the size of a
postage stamp.

Everything goes through the `/m/` proxy, like all other media. Without that a
comment thread would have the reader's browser fetch dozens of files straight
from Reddit, which is the one thing this whole app exists to avoid.

Placing the image is the renderer's job, not the Markdown parser's, because
the proxy is a render-time concern and markdown.ts knows nothing about it --
so it takes an optional `ImageRenderer` and, without one, an image stays a
link exactly as before. The parser also learned `![...]` proper: the link rule
was matching from the `[` and stranding the `!` as text.

Verified on r/aww/comments/171dxph, which carries one Giphy and 77 pasted
images: 35 render on the first page, all 35 load, all 35 through the proxy,
no upstream address reaches the page, no token is left unresolved, and
nothing overflows the column or scrolls the page sideways.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu
2026-08-27 16:14:06 -03:00
thatguygriff 9abc0eaf62 Merge pull request 'Move the working version on after a release' (#4) from bump-version-after-release into main
CI / Typecheck, test, build (push) Waiting to run
Publish / Build and push (push) Waiting to run
Publish / Move the working version on (push) Blocked by required conditions
Reviewed-on: #4
2026-08-27 16:27:51 +00:00
thatguygriffandClaude Opus 5 24db9be5af Make the rewrite rules copyable from the file, not just the render
Publish / Build and push (pull_request) Successful in 26s
Publish / Move the working version on (pull_request) Skipped
CI / Typecheck, test, build (pull_request) Successful in 30s
The rules were a Markdown table, and a table cell cannot hold a bare `|`. It
has to be written `\|`, which renders as a pipe and copies as a backslash and
a pipe. Every one of these rules is an alternation full of pipes, so anyone
reading README.md rather than a rendered view -- which, for a self-hosted
thing, is most of the time -- got a regex whose alternation had quietly become
literal characters. It matches nothing, and nothing about it looks wrong.

That is not hypothetical: it is how this came up. The Reddit rule is the
longest row, and reading it out of the file gave something that plainly did
not work, so the backslashes came out. Which fixed the copy and broke the
table -- four pipes turned into column separators, the row became seven cells,
the separator row was widened to seven to match, and `(.*)` picked up an
escape on the way past.

Restoring the row would have left the trap exactly where it was, for the next
person or the same one. So the table is now a code block: each rule is a
comment naming the platform, then the find field, then the replace field, one
per line. Nothing is escaped, the file and the render agree, and each field is
a whole line to select.

Every rule is checked from both directions. Out of README.md byte for byte
with no unescaping step, which is the copy-from-the-file path; and out of the
rendered HTML with entities decoded, which is the copy-from-the-page path.
Both give all seven rules, both compile, and both rewrite all seventeen sample
URLs correctly -- mobile.x.com, threads.net, the TikTok vm. and vt. hosts,
five Reddit subdomains, a /s/ share link and a redd.it short code. The two
extractions are byte-identical to each other, which is the property that was
missing before.

Also reformats the rest of the file, and stops calling it a table on the way
past.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu
2026-08-27 13:27:13 -03:00
thatguygriffandClaude Opus 5 503d8a8dec Move the working version on after a release
package.json has sat at 0.1.0 since the first commit, through three releases,
because nothing read it. That is fine right up until something does — an image
label, a health endpoint, a bug report quoting a version — at which point the
tree claims to be a version that shipped long ago.

A new `bump` job takes the tag that was just published, works out the next
patch from it, and commits that to main. After 1.2.0 the tree says 1.2.1: not
a version that exists, which is the point. A build from main is then legible
as "after 1.2.0" rather than as 1.2.0 itself.

It sits in publish.yml rather than a workflow of its own so that it can say
`needs: build`. A version that failed to publish has not been released, and
moving past it would say that it had. Prereleases are skipped for the same
reason -- 1.2.3-rc1 is a candidate for a version that has not shipped, so
there is nothing yet to move past.

The bump goes through `npm version` rather than editing the file. The version
is in the lockfile too, in two places, and a tree where those disagree is worse
than one that is merely out of date.

Three smaller things. The patch arithmetic forces base ten, because a patch
number written 08 is otherwise read as octal and kills the job. The commit
carries `[skip ci]`, or pushing it starts another build of the image that was
just published. And the committer is a name that is not a person at a reserved
address that can never become one, so nothing here names the instance it runs
on.

Pushing to main needs a token that may write to the repository. The Actions
task token can where the instance allows it; where it does not, setting a
VERSION_BUMP_TOKEN secret overrides it. A push that is refused fails the job
with both of those as the suggestion rather than a bare 403.

package.json goes to 1.2.1 here, which is where the job would have left it had
it existed when 1.2.0 went out.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu
2026-08-27 13:19:05 -03:00
12 changed files with 495 additions and 47 deletions
+97
View File
@@ -141,3 +141,100 @@ jobs:
- name: Log out - name: Log out
if: always() && github.event_name != 'pull_request' if: always() && github.event_name != 'pull_request'
run: docker logout "${{ vars.REGISTRY }}" || true run: docker logout "${{ vars.REGISTRY }}" || true
# Once a release is out, the version in package.json has already shipped.
# Moving it on to the next patch means the working tree is never sitting on
# a number that is published and immutable, and that a build from main is
# always identifiable as "after 1.2.0" rather than "1.2.0, but not really".
#
# `needs: build` is the point of putting this here rather than in a workflow
# of its own: a version that failed to publish has not been released, and
# bumping past it would say it had.
bump:
name: Move the working version on
needs: build
# Tags only, and only final ones. A prerelease has not shipped the version
# it is a candidate for, so there is nothing yet to move past.
if: github.ref_type == 'tag' && !contains(github.ref_name, '-')
runs-on: ubuntu-latest
container:
image: node:22
steps:
# The tag names a commit in main's history, but the bump belongs on the
# branch, so this checks out main rather than the tag.
- uses: actions/checkout@v4
with:
ref: main
# The task token can push only if the instance allows Actions to
# write to the repository. Where it does not, set VERSION_BUMP_TOKEN
# to a personal access token with write access and it is used
# instead.
token: ${{ secrets.VERSION_BUMP_TOKEN || secrets.GITEA_TOKEN }}
- name: Work out the next patch version
id: next
run: |
set -euo pipefail
version="${{ github.ref_name }}"
version="${version#v}"
major="${version%%.*}"
rest="${version#*.}"
minor="${rest%%.*}"
patch="${rest##*.}"
# `10#` forces base ten: a patch number written 08 would otherwise be
# read as octal and fail to parse.
next="${major}.${minor}.$((10#${patch} + 1))"
echo "next=${next}" >> "$GITHUB_OUTPUT"
echo "Released ${version}; the working version becomes ${next}"
- name: Bump package.json
id: bump
env:
NEXT: ${{ steps.next.outputs.next }}
run: |
set -euo pipefail
current="$(node -p "require('./package.json').version")"
if [ "${current}" = "${NEXT}" ]; then
echo "package.json is already ${NEXT}; nothing to do."
echo "changed=false" >> "$GITHUB_OUTPUT"
exit 0
fi
# npm rather than editing the file: the version is in the lockfile
# too, in more than one place, and they have to agree.
npm version "${NEXT}" --no-git-tag-version --allow-same-version >/dev/null
echo "changed=true" >> "$GITHUB_OUTPUT"
echo "package.json ${current} -> ${NEXT}"
- name: Commit it to main
if: steps.bump.outputs.changed == 'true'
env:
NEXT: ${{ steps.next.outputs.next }}
run: |
set -euo pipefail
# A name that is not a person, and a reserved address that can never
# resolve to one. Nothing here names the instance it runs on.
git config user.name 'Release bot'
git config user.email '[email protected]'
git add package.json package-lock.json
# `[skip ci]` because this commit is a number and nothing else:
# without it the push to main starts another build of the very image
# that was just published.
git commit -m "Set the working version to ${NEXT} [skip ci]"
if ! git push origin HEAD:main; then
echo >&2
echo "Could not push the version bump to main. Either the Actions" >&2
echo "token has no write access to this repository, or main is" >&2
echo "protected against direct pushes. Set VERSION_BUMP_TOKEN to a" >&2
echo "token that may push to main, or allow that token past the" >&2
echo "branch protection." >&2
exit 1
fi
+21 -1
View File
@@ -147,7 +147,16 @@ Things worth knowing before editing:
timeline thumbnails, and the still is in `preview.images`. A gallery's pictures are in timeline thumbnails, and the still is in `preview.images`. A gallery's pictures are in
`media_metadata`, keyed and unordered; their order is only in `gallery_data`. Comment `media_metadata`, keyed and unordered; their order is only in `gallery_data`. Comment
bodies are Markdown, rendered by `src/render/markdown.ts` — escape first, then put bodies are Markdown, rendered by `src/render/markdown.ts` — escape first, then put
back the constructs we chose to support, never `body_html`. back the constructs we chose to support, never `body_html`. An image in a comment is
written as a token rather than an address — `![gif](giphy|Ve7wX45)`,
`![img](emote|t5_2th52|4358)`, `![img](jo8gf0ca92zd1)` — and in every case the token is
a key in that same comment's own `media_metadata`, so `resolveInlineImages` is one
lookup rather than three special cases. A bare `preview.redd.it` address pasted into a
comment is in there too, keyed by the id inside the URL. Prefer `s.gif` over `s.mp4`
for an animated one: a GIF moves in an `<img>` and an MP4 needs a player. An address
typed without a scheme counts as well, but only when it ends in an image extension —
the rule wants a host, a path *and* that extension, because comments are full of
dotted, slashed prose that must not turn into links.
- **Threads** — same media schema as Instagram (`src/platforms/meta-media.ts`). Its - **Threads** — same media schema as Instagram (`src/platforms/meta-media.ts`). Its
payloads are full of empty stub nodes, so the finder only accepts a node with actual payloads are full of empty stub nodes, so the finder only accepts a node with actual
candidates in it. The page ships the linked post, the author's follow-ups, other candidates in it. The page ships the linked post, the author's follow-ups, other
@@ -201,6 +210,17 @@ The registry comes from the `REGISTRY` repository variable, the image name from
`IMAGE_NAME` or the repository name, and credentials from `REGISTRY_USER` and the `IMAGE_NAME` or the repository name, and credentials from `REGISTRY_USER` and the
`REGISTRY_TOKEN` secret. Nothing about any particular deployment is committed here. `REGISTRY_TOKEN` secret. Nothing about any particular deployment is committed here.
A release also moves `package.json` on to the next patch version, committed to main by
the `bump` job — so the number in the tree is never one that has already shipped and
been made immutable. It lives in `publish.yml` rather than a workflow of its own so it
can say `needs: build`: a version that failed to publish has not been released, and
bumping past it would claim otherwise. Prereleases are skipped, being candidates for a
version that has not shipped. The bump goes through `npm version` rather than an edit in
place, because the version is in the lockfile too, in more than one place, and the two
have to agree. The commit carries `[skip ci]`, or pushing it would rebuild the image
that was just published. Pushing to main needs a token with write access —
`VERSION_BUMP_TOKEN` overrides the task token where that one cannot.
Two things any deployment has to get right, both learned the hard way: Two things any deployment has to get right, both learned the hard way:
- **Chromium needs more than the default 64Mi `/dev/shm`** or it crashes. Mount a - **Chromium needs more than the default 64Mi `/dev/shm`** or it crashes. Mount a
+48 -12
View File
@@ -21,15 +21,43 @@ readable in your history.
Replace `antisocial.example.com` with wherever you are running it. Replace `antisocial.example.com` with wherever you are running it.
| Platform | Find | Replace | Each rule is two fields. Both are on their own line below, and neither needs any
| --------- | ------------------------------------------------------------- | ------------------------------------------- | escaping — copy them straight out of this file.
| X | `/^https:\/\/(?:www\.\|mobile\.)?(?:x\|twitter)\.com\/(.*)$/` | `https://antisocial.example.com/x/$1` |
| Threads | `/^https:\/\/(?:www\.)?threads\.(?:net\|com)\/(.*)$/` | `https://antisocial.example.com/threads/$1` | ```text
| Instagram | `/^https:\/\/(?:www\.)?instagram\.com\/(.*)$/` | `https://antisocial.example.com/ig/$1` | # X
| TikTok | `/^https:\/\/(?:www\.\|vm\.\|vt\.)?tiktok\.com\/(.*)$/` | `https://antisocial.example.com/tiktok/$1` | /^https:\/\/(?:www\.|mobile\.)?(?:x|twitter)\.com\/(.*)$/
| Bluesky | `/^https:\/\/bsky\.app\/(.*)$/` | `https://antisocial.example.com/bsky/$1` | https://antisocial.example.com/x/$1
| Reddit | `/^https:\/\/(?:www\.\|old\.\|new\.\|np\.\|m\.)?reddit\.com\/(.*)$/` | `https://antisocial.example.com/reddit/$1` |
| Reddit | `/^https:\/\/redd\.it\/(.*)$/` | `https://antisocial.example.com/reddit/$1` | # Threads
/^https:\/\/(?:www\.)?threads\.(?:net|com)\/(.*)$/
https://antisocial.example.com/threads/$1
# Instagram
/^https:\/\/(?:www\.)?instagram\.com\/(.*)$/
https://antisocial.example.com/ig/$1
# TikTok
/^https:\/\/(?:www\.|vm\.|vt\.)?tiktok\.com\/(.*)$/
https://antisocial.example.com/tiktok/$1
# Bluesky
/^https:\/\/bsky\.app\/(.*)$/
https://antisocial.example.com/bsky/$1
# Reddit
/^https:\/\/(?:www\.|old\.|new\.|np\.|m\.)?reddit\.com\/(.*)$/
https://antisocial.example.com/reddit/$1
# Reddit short links
/^https:\/\/redd\.it\/(.*)$/
https://antisocial.example.com/reddit/$1
```
A code block rather than a table, because a table cell cannot hold a bare `|` — it has
to be written `\|`, which renders correctly and copies wrongly. The alternation in these
rules is full of them, and a regex whose pipes arrive as literal pipes matches nothing
and says nothing about why.
So `https://x.com/user/status/123` becomes So `https://x.com/user/status/123` becomes
`https://antisocial.example.com/x/user/status/123`. `https://antisocial.example.com/x/user/status/123`.
@@ -42,7 +70,8 @@ where it came from Reddit. A Reddit `/r/<sub>/s/<code>` share link is followed t
post it points at, and that permalink — not the opaque share code — is what the copy post it points at, and that permalink — not the opaque share code — is what the copy
button hands back. button hands back.
`/` serves this table with the live hostnames, if you'd rather read it there. `/` serves these rules with the live hostname already filled in, if you'd rather copy
them from there.
## How it works ## How it works
@@ -60,7 +89,7 @@ Each adapter layers its extraction, most structured first:
4. **Open Graph tags** — the floor, and enough to show something. 4. **Open Graph tags** — the floor, and enough to show something.
| Platform | Loads | Reads | | Platform | Loads | Reads |
| --------- | ---------------------------- | -------------------------------------------------------- | | --------- | ---------------------------- | --------------------------------------------------------------- |
| Bluesky | the public AT Protocol API | `getPostThread`; falls back to the post page | | Bluesky | the public AT Protocol API | `getPostThread`; falls back to the post page |
| X | `platform.twitter.com` embed | the `cdn.syndication.twimg.com/tweet-result` response | | X | `platform.twitter.com` embed | the `cdn.syndication.twimg.com/tweet-result` response |
| Instagram | `/embed/captioned/` | `shortcode_media`, then the rendered `<video>`/`<img>` | | Instagram | `/embed/captioned/` | `shortcode_media`, then the rendered `<video>`/`<img>` |
@@ -93,10 +122,17 @@ Reddit posts come with it: every comment the first page carried, nested the way
written. Each comment is a `<details>` element, so folding one takes its whole subtree written. Each comment is a `<details>` element, so folding one takes its whole subtree
with it, works without JavaScript and works from the keyboard; a collapsed comment says with it, works without JavaScript and works from the keyboard; a collapsed comment says
how many replies it is hiding. "Collapse all" is the one piece that needs the script, how many replies it is hiding. "Collapse all" is the one piece that needs the script,
which is why it only appears once the script has run. What was behind a *load more* is which is why it only appears once the script has run. What was behind a _load more_ is
not fetched — that is a second page and often a third — but it is counted and said out not fetched — that is a second page and often a third — but it is counted and said out
loud rather than quietly dropped. loud rather than quietly dropped.
Pictures inside comments are shown as pictures. Reddit writes them as a token rather
than an address — a Giphy id, a subreddit emote, or an image uploaded to the comment —
and all three are looked up in the comment's own metadata to find the real file. An
image address someone simply pasted is shown too, which on Reddit is how most of them
arrive. All of it goes through the same `/m/` proxy as everything else, so reading a
comment thread never has your browser talking to Reddit.
Media never gets linked straight at a CDN. Instagram and TikTok reject requests without Media never gets linked straight at a CDN. Instagram and TikTok reject requests without
a matching `Referer` (and sometimes cookies), and proxying keeps your browser from a matching `Referer` (and sometimes cookies), and proxying keeps your browser from
talking to the platform at all. Every asset is registered under an opaque `/m/<id>` and talking to the platform at all. Every asset is registered under an opaque `/m/<id>` and
+2 -2
View File
@@ -1,12 +1,12 @@
{ {
"name": "antisocial", "name": "antisocial",
"version": "0.1.0", "version": "1.2.1",
"lockfileVersion": 3, "lockfileVersion": 3,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "antisocial", "name": "antisocial",
"version": "0.1.0", "version": "1.2.1",
"license": "UNLICENSED", "license": "UNLICENSED",
"dependencies": { "dependencies": {
"@fastify/static": "10.1.3", "@fastify/static": "10.1.3",
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "antisocial", "name": "antisocial",
"version": "0.1.0", "version": "1.2.1",
"private": true, "private": true,
"description": "Reads social posts back to you without the app.", "description": "Reads social posts back to you without the app.",
"license": "UNLICENSED", "license": "UNLICENSED",
+19
View File
@@ -556,3 +556,22 @@ main { max-width: 680px; margin: 0 auto; }
max-height: 45dvh; max-height: 45dvh;
min-height: 0; min-height: 0;
} }
/* A picture someone put in a comment. Capped hard: it is a remark inside a
conversation, not the thing the page is about. */
.c__img {
display: block;
max-width: min(100%, 420px);
max-height: 40vh;
max-height: 40dvh;
width: auto;
height: auto;
margin: 8px 0;
border: 1px solid var(--line);
border-radius: 8px;
background: color-mix(in srgb, var(--ink) 4%, transparent);
}
/* A lone image is the whole comment more often than not, so it should not
carry a paragraph's worth of space above it as well as its own. */
.c__body > p:first-child > .c__img:first-child { margin-top: 2px; }
+33 -1
View File
@@ -62,6 +62,7 @@ type Link = {
type CommentData = { type CommentData = {
author?: string; author?: string;
body?: string; body?: string;
media_metadata?: Record<string, MediaMeta>;
created_utc?: number; created_utc?: number;
score?: number; score?: number;
score_hidden?: boolean; score_hidden?: boolean;
@@ -188,6 +189,37 @@ function bodyOf(link: Link): string | undefined {
return undefined; return undefined;
} }
/** The whole of `![...](...)`, with the target captured. */
const INLINE_IMAGE = /!\[([^\]\n]*)\]\(([^)\s]+)\)/g;
/**
* Point a comment's inline images at something fetchable.
*
* Reddit writes them as `![gif](giphy|Ve7wX45)`, `![img](emote|t5_2th52|4358)`
* or `![img](jo8gf0ca92zd1)` — a token rather than an address. In every case
* the token is a key in that same comment's `media_metadata`, which is where
* the real URL is, so one lookup covers all three and none of them needs
* naming here.
*
* A target that is already an address is not a key, so it falls through
* untouched.
*/
export function resolveInlineImages(
body: string,
meta: Record<string, MediaMeta> | undefined,
): string {
if (!meta) return body;
return body.replace(INLINE_IMAGE, (whole, alt: string, token: string) => {
const entry = meta[token];
if (!entry || entry.status !== 'valid') return whole;
// An animated one has both; the GIF plays in an `<img>` on its own, which
// an MP4 does not.
const url = entry.s?.gif ?? entry.s?.u;
return url ? `![${alt}](${url})` : whole;
});
}
export function commentsFrom(listing: Listing<CommentData> | undefined): { export function commentsFrom(listing: Listing<CommentData> | undefined): {
comments: Comment[]; comments: Comment[];
more: number; more: number;
@@ -209,7 +241,7 @@ export function commentsFrom(listing: Listing<CommentData> | undefined): {
comments.push({ comments.push({
author: authorName(data.author), author: authorName(data.author),
...(data.body ? { text: data.body } : {}), ...(data.body ? { text: resolveInlineImages(data.body, data.media_metadata) } : {}),
...(isoFrom(data.created_utc) ? { postedAt: isoFrom(data.created_utc) } : {}), ...(isoFrom(data.created_utc) ? { postedAt: isoFrom(data.created_utc) } : {}),
// Reddit hides the score on a new comment so an early downvote cannot // Reddit hides the score on a new comment so an early downvote cannot
// steer the rest. Showing a placeholder 1 would be a lie. // steer the rest. Showing a placeholder 1 would be a lie.
+62 -22
View File
@@ -13,6 +13,20 @@ import { escapeHtml, raw, type Raw } from './html.ts';
const REDDIT = 'https://www.reddit.com'; const REDDIT = 'https://www.reddit.com';
/**
* How an image in a comment becomes markup.
*
* Supplied by the caller rather than decided here, because the address has to
* go through the media proxy and this file knows nothing about that. Without
* one an image degrades to a link, which is what it was before.
*/
export type ImageRenderer = (url: string, alt: string) => string;
/** Worth showing as a picture rather than as a link to one. */
function looksLikeImage(url: string): boolean {
return /\.(jpe?g|png|gif|webp|avif)(\?|$)/i.test(url);
}
/** Absolute http(s) only. `javascript:` and friends never become links. */ /** Absolute http(s) only. `javascript:` and friends never become links. */
function safeHref(url: string): string | undefined { function safeHref(url: string): string | undefined {
try { try {
@@ -46,44 +60,69 @@ function trimUrlTail(url: string): string {
const INLINE = new RegExp( const INLINE = new RegExp(
[ [
'`([^`\\n]+)`', // 1 code '`([^`\\n]+)`', // 1 code
'\\[([^\\]\\n]+)\\]\\(([^)\\s]+)\\)', // 2 label, 3 href // Before the link rule, or the `[` of an image matches as a link and
'\\*\\*([^*\\n]+)\\*\\*', // 4 strong // leaves its `!` behind as text.
'~~([^~\\n]+)~~', // 5 strike '!\\[([^\\]\\n]*)\\]\\(([^)\\s]+)\\)', // 2 alt, 3 src
'(?<![\\w*])\\*([^*\\n]+)\\*(?![\\w*])', // 6 em with asterisks '\\[([^\\]\\n]+)\\]\\(([^)\\s]+)\\)', // 4 label, 5 href
'(?<![\\w_])_([^_\\n]+)_(?![\\w_])', // 7 em with underscores '\\*\\*([^*\\n]+)\\*\\*', // 6 strong
'(https?://[^\\s<>]+)', // 8 bare url '~~([^~\\n]+)~~', // 7 strike
'(?<![\\w/])(/?[ru]/[A-Za-z0-9_][A-Za-z0-9_-]{1,30})', // 9 r/sub and u/name '(?<![\\w*])\\*([^*\\n]+)\\*(?![\\w*])', // 8 em with asterisks
'(?<![\\w_])_([^_\\n]+)_(?![\\w_])', // 9 em with underscores
'(https?://[^\\s<>]+)', // 10 bare url
// 11 the same thing with the scheme left off, which is how people type
// them. Narrow on purpose: a host, a path, and an image extension. Prose
// is full of dotted words, and `src/render/post.ts` must not become a
// link to a website in Tonga.
'(?<![\\w@/.])((?:[a-z0-9-]+\\.)+[a-z]{2,}/[^\\s<>]*\\.(?:jpe?g|png|gif|webp|avif)(?:\\?[^\\s<>]*)?)',
'(?<![\\w/])(/?[ru]/[A-Za-z0-9_][A-Za-z0-9_-]{1,30})', // 12 r/sub and u/name
].join('|'), ].join('|'),
'g', 'g',
); );
/** One line of body text: escaped, with the inline constructs put back. */ /** One line of body text: escaped, with the inline constructs put back. */
function inline(text: string): string { function inline(text: string, image?: ImageRenderer): string {
let out = ''; let out = '';
let cursor = 0; let cursor = 0;
for (const match of text.matchAll(INLINE)) { for (const match of text.matchAll(INLINE)) {
const [whole, code, label, href, strong, strike, emStar, emScore, url, subOrUser] = match; const [whole, code, alt, src, label, href, strong, strike, emStar, emScore, url, schemeless,
subOrUser] = match;
out += escapeHtml(text.slice(cursor, match.index)); out += escapeHtml(text.slice(cursor, match.index));
cursor = match.index + whole.length; cursor = match.index + whole.length;
if (code !== undefined) { if (code !== undefined) {
out += `<code>${escapeHtml(code)}</code>`; out += `<code>${escapeHtml(code)}</code>`;
} else if (src !== undefined) {
const safe = safeHref(src);
// Without a renderer to place it, an image is still a link to one.
out += safe ? (image ? image(safe, alt ?? '') : anchor(safe, alt || safe)) : escapeHtml(whole);
} else if (label !== undefined && href !== undefined) { } else if (label !== undefined && href !== undefined) {
const safe = safeHref(href); const safe = safeHref(href);
out += safe ? anchor(safe, label) : escapeHtml(whole); out += safe ? anchor(safe, label) : escapeHtml(whole);
} else if (strong !== undefined) { } else if (strong !== undefined) {
out += `<strong>${inline(strong)}</strong>`; out += `<strong>${inline(strong, image)}</strong>`;
} else if (strike !== undefined) { } else if (strike !== undefined) {
out += `<del>${inline(strike)}</del>`; out += `<del>${inline(strike, image)}</del>`;
} else if (emStar !== undefined || emScore !== undefined) { } else if (emStar !== undefined || emScore !== undefined) {
out += `<em>${inline(emStar ?? emScore ?? '')}</em>`; out += `<em>${inline(emStar ?? emScore ?? '', image)}</em>`;
} else if (url !== undefined) { } else if (url !== undefined) {
const trimmed = trimUrlTail(url); const trimmed = trimUrlTail(url);
const safe = safeHref(trimmed); const safe = safeHref(trimmed);
out += safe const tail = escapeHtml(url.slice(trimmed.length));
? anchor(safe, trimmed.replace(/^https?:\/\/(www\.)?/, '')) + escapeHtml(url.slice(trimmed.length)) if (!safe) {
: escapeHtml(whole); out += escapeHtml(whole);
} else if (image && looksLikeImage(trimmed)) {
// People paste the address of a picture and mean the picture. On
// Reddit that is most of what an image in a comment even is.
out += image(safe, '') + tail;
} else {
out += anchor(safe, trimmed.replace(/^https?:\/\/(www\.)?/, '')) + tail;
}
} else if (schemeless !== undefined) {
// Assumed https: every host that serves these redirects to it anyway,
// and a picture is the one thing worth guessing a scheme for.
const safe = safeHref(`https://${schemeless}`);
out += safe ? (image ? image(safe, '') : anchor(safe, schemeless)) : escapeHtml(whole);
} else if (subOrUser !== undefined) { } else if (subOrUser !== undefined) {
const path = subOrUser.startsWith('/') ? subOrUser : `/${subOrUser}`; const path = subOrUser.startsWith('/') ? subOrUser : `/${subOrUser}`;
out += anchor(`${REDDIT}${path}`, subOrUser); out += anchor(`${REDDIT}${path}`, subOrUser);
@@ -101,7 +140,7 @@ const NUMBERED = /^\s{0,3}\d+[.)]\s+/;
* from the reply to it, and a comment that loses that separation reads as * from the reply to it, and a comment that loses that separation reads as
* though the commenter said both halves. * though the commenter said both halves.
*/ */
function blocks(lines: string[]): string { function blocks(lines: string[], image?: ImageRenderer): string {
let out = ''; let out = '';
let at = 0; let at = 0;
@@ -139,7 +178,7 @@ function blocks(lines: string[]): string {
if (/^\s*>/.test(line)) { if (/^\s*>/.test(line)) {
const body = takeWhile((l) => /^\s*>/.test(l)); const body = takeWhile((l) => /^\s*>/.test(l));
// Nested, so a quote of a quote keeps its shape. // Nested, so a quote of a quote keeps its shape.
out += `<blockquote>${blocks(body.map((l) => l.replace(/^\s*>\s?/, '')))}</blockquote>`; out += `<blockquote>${blocks(body.map((l) => l.replace(/^\s*>\s?/, '')), image)}</blockquote>`;
continue; continue;
} }
@@ -148,20 +187,21 @@ function blocks(lines: string[]): string {
const pattern = ordered ? NUMBERED : BULLET; const pattern = ordered ? NUMBERED : BULLET;
const items = takeWhile((l) => pattern.test(l)); const items = takeWhile((l) => pattern.test(l));
const tag = ordered ? 'ol' : 'ul'; const tag = ordered ? 'ol' : 'ul';
out += `<${tag}>${items.map((l) => `<li>${inline(l.replace(pattern, ''))}</li>`).join('')}</${tag}>`; out += `<${tag}>${items.map((l) => `<li>${inline(l.replace(pattern, ''), image)}</li>`).join('')}</${tag}>`;
continue; continue;
} }
const paragraph = takeWhile( const paragraph = takeWhile(
(l) => l.trim() !== '' && !/^\s*>/.test(l) && !BULLET.test(l) && !NUMBERED.test(l) && !/^\s*```/.test(l), (l) => l.trim() !== '' && !/^\s*>/.test(l) && !BULLET.test(l) && !NUMBERED.test(l) && !/^\s*```/.test(l),
); );
out += `<p>${paragraph.map((l) => inline(l)).join('<br>')}</p>`; out += `<p>${paragraph.map((l) => inline(l, image)).join('<br>')}</p>`;
} }
return out; return out;
} }
/** Comment text, as safe markup. */ /** Comment text, as safe markup. `image` places the pictures; without it
export function renderMarkdown(text: string): Raw { * they stay links, which is what they were before. */
return raw(blocks(text.replace(/\r\n?/g, '\n').split('\n'))); export function renderMarkdown(text: string, image?: ImageRenderer): Raw {
return raw(blocks(text.replace(/\r\n?/g, '\n').split('\n'), image));
} }
+19 -1
View File
@@ -145,6 +145,22 @@ function shortWhen(postedAt: string | undefined): Raw {
})}</time>`; })}</time>`;
} }
/**
* A picture inside a comment.
*
* Through the proxy like everything else a comment full of `preview.redd.it`
* addresses would otherwise have the viewer's browser fetch every one of them
* straight from Reddit, which is the thing this whole app exists to avoid.
*
* No dimensions to reserve space with: the size is in the payload but not in
* the Markdown, so these are capped by the stylesheet and load at whatever
* shape they are.
*/
function renderCommentImage(url: string, alt: string): string {
return html`<img class="c__img" src="${proxyUrlFor({ url })}" alt="${alt}" loading="lazy" decoding="async">`
.value;
}
/** Everything hanging off a comment, however deep. Shown only while it is /** Everything hanging off a comment, however deep. Shown only while it is
* collapsed, so what a fold is hiding is never a mystery. */ * collapsed, so what a fold is hiding is never a mystery. */
function descendantsOf(comment: Comment): number { function descendantsOf(comment: Comment): number {
@@ -179,7 +195,9 @@ function renderComment(comment: Comment, depth: number): Raw {
}</span>` }</span>`
: ''} : ''}
</summary> </summary>
${comment.text ? html`<div class="c__body">${renderMarkdown(comment.text)}</div>` : ''} ${comment.text
? html`<div class="c__body">${renderMarkdown(comment.text, renderCommentImage)}</div>`
: ''}
${comment.replies.length || comment.moreReplies ${comment.replies.length || comment.moreReplies
? html`<div class="c__replies"> ? html`<div class="c__replies">
${comment.replies.map((reply) => renderComment(reply, depth + 1))} ${comment.replies.map((reply) => renderComment(reply, depth + 1))}
+89
View File
@@ -1,9 +1,16 @@
import assert from 'node:assert/strict'; import assert from 'node:assert/strict';
import { test } from 'node:test'; import { test } from 'node:test';
import { escapeHtml } from '../src/render/html.ts';
import { renderMarkdown } from '../src/render/markdown.ts'; import { renderMarkdown } from '../src/render/markdown.ts';
const md = (text: string): string => String(renderMarkdown(text)); const md = (text: string): string => String(renderMarkdown(text));
/** Stands in for the real one, which proxies. Escapes the way that one does:
* placing the image is the renderer's job, and so is making it safe. */
const img = (url: string, alt: string): string =>
`<img src="${escapeHtml(url)}" alt="${escapeHtml(alt)}">`;
const mdi = (text: string): string => String(renderMarkdown(text, img));
test('markup a commenter typed is text, not markup', () => { test('markup a commenter typed is text, not markup', () => {
const out = md('<script>alert(1)</script> & "quoted"'); const out = md('<script>alert(1)</script> & "quoted"');
assert.ok(!out.includes('<script>')); assert.ok(!out.includes('<script>'));
@@ -64,3 +71,85 @@ test('a single newline inside a paragraph is a line break, a blank line is a new
assert.equal(md('one\ntwo'), '<p>one<br>two</p>'); assert.equal(md('one\ntwo'), '<p>one<br>two</p>');
assert.equal(md('one\n\ntwo'), '<p>one</p><p>two</p>'); assert.equal(md('one\n\ntwo'), '<p>one</p><p>two</p>');
}); });
test('an image is a picture when there is something to place it with', () => {
assert.equal(mdi('![a cat](https://i.redd.it/x.jpg)'),
'<p><img src="https://i.redd.it/x.jpg" alt="a cat"></p>');
});
test('an image degrades to a link when there is not', () => {
const out = md('![a cat](https://i.redd.it/x.jpg)');
assert.ok(out.includes('<a href="https://i.redd.it/x.jpg"'));
assert.ok(!out.includes('<img'));
});
test("an image's `!` is not left behind as text", () => {
// The link rule would otherwise match from the `[` and strand the bang.
assert.ok(!mdi('![](https://i.redd.it/x.png)').includes('!'));
});
test('a pasted image address becomes the picture, not a link to it', () => {
// Which is how most images in a Reddit comment arrive.
const out = mdi('look\n\nhttps://preview.redd.it/abc.jpeg?width=1274&s=deadbeef');
assert.ok(out.includes('<img src="https://preview.redd.it/abc.jpeg?width=1274&amp;s=deadbeef"'));
assert.ok(!out.includes('<a href'));
});
test('a link that is not an image is still a link', () => {
const out = mdi('see https://example.com/article');
assert.ok(out.includes('<a href="https://example.com/article"'));
assert.ok(!out.includes('<img'));
});
test('a sentence after a pasted image keeps its punctuation out of the address', () => {
const out = mdi('here https://i.redd.it/x.jpg.');
assert.ok(out.includes('src="https://i.redd.it/x.jpg"'), out);
assert.ok(out.endsWith('.</p>'), out);
});
test('only http and https become pictures', () => {
const out = mdi('![x](javascript:alert(1))');
assert.ok(!out.includes('<img'));
assert.ok(out.includes('![x]'));
});
test("an image's alt text is escaped like anything else a stranger wrote", () => {
const out = mdi('![" onerror=alert(1)](https://i.redd.it/x.jpg)');
assert.ok(!out.includes('onerror=alert(1)>'), out);
assert.ok(out.includes('&quot;'));
});
test('images inside a quote are still placed', () => {
assert.match(mdi('> ![](https://i.redd.it/x.gif)'), /<blockquote><p><img/);
});
test('an image address typed without a scheme is still the picture', () => {
// Which is how people type them: no https, straight from the address bar.
const out = mdi('preview.redd.it/lz4drsqh0clh1.jpeg?width=1290&s=b27e');
assert.ok(out.includes('<img src="https://preview.redd.it/lz4drsqh0clh1.jpeg?width=1290&amp;s=b27e"'), out);
});
test('a schemeless image address mid-sentence keeps the sentence', () => {
const out = mdi('look at i.redd.it/x.png nice one');
assert.ok(out.startsWith('<p>look at <img'), out);
assert.ok(out.endsWith(' nice one</p>'), out);
});
test('prose full of dots and slashes is not mistaken for an address', () => {
// The reason this rule insists on a host, a path and an image extension.
for (const text of [
'the file is at src/render/post.ts',
'see node_modules/foo/bar.js',
'a path like ./images/cat.jpg',
'C:/Users/x/cat.png',
'email [email protected]/nope.jpg',
'version 1.2.3/4.png',
]) {
assert.ok(!mdi(text).includes('<img'), `treated as an image: ${text}`);
}
});
test('a schemeless address that is not an image is left alone', () => {
// Guessing a scheme is worth it for a picture and not for prose.
assert.equal(mdi('example.com/article'), '<p>example.com/article</p>');
});
+80 -1
View File
@@ -1,6 +1,6 @@
import assert from 'node:assert/strict'; import assert from 'node:assert/strict';
import { test } from 'node:test'; import { test } from 'node:test';
import { commentsFrom, mediaFromLink, toPost, treeFromDepths } from '../src/platforms/reddit.ts'; import { commentsFrom, mediaFromLink, resolveInlineImages, toPost, treeFromDepths } from '../src/platforms/reddit.ts';
import { reddit } from '../src/platforms/reddit.ts'; import { reddit } from '../src/platforms/reddit.ts';
import { originalUrlFor } from '../src/platforms/index.ts'; import { originalUrlFor } from '../src/platforms/index.ts';
import { fixture } from './helpers.ts'; import { fixture } from './helpers.ts';
@@ -168,3 +168,82 @@ test('the page fallback rebuilds nesting from the depth on each comment', () =>
test('a comment the page gave no text for is dropped rather than shown empty', () => { test('a comment the page gave no text for is dropped rather than shown empty', () => {
assert.deepEqual(treeFromDepths([{ depth: 0, author: 'a', score: 1, created: '', text: '' }]), []); assert.deepEqual(treeFromDepths([{ depth: 0, author: 'a', score: 1, created: '', text: '' }]), []);
}); });
// Real shapes, captured from comments carrying each kind.
const GIPHY = {
'giphy|Ve7wX45gaOFmw8eeEM': {
status: 'valid',
e: 'AnimatedImage',
m: 'image/gif',
s: {
y: 200,
x: 304,
gif: 'https://external-preview.redd.it/CTp8.gif?width=304&height=200&s=b0e9',
mp4: 'https://external-preview.redd.it/CTp8.gif?width=304&height=200&format=mp4&s=a389',
},
},
};
const UPLOAD = {
jo8gf0ca92zd1: {
status: 'valid',
e: 'Image',
m: 'image/jpeg',
s: { y: 1270, x: 1274, u: 'https://preview.redd.it/jo8gf0ca92zd1.jpeg?width=1274&s=c226' },
},
};
test('a giphy comment points at the gif rather than at a token', () => {
// `![gif](giphy|ID)` is not an address, and renders as nothing at all until
// it is looked up in the comment's own media_metadata.
assert.equal(
resolveInlineImages('![gif](giphy|Ve7wX45gaOFmw8eeEM)', GIPHY),
'![gif](https://external-preview.redd.it/CTp8.gif?width=304&height=200&s=b0e9)',
);
});
test('the animated form takes the gif, which plays on its own', () => {
const out = resolveInlineImages('![gif](giphy|Ve7wX45gaOFmw8eeEM)', GIPHY);
assert.ok(out.includes('.gif?'), out);
assert.ok(!out.includes('format=mp4'), 'an mp4 would need a player to move');
});
test('an uploaded image resolves through the same lookup', () => {
// Giphy, emotes and uploads are all a token that is a key in the same map,
// so none of them needs naming.
assert.equal(
resolveInlineImages('![img](jo8gf0ca92zd1)', UPLOAD),
'![img](https://preview.redd.it/jo8gf0ca92zd1.jpeg?width=1274&s=c226)',
);
});
test('a target that is already an address is left alone', () => {
const already = '![a](https://i.redd.it/x.jpg)';
assert.equal(resolveInlineImages(already, UPLOAD), already);
assert.equal(resolveInlineImages(already, undefined), already);
});
test('a token with no entry, or a broken one, is not invented', () => {
assert.equal(resolveInlineImages('![gif](giphy|missing)', GIPHY), '![gif](giphy|missing)');
assert.equal(
resolveInlineImages('![img](gone)', { gone: { status: 'failed', e: 'Image' } }),
'![img](gone)',
);
});
test('inline images survive the walk into the comment tree', () => {
const { comments } = commentsFrom({
data: {
children: [{
kind: 't1',
data: {
author: 'a',
body: 'ha ![gif](giphy|Ve7wX45gaOFmw8eeEM)',
media_metadata: GIPHY,
replies: '',
},
}],
},
});
assert.match(comments[0]?.text ?? '', /external-preview\.redd\.it/);
});
+18
View File
@@ -237,3 +237,21 @@ test('markup in a quoted post is escaped like any other stranger\'s text', () =>
assert.ok(!html_.includes('<script>alert(1)</script>')); assert.ok(!html_.includes('<script>alert(1)</script>'));
assert.ok(!html_.includes('<img src=x')); assert.ok(!html_.includes('<img src=x'));
}); });
test('a picture in a comment is proxied, and its alt text cannot break out', () => {
const page = renderPost(redditPost({
comments: [{
author: 'u/a',
text: '![" onerror=alert(1) x="](https://preview.redd.it/x.jpeg?s=abc)\n\nhttps://i.redd.it/y.png',
replies: [],
}],
}));
assert.ok(!page.includes('https://preview.redd.it/x.jpeg'), 'upstream URLs must not reach the page');
assert.ok(!page.includes('https://i.redd.it/y.png'), 'a pasted address is proxied too');
assert.equal((page.match(/<img class="c__img" src="\/m\//g) ?? []).length, 2);
// The payload survives as text inside the attribute, which is the point:
// its quotes are neutered, so it cannot close `alt="` and become markup.
assert.ok(page.includes('alt="&quot; onerror=alert(1) x=&quot;"'), 'alt text must be escaped');
assert.ok(!/alt="" onerror/.test(page), 'the attribute must not be closable');
});