96c13b8e0f7e15c2d8373737291e31e2253329ab
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ab23c2b9fe
|
Add Facebook
CI / Typecheck, test, build (pull_request) Successful in 29s
A shared Facebook link lands on a page that carries the whole post logged
out -- the caption, the author, the files, the dimensions -- in ScheduledServerJS
payloads that are plain JSON in ordinary script tags. What it does not carry
is only that post. A reel arrives with the next five reels of the feed
attached under viewer.lasso_blue_feed, a video with its related videos, and
every one of them has the same fields in the same shape as the real one.
Reading the first node with media on it gets a stranger's reel under someone
else's name, which is the same false attribution a lifted quote-post picture
used to be.
So nothing is read until the post has been picked out. partsOfPost matches the
id in the address against every node that names one; an address with no id in
it -- a pfbid permalink -- is matched on permalink_url instead. Only with
nothing to match on at all does it fall back to the route's query results,
which is still narrower than the whole payload: the page ships its entire
client configuration alongside the post, thousands of nodes carrying a name or
an id, and a plain search finds a video player setting long before it finds the
author.
An id that matches nothing is a failure rather than a best guess. Facebook
answers a link to something it no longer has by quietly serving something else
-- /watch/<id> for a video that is gone comes back as the Watch home page, feed
and all -- so the error card, which still carries the link and the copy button,
is the honest answer.
One post's pieces are spread across several payload blocks: a video post keeps
its files in one, its caption in another, its author's avatar in a third and
its timestamp in a fourth. Every claiming node is collected, not just the
first, and the author's gaps are filled only from nodes carrying the same id.
The author is whatever the payload calls the owner -- actors, owner,
video_owner, owner_as_page. `author` on a Facebook page means the author of a
comment, which sits right there in the same shape with a name and a picture of
its own.
Media is videoDeliveryLegacyFields.browser_native_hd_url with
preferred_thumbnail as its poster, photo_image or image for a picture, and
all_subattachments.nodes for a post of several -- which Facebook ships empty on
every single-picture post, so only a populated one is a carousel. The CDN is
signed and serves Range without asking for a referrer, but the assets are
proxied like everything else.
/share/{r,v,p,g} links are stubs, so the adapter follows one and hands back
where it landed: the share code says nothing about what it opens, and rdid,
share_url and fs are what the redirect leaves behind. m.facebook.com is a login
wall logged out, so the rewrite rule rebuilds on www.
Fixtures are real captures of a reel, a photo post and a video post, trimmed of
the DASH manifests and tracking blobs. The reel keeps its recommendations and
the video post keeps its comments, because those are the two things that have
to survive being read past.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
4c78ce8a27
|
Ask iOS for the playback audio session on a page with a video
CI / Typecheck, test, build (pull_request) Successful in 12s
A TikTok post played perfectly and said nothing on an iPhone. The file was not the problem: every rendition TikTok offers for it carries a full AAC track, and the page as served measures loud audio in WebKit and Chromium alike, straight through the /m/ proxy. iOS is the problem. A video playing inline gets the "ambient" audio session, which the Ring/Silent switch mutes; only going fullscreen gets the sound back. Claiming "playback" says what is true of this page -- the sound is the point, not decoration -- and the switch stops applying. Declared only where there is a video, so an ordinary text post never claims it, and the session activates when something plays rather than on load, so it interrupts nothing on a page nobody presses play on. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01B3RUMiXa6eAW3okw9nB9NF |
||
|
|
5b4378a838
|
Give a video its shape before it has any data
CI / Typecheck, test, build (pull_request) Successful in 26s
A Reddit video sat in the wrong box until you pressed play, for two reasons that looked like one. The renderer put only `aspect-ratio` on the `<video>`. A video with no data has a natural size of 300x150, and WebKit sizes a replaced element from that rather than from the ratio, so a portrait video got a squat landscape box and kept it until playback supplied real dimensions. Chromium stretch-fits instead and gets it right, which is why this only showed on Safari. A video's size before its data arrives is its poster's, so one with no poster of its own now gets an empty SVG of the right shape as a stand-in: a data URI, so it costs no request. Measured in both engines across portrait, landscape, square and small. And Reddit's videos with sound had no poster to be sized by, because `fromRedditVideo` attached the still from `preview.images` to the MP4 branch alone. The still belongs to the video, not to the format it is served in, so every post with sound showed an empty box where a silent one showed a frame. It now goes on both. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01PLkmgp1fWbA4XbarxKdRKt |
||
|
|
4669fe0b6a
|
Resolve a giphy token with no metadata to look it up in
`` is a token, not an address, and the only thing that turned it into one was a lookup in the comment's own `media_metadata`. Reddit ships plenty of comments carrying such a token and no `media_metadata` at all, and with nothing to look it up in the token itself was what the comment showed. Giphy is the one of the three token kinds whose id means something off Reddit, so that one can be resolved without the lookup. A variant name after the id is dropped: Giphy does not serve every variant of every gif, but the full one is always there. The other two still resolve only through the metadata. An emote id and an upload id name nothing outside Reddit, so with no entry for them there is still nothing to point them at. Closes #10 Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01PLkmgp1fWbA4XbarxKdRKt |
||
|
|
abf8ec317c
|
Let the "Open on" button choose a browser
CI / Typecheck, test, build (pull_request) Successful in 28s
The StopTheMadness rules are indiscriminate, which is the point, but they catch the link on the way back out as well: in the browser they are installed in, "Open on <platform>" redirects straight back here. The one button meant to reach the app is the one that cannot. Handing the address to a different browser is the way past it, and the only way to do that from a page is that browser's own URL scheme. `/` gets a picker for which one; the choice lives in that browser's localStorage, because which browsers are installed is a fact about the device and the phone's answer is not the Mac's. The scheme table is its own module so a test can pin it -- every browser spells it differently, and Edge differs between macOS and iOS. Firefox has no scheme on macOS, so it is only offered on iOS, and an unknown or schemeless choice keeps the plain link rather than producing a dead one. The markup still carries the plain https address and the script swaps it afterwards, so nothing changes without JavaScript. Once a browser is chosen the button says which, since a scheme for a browser that is not installed opens nothing and the tap would otherwise be silent. Closes #7 Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01KF5YF3iZVKbezwALap8LYd |
||
|
|
b94c43a10c
|
Open the version bump as a pull request
The bump job has never once worked. Its first step died in two seconds on
`set: Illegal option -o pipefail` -- it is the only job here that runs in a
container, and a container job is handed `sh -e {0}` rather than the bash the
runner gives its own jobs. Dash has no `pipefail`, and no `10#` either, so the
arithmetic on the line after would have gone the same way. It now asks for
bash, which node:22 carries.
That would have got the job as far as its last step, which pushed straight to
main. Nothing had ever exercised that, and it needs the token to be allowed
past whatever protects the branch -- hence the second token, VERSION_BUMP_TOKEN,
standing by for where it is not. unsupervised-scheduler has been bumping its
version on every release for a while by pushing a branch and opening a pull
request with the ordinary task token, so that is what this does now. The second
token is no longer needed, and neither is `[skip ci]`: main is never pushed, so
there is no build of the just-published image to suppress. The pull request
puts the changed package.json through a build before it lands.
scheduler builds the request body with jq because its CI image carries jq. This
one is node:22, where node is the thing that certainly is there.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01XBT25ZDzm453A8XViSRqRB
|
||
|
|
031101c382
|
Show a picture whose address was typed without a scheme
CI / Typecheck, test, build (pull_request) Successful in 26s
`preview.redd.it/lz4drsqh0clh1.jpeg?width=1290&...` was rendering as plain text -- not a picture, and not even a link. The bare-address rule has always insisted on `https://`, so anything copied out of an address bar, where the browser hides the scheme, fell through to nothing at all. That predates the inline images from the last commit; it just did not matter until pictures started being worth placing. The rule is deliberately narrow: a host with a dot, a path, and an image extension. Widening it to every schemeless address would be the obvious move and is wrong, because a comment thread is full of dotted, slashed prose -- `src/render/post.ts` would become a link to a website in Tonga, and `node_modules/foo/bar.js` a website in Jersey. Requiring the extension costs nothing here, since the only thing worth guessing a scheme for is a picture. https is assumed. Every host that serves these redirects to it anyway. Checked both ways: the four address shapes that should become pictures do, and nine pieces of ordinary prose that must not -- file paths, a Windows drive letter, a relative path, an email address followed by a filename, a version number -- still do not. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu |
||
|
|
0db18547c8
|
Show the pictures inside Reddit comments
CI / Typecheck, test, build (pull_request) Successful in 11s
A comment that was a picture rendered as either a link or, for a Giphy, the literal text ``. On r/aww that is most of the thread. Reddit writes an inline image as a token rather than an address, in three shapes: `` for a Giphy, `` for a subreddit emote, and `` for an image uploaded straight to the comment. The useful part is that all three tokens are keys in that same comment's own `media_metadata`, so this is one lookup and not three special cases. Nothing in the adapter has to know what Giphy is. The fourth shape is someone pasting the address of a picture, which on Reddit is how most images in comments actually arrive -- 117 of them against 21 Giphys in the sample I scanned. Those are shown as pictures too, decided by the file extension. A link that is not to an image stays a link. Animated ones take `s.gif` over `s.mp4` even though the MP4 is several times smaller: a GIF moves on its own in an `<img>`, and an MP4 would need a player element with autoplay, loop and muted set, for something the size of a postage stamp. Everything goes through the `/m/` proxy, like all other media. Without that a comment thread would have the reader's browser fetch dozens of files straight from Reddit, which is the one thing this whole app exists to avoid. Placing the image is the renderer's job, not the Markdown parser's, because the proxy is a render-time concern and markdown.ts knows nothing about it -- so it takes an optional `ImageRenderer` and, without one, an image stays a link exactly as before. The parser also learned `![...]` proper: the link rule was matching from the `[` and stranding the `!` as text. Verified on r/aww/comments/171dxph, which carries one Giphy and 77 pasted images: 35 render on the first page, all 35 load, all 35 through the proxy, no upstream address reaches the page, no token is left unresolved, and nothing overflows the column or scrolls the page sideways. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu |
||
|
|
503d8a8dec
|
Move the working version on after a release
package.json has sat at 0.1.0 since the first commit, through three releases, because nothing read it. That is fine right up until something does — an image label, a health endpoint, a bug report quoting a version — at which point the tree claims to be a version that shipped long ago. A new `bump` job takes the tag that was just published, works out the next patch from it, and commits that to main. After 1.2.0 the tree says 1.2.1: not a version that exists, which is the point. A build from main is then legible as "after 1.2.0" rather than as 1.2.0 itself. It sits in publish.yml rather than a workflow of its own so that it can say `needs: build`. A version that failed to publish has not been released, and moving past it would say that it had. Prereleases are skipped for the same reason -- 1.2.3-rc1 is a candidate for a version that has not shipped, so there is nothing yet to move past. The bump goes through `npm version` rather than editing the file. The version is in the lockfile too, in two places, and a tree where those disagree is worse than one that is merely out of date. Three smaller things. The patch arithmetic forces base ten, because a patch number written 08 is otherwise read as octal and kills the job. The commit carries `[skip ci]`, or pushing it starts another build of the image that was just published. And the committer is a name that is not a person at a reserved address that can never become one, so nothing here names the instance it runs on. Pushing to main needs a token that may write to the repository. The Actions task token can where the instance allows it; where it does not, setting a VERSION_BUMP_TOKEN secret overrides it. A push that is refused fails the job with both of those as the suggestion rather than a bare 403. package.json goes to 1.2.1 here, which is where the job would have left it had it existed when 1.2.0 went out. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu |
||
|
|
6325f0ff32
|
Add Reddit with its comment threads, and show quoted posts whole
CI / Typecheck, test, build (pull_request) Successful in 9s
Two changes. They share the `Segment` model, which is why they arrive
together.
## Reddit
A new adapter under /reddit, plus the thread beneath the post -- on Reddit the
conversation is usually the reason the link was shared, so a viewer that showed
only the post would be showing the wrong half.
The `.json` twin of a post URL is the post and the whole first page of comments
in one response, far better than anything the page gives up, so it is the only
layer that normally runs. Reddit refuses it to a browser it has never seen and
answers with a JavaScript challenge, which any ordinary navigation solves by
itself; the adapter navigates once and retries, and the cookie left behind
serves every later post. Below that, `shreddit-comment` elements are read from
the rendered page -- flat, each carrying its own `depth`, so `treeFromDepths`
rebuilds the nesting.
`Post.comments` is a tree rather than a flat list with depths, because folding
a comment has to take everything under it along and nesting is what makes that
free. Each comment renders as a `<details open>`, so collapsing works with the
stylesheet off and from the keyboard, and a collapsed one says how many replies
it is hiding. "Collapse all" is the only part that needs the script, so it
ships hidden and appears once the script has run. What sits behind a "load
more" is not fetched -- that is a second page and often a third -- but it is
counted and said out loud rather than quietly dropped.
Four things the payloads got wrong on the first try, each now with a fixture:
`fallback_url` is the video track alone whenever `has_audio` is true, so a post
with sound has to use `hls_url` and only a silent one gets the proxied MP4;
`scrubber_media_url` looks like a poster and is a second MP4 for the timeline
thumbnails, while the still is in `preview.images`; a gallery's pictures live
in `media_metadata` keyed and unordered, with their order only in
`gallery_data`; and `replies` is the string "" rather than an object when there
are none. Comment bodies are Markdown, rendered by a new render/markdown.ts
that escapes first and then puts back only the constructs we chose to support
-- never Reddit's own `body_html`, which would mean trusting markup a stranger
caused to be generated.
An app share link (/r/<sub>/s/<code>) is a plain 301, so one request told not
to follow it is enough. The permalink it resolves to is what the copy button
hands back, since an opaque share code is a tracking parameter by another name.
## Quoted posts
Both X and Bluesky lifted the quoted post's media out and showed it as the
quoter's own, dropping the quoted words and the quoted author entirely. A quote
of a photo post therefore rendered as somebody else's picture under the wrong
name with nothing to say so, and a quote that had a picture of its own dropped
the quoted one instead -- the two could never both appear. Half the quote posts
people share are someone answering a stranger and the other half are someone
continuing a thought from an earlier post; neither reads with only one side of
it on the page.
`Segment.quoted` now carries the whole thing -- author, words, pictures, time
and a link to it -- and renders as a post inside the post. Neither payload
carries a usable address for it: X has no permalink and it is rebuilt from the
handle and `id_str`, Bluesky has an `at://` URI nobody can open and it is
rebuilt from the handle and the record key. On Bluesky the record sits at
`embed.record` for a plain quote and at `embed.record.record` when the quoting
post has media of its own, and a quote can also point at a feed, a list or a
post since deleted, which arrive in the same slot under a different `$type` --
only `app.bsky.embed.record#viewRecord` is taken.
Two things about X's text, both visible on any post and not only a quote. It
arrives pre-escaped, so an ampersand someone typed was reaching the page as the
literal `&`; it is decoded in the adapter, where the encoding comes from,
leaving the escape-on-the-way-out rule alone. And every link is a `t.co`, which
tells the reader nothing and routes them through X's click tracker to find out
-- `entities.urls` carries the real address alongside, so it is put back. The
shortlink X staples onto the end of a quote post is dropped rather than
expanded, since the post it points at is already on the page;
`display_text_range` is where that boundary is and it keeps a link the author
put there deliberately. Its indices are UTF-16 units into the escaped text, so
the slice happens before decoding and before expanding, and splitting to
codepoints first overshoots past an emoji -- checked against a post carrying
one.
A quote is context, and context that fills the screen has stopped being
context, so a segment carrying one gives up the window-filling cap on its own
media.
## Along the way
setupMedia only ever wired `document.querySelector('.media')`, the first rail
on the page. That was already wrong for a Bluesky or Threads chain with media
in more than one post, and a quoted carousel would have hit it too. It now runs
per rail.
## Verified
108 tests, typecheck and build clean. Resolved end to end against the live
platforms: Reddit self, gallery, video, link and share-link posts; X quotes
with no media, with media on the quoted side, and with media on both; Bluesky
quotes in both embed shapes, including a ten-post chain where every post quotes
a different account and all nine quotes come back under the right name.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017nMQ2eDKnqALYhAibpTKTu
|
||
|
|
60a9468875
|
Show the author's own chain on Bluesky and Threads
CI / Typecheck, test, build (pull_request) Successful in 46s
People write in chains on both, and a link into one arrives pointing at a single post out of several. Showing only that post loses the thing that was being said. Other people's replies are a different matter: they are a conversation rather than the thing that was shared, and on a busy post there are hundreds of them. A Post is now a list of Segments instead of one body. Most platforms produce exactly one and say so through oneSegment(); the two that thread produce the whole chain, with isAnchor marking the post that was actually linked, which need not be the first. Bluesky walks parent upward and the author's own replies downward, stopping at the first post by anyone else. That needs depth and parentHeight on getPostThread, which drags the entire reply tree along -- a few hundred KB on a popular post -- because there is no way to ask the API for one author's branch. Threads is harder to read. The page ships the linked post, the author's follow-ups, other people's replies and a pile of unrelated recommendations, all as flat thread_items containers with no nesting to go on. What separates a follow-up from a stranger's reply is that a follow-up is the author replying to themselves; a reply from someone else carries the same reply_to_author with a different name on it. The first post of a chain replies to nothing at all, so it is reachable only by walking backwards from the post that answers it -- a test caught that, when linking the second post of a thread returned just the one post. Fixtures for both are real captures. The Bluesky one keeps two of every level's outside replies rather than pruning them away, because a filter is only worth testing against the thing it is supposed to exclude. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01BGkRmLfiWuJHx6tQ12EELY |
||
|
|
fb12eb6c9f
|
Slim the image, publish on version tags, drop deployment specifics
The first publish failed partway through the push with 413 Payload Too Large: one layer was bigger than the proxy in front of the registry would accept. Three changes, only one of which is that fix. Keep deployment out of a public repo. The registry, image name and credentials now come from repository variables and secrets rather than being written down here, and the docs describe how to run the thing rather than where one particular instance runs. PUBLIC_ORIGIN defaults to localhost. The 413 is a proxy limit, so the fix is pointing REGISTRY at a host the runner reaches directly; the workflow explains itself if that host is plain HTTP and the builder's daemon has not been told to allow it. Publish on version tags. A tag like 1.2.3 publishes :1.2.3, :1.2, :1 and :latest; a prerelease publishes only its exact version and leaves :latest alone. Pushes to main publish :main and :sha-<short> and no longer move :latest, so what is deployed moves when a release says so. Shrink the image from over 1.2GB to 353MB. The Playwright base image carries Firefox and WebKit, which this never launches. Installing just the browser it does launch onto a slim Node base drops two thirds of the weight, which is worth having on a Raspberry Pi even though it does not get any single layer under a proxy limit. That last change surfaced something worth naming: a headless launch resolves to Playwright's headless shell, not the full browser, so that is what every test so far has actually been running. The image now installs exactly that binary and pool.ts names the channel, so the two cannot drift apart. Verified in the container: Bluesky, Instagram, X and Threads all resolve identically on the slim image. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01BGkRmLfiWuJHx6tQ12EELY |
||
|
|
b5f9483615
|
Initial commit: read social posts back without the app
antisocial is the other half of a StopTheMadness redirect rule. Links to X, Threads, Instagram, TikTok and Bluesky get rewritten to /<prefix>/<original path>, and this resolves the post and shows the media and the words, with a badge saying where it came from and a button to copy the original URL. Every request drives a real headless Chromium, logged out, from a residential IP. One code path, and it survives markup changes better than parsing from the outside would. Extraction is layered, most structured first: the platform's own API response caught in flight, then an inline payload, then the rendered DOM, then Open Graph tags. Media is never linked straight at a CDN. Instagram and TikTok reject requests without a matching Referer and cookies, and proxying keeps the viewer's browser from talking to the platform at all. Range is forwarded so the native video scrubber can seek. HLS is the exception, since proxying it would mean rewriting playlists. TikTok sometimes answers with a slider puzzle. Rather than reporting that as a failure, the page is parked and the viewer is handed the puzzle: screenshots stream out, pointer events are replayed back. Solving it leaves the cookie in the shared browser context, so the retry is an ordinary request. A failed resolve is never a blank error page. The card carries the platform, the original URL and the copy button, so a broken adapter still leaves the link one tap away. Verified end to end against real shared links on all five platforms, in the container, including multi-image carousels, reels, TikTok short links and photo posts. 49 tests run the adapters against captured payloads with no network. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01BGkRmLfiWuJHx6tQ12EELY |