Fetches a page, discards everything that is not the article, and renders what is left as a self-contained HTML document with no scripts, no stylesheets, no webfonts and no third-party requests. The reader rule in the starter config already points at it. Extraction follows Readability: score p/pre/td/blockquote by prose weight, propagate to ancestors with decay, adjust by class and id names, scale by one minus link density, then take the winner plus sibling nodes that also read like body copy. Serialisation runs against a tag whitelist, with unlisted elements contributing their children but no tag of their own, so wrapper divs disappear. Handling for what real pages actually do: - follow meta http-equiv=refresh stubs, including inside noscript - fall back to data-src when src holds a lazy-load placeholder - take the first srcset candidate, the smallest, not the last - decode via Content-Type charset, then meta charset, then UTF-8 - resolve links against the post-redirect URL so file:// output works - ignore script and style text so a page cannot inflate its own score Output goes to a cache file named by a hash of the URL and opens in $FURST_BROWSER, $BROWSER, or the first light browser on $PATH; --html, --text, --out and --stdin cover the other uses. Restructures the repository as a workspace so the router keeps its two dependencies and its fast build.
3.5 KiB
furst
A URL router. It sits where your default browser used to, and sends each URL to
the cheapest tool that can actually handle it — mpv for video, a pager for
PDFs, a light WebKit browser for reading, and Firefox only when nothing else
will do.
Built for a Core2 Duo with 4GB of RAM, where the browser is the problem.
Why
A news page is 2–5MB across 80+ requests with 1–3MB of JavaScript to parse and
JIT. The same article extracted is ~20KB. Choosing a lighter engine buys
2–3×; not loading the payload at all buys 10–100×. furst is the dispatcher
that decides which of those you get, per URL.
Install
cargo build --release
install -Dm755 target/release/furst ~/.local/bin/furst
furst --init # writes ~/.config/furst/rules.toml, probes for a light browser
furst --install # registers furst as the system default browser
--install writes ~/.local/share/applications/furst.desktop and points
xdg-settings at it, so every link click in every application routes here.
Use
furst <url> # match a rule and exec its command
furst --explain <url> # show what would run, and why; run nothing
furst --list # show the loaded rules
--explain is the one you want when a URL goes somewhere surprising:
$ furst --explain https://youtu.be/abc123
scheme https
host youtu.be
path /abc123
-> [video] mpv --ytdl-format=bestvideo[vcodec^=avc1][height<=?720]+... https://youtu.be/abc123
[default] surf https://youtu.be/abc123
Rules
~/.config/furst/rules.toml. First matching rule wins. If its command is
missing from $PATH, furst falls through to the next matching rule, and
finally to default — which is what lets you name tools you have not written
yet and have the config stay working today.
default = ["surf", "{url}"]
[[rule]]
name = "video"
hosts = ["youtube.com", "youtu.be"]
run = ["mpv", "--ytdl-format=bestvideo[vcodec^=avc1][height<=?720]+bestaudio/best", "{url}"]
[[rule]]
name = "hn"
hosts = ["news.ycombinator.com"]
terminal = true # wrap in $TERMINAL -e
run = ["furst-hn", "{url}"]
A rule matches when every criterion it states is satisfied; a criterion is satisfied by any one of its patterns. A rule that states nothing matches everything.
| Key | Matches against |
|---|---|
schemes |
https, mailto, magnet, … |
hosts |
example.com = apex and every subdomain; *.example.com = same; =example.com = that host exactly; * = any |
paths |
path component only, glob with *, case-insensitive |
contains |
substring of the whole raw URL |
| Placeholder | |
|---|---|
{url} {url_enc} |
the URL, raw or percent-encoded |
{host} {path} {scheme} |
parsed components |
If no argument mentions {url} or {url_enc}, the URL is appended last.
Notes for old hardware
- Force H.264 for video. A Core2 handles 720p
avc1in software but stalls on VP9/AV1, which is what YouTube serves by default. That format string is doing more work than the resolution cap. - Host matching strips userinfo with
rfind('@'), sohttps://bank.example@evil.example/routes onevil.example. furstexecs the handler rather than forking, so it leaves no process behind.
Companion tools
furst-read— fetch a page, extract the article, and render it as minimal HTML with no scripts, stylesheets, or webfonts. Thereaderrule in the starter config points at it.
License
MIT