Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

914 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

⬇️ abx-dl

A simple all-in-one CLI tool to auto-detect and download everything available from a URL.

uvx abx-dl --plugins=title,wget 'https://example.com'
set -Eeuo pipefail
output_dir="$(mktemp -d)"
image="${ABXDL_IMAGE:-archivebox/abx-dl:latest}"
trap 'rm -rf "$output_dir"' EXIT
docker run --rm \
  --env OUTPUT_UID="$(id -u)" \
  --env OUTPUT_GID="$(id -g)" \
  --volume "$output_dir:/out" \
  --entrypoint bash \
  "$image" \
  -c 'set -Eeuo pipefail
cleanup() { chown -R "$OUTPUT_UID:$OUTPUT_GID" /out; }
trap cleanup EXIT
/venv/bin/abx-dl "$@"' \
  -- --no-install --max-urls=1 --plugins=title,wget 'https://example.com'
test -s "$output_dir/index.jsonl"
test -s "$output_dir/title/title.txt"
test -s "$output_dir/wget/example.com/index.html"
grep -q 'Example Domain' "$output_dir/title/title.txt"
grep -q 'Example Domain' "$output_dir/wget/example.com/index.html"

✨ Ever wish you could yt-dlp, gallery-dl, wget, curl, puppeteer, etc. all in one command?

abx-dl is an all-in-one CLI tool for downloading URLs "by any means necessary".

It's useful for scraping, downloading, OSINT, digital preservation, and more. abx-dl provides a simpler one-shot CLI interface to the ArchiveBox plugin ecosystem.

Screenshot 2026-03-11 at 6 53 03β€―AM

🍜 What does it save?

abx-dl --plugins=wget,title,screenshot,pdf,readability 'https://example.com'

abx-dl runs all plugins by default (and auto installs dependencies). You can specify --plugins=wget,favicon,title or filters like --output=html,pdf,ico,text/ to limit plugin selection.

  • HTML, JS, CSS, images, etc. rendered with a headless browser
  • title, favicon, headers, outlinks, and other metadata
  • audio, video, subtitles, playlists, comments
  • snapshot of the page as a PDF, screenshot, and Singlefile HTML
  • article text, git source code
  • and much more...

🧩 How does it work?

abx-dl uses the Plugin Library (shared with ArchiveBox) to run a collection of downloading and scraping tools.

Plugins are loaded from the installed abx-plugins package (or from ABX_PLUGINS_DIR if you override it) and execute in distinct phases:

  1. Install phase runner reads plugins config.json: required_binaries and emits BinaryRequestEvents for abxpkg.binary_service.BinaryService, which resolves or installs binaries using built-in providers such as env, pip, npm, brew, apt, cargo, and browser-specific providers. BinaryCacheService and the abx-dl cache backend then project resolved state into derived.env.
  2. CrawlSetup hooks (on_CrawlSetup__*) launch/configure expensive crawl-scoped processes like chrome, or trigger side effects. background hooks use their first stdout line as the readiness boundary and emit no stdout JSONL records.
  3. Snapshot hooks (on_Snapshot__*) run per URL to extract content. background hooks use their first stdout line as the readiness boundary; JSONL records after that are ArchiveResult, Snapshot, and Tag.

βš™οΈ Configuration

Configuration is handled via environment variables plus a user config file under the platformdirs user config path (<user-config>/abx/config.env). Runtime-derived cache entries such as resolved binary paths are stored separately in <user-config>/abx/derived.env:

abx-dl config                        # show all config (global + per-plugin)
abx-dl config --get WGET_TIMEOUT     # get a specific value
abx-dl config --set TIMEOUT=120      # set persistently (resolves aliases)

Output is grouped by section:

# GLOBAL
TIMEOUT=60
USER_AGENT="Mozilla/5.0 ..."
...

# plugins/wget
WGET_BINARY="wget"
WGET_TIMEOUT=60
...

# plugins/chrome
CHROME_BINARY="chromium"
...

Common options:

  • TIMEOUT=60 - default timeout for hooks
  • USER_AGENT - default user agent string
  • {PLUGIN}_BINARY - path or name of the binary to use (e.g. WGET_BINARY=wget or CHROME_BINARY=/usr/bin/chromium)
  • {PLUGIN}_ENABLED=True/False - enable/disable specific plugins
  • {PLUGIN}_TIMEOUT=120 - per-plugin timeout overrides

Aliases are automatically resolved (e.g. --set USE_WGET=false saves as WGET_ENABLED=false).

One-off config is easy via env vars or CLI args:

env \
  TIMEOUT=120 \
  WGET_TIMEOUT=120 \
  abx-dl \
    --dir=./config-example \
    --plugins=title,wget \
    --timeout=90 \
    'https://example.com'



πŸ“¦ Install

uv tool install abx-dl
abx-dl version
uvx abx-dl version
abx-dl install wget title

πŸ”  Usage

abx-dl --plugins=title,wget --dir=./downloads --timeout=120 'https://example.com'
# Default command - a bare URL archives with all enabled plugins:
abx-dl 'https://example.com'

# Select plugins by output type (mimetypes, categories, or file extensions):
abx-dl --output=html,pdf,video/ 'https://example.com'
abx-dl -o text -o image -o mp4 'https://example.com'

# Limit work to a subset of plugins by name:
abx-dl --plugins=wget,title,screenshot,pdf 'https://example.com'

# Skip auto-installing missing dependencies (emit warnings instead):
abx-dl --no-install 'https://example.com'

# Specify output directory (default is current working dir):
abx-dl --dir=./downloads 'https://example.com'

# Set timeout:
abx-dl --timeout=120 'https://example.com'

Commands

abx-dl <url>                              # Download URL (default shorthand)
abx-dl plugins                            # Check + show info for all plugins
abx-dl plugins wget ytdlp git             # Check + show info for specific plugins
abx-dl install wget ytdlp git             # Pre-install plugin dependencies
abx-dl config                             # Show all config values
abx-dl config --get TIMEOUT               # Get a specific config value
abx-dl config --set TIMEOUT=120           # Set a config value persistently

Installing Dependencies

Many plugins require external binaries (e.g., wget, chrome, yt-dlp, single-file).

By default, abx-dl lazily installs missing dependencies as needed when you download a URL. Use --no-install to skip plugins with missing dependencies instead. install runs only the pre-run dependency pipeline (required_binaries β†’ BinaryRequestEvent β†’ BinaryEvent) without starting crawl setup or snapshot extraction:

abx-dl install wget title
abx-dl plugins wget title
abx-dl 'https://example.com'              # auto-installs missing deps on-the-fly
abx-dl --no-install 'https://example.com' # skips plugins with missing deps and emits warnings
abx-dl install wget singlefile ytdlp      # installs dependencies for specific plugins only
abx-dl plugins                            # checks which dependencies are available/missing

Every hook executable declares its dependencies in an abxpkg run --script --deps-from=... shebang. Compatible host binaries are selected first and projected into ABXPKG_LIB_DIR/env/bin; otherwise the configured managed provider installs and projects the dependency. The explicit install command and --no-install dependency check use the same abxpkg resolution path.

The normal runtime flow is:

  • CrawlEvent (internal lifecycle root)
  • CrawlSetupEvent β†’ plugin on_CrawlSetup__* hooks
  • CrawlStartEvent β†’ SnapshotEvent
  • SnapshotEvent β†’ plugin on_Snapshot__* hooks
  • SnapshotCleanupEvent / CrawlCleanupEvent

Hook output contract:

  • hook dependencies are driven by plugin required_binaries and resolved by each hook's abxpkg shebang
  • on_CrawlSetup__* background hooks emit a first stdout readiness line, but no stdout JSONL records
  • on_Snapshot__* background hooks emit a first stdout readiness line; hook JSONL records after that are only ArchiveResult, Snapshot, and Tag
  • the TUI and services consume structured events derived from those hook records

Dependencies are installed to <user-config>/abx/lib/{arch}/ using the appropriate package manager:

  • pip packages β†’ <user-config>/abx/lib/{arch}/pip/venv/
  • npm packages β†’ <user-config>/abx/lib/{arch}/npm/
  • brew/apt packages β†’ system locations

You can override the install location with ABXPKG_LIB_DIR=/path/to/lib abx-dl install wget.




Output Structure

By default, abx-dl writes results into the current working directory. Each run creates an index.jsonl manifest plus one subdirectory per plugin that produced output. If you want to keep runs isolated, cd into a scratch directory first or pass --dir=/path/to/run.

mkdir -p /tmp/abx-run && cd /tmp/abx-run
uvx --from abx-dl abx-dl --plugins=title,wget 'https://example.com'
./
β”œβ”€β”€ index.jsonl             # Snapshot metadata and results (JSONL format)
β”œβ”€β”€ title/
β”‚   └── title.txt
β”œβ”€β”€ favicon/
β”‚   └── favicon.ico
β”œβ”€β”€ screenshot/
β”‚   └── screenshot.png
β”œβ”€β”€ pdf/
β”‚   └── output.pdf
β”œβ”€β”€ dom/
β”‚   └── output.html
β”œβ”€β”€ wget/
β”‚   └── example.com/
β”‚       └── index.html
β”œβ”€β”€ singlefile/
β”‚   └── output.html
└── ...

All Outputs

  • index.jsonl - snapshot metadata and plugin results (JSONL format, ArchiveBox-compatible)
  • title/title.txt - page title
  • favicon/favicon.ico - site favicon
  • screenshot/screenshot.png - full page screenshot (Chrome)
  • pdf/output.pdf - page as PDF (Chrome)
  • dom/output.html - rendered DOM (Chrome)
  • wget/example.com/... - mirrored site files
  • singlefile/output.html - single-file HTML snapshot
  • ... and more via plugin library ...

Available Plugins

See the abx-plugins marketplace.

Snapshot / Extraction Plugins

  • ytdlp - downloads media plus sidecars: audio, video, images/thumbnails, subtitles (.srt, .vtt), JSON metadata, and text descriptions.
  • gallerydl - downloads gallery/media sets as images, videos, JSON sidecars, text sidecars, and ZIP archives.
  • forumdl - exports forum/thread archives as JSONL, WARC, and mailbox-style message archives.
  • git - clones repository contents including text, binaries, images, audio, video, fonts, and other tracked files.
  • wget - mirrors pages and requisites as HTML, WARC, images, CSS, JavaScript, fonts, audio, and video.
  • archivedotorg - saves a Wayback Machine archive link as plain text.
  • favicon - saves site favicons and touch icons as image files.
  • modalcloser - setup helper only; no direct archive files.
  • consolelog - saves browser console events as JSONL.
  • dns - saves observed DNS activity as JSONL.
  • ssl - saves TLS certificate/connection metadata as JSONL.
  • responses - saves HTTP response metadata as JSONL and can record referenced text, images, audio, video, apps, and fonts.
  • redirects - saves redirect chains as JSONL.
  • staticfile - saves non-HTML direct file responses such as PDF, EPUB, images, audio, video, JSON, XML, CSV, ZIP, and generic binary files.
  • headers - saves main-document HTTP headers as JSON.
  • chrome - manages shared browser state and emits plain-text and JSON runtime metadata.
  • seo - saves SEO metadata such as meta tags and Open Graph fields as JSON.
  • accessibility - saves the browser accessibility tree as JSON.
  • infiniscroll - page-expansion helper only; no direct archive files.
  • claudechrome - saves Claude-computer-use interaction results as JSON plus PNG screenshots.
  • singlefile - saves a full self-contained page snapshot as HTML.
  • screenshot - saves rendered page screenshots as PNG.
  • pdf - saves rendered pages as PDF.
  • dom - saves fully rendered DOM output as HTML.
  • title - saves the final page title as plain text.
  • readability - extracts article HTML, plain text, and JSON metadata.
  • defuddle - extracts cleaned article HTML, plain text, and JSON metadata.
  • mercury - extracts article HTML, plain text, and JSON metadata.
  • claudecodeextract - generates cleaned Markdown from other extractor outputs.
  • htmltotext - converts archived HTML into plain text.
  • trafilatura - extracts article content as plain text, Markdown, HTML, CSV, JSON, and XML/TEI.
  • papersdl - downloads academic papers as PDF.
  • parse_html_urls - emits discovered links from HTML as JSONL records.
  • parse_txt_urls - emits discovered links from text files as JSONL records.
  • parse_rss_urls - emits discovered feed entry URLs from RSS/Atom as JSONL records.
  • parse_netscape_urls - emits discovered bookmark URLs from Netscape bookmark exports as JSONL records.
  • parse_jsonl_urls - emits discovered bookmark URLs from JSONL exports as JSONL records.
  • parse_dom_outlinks - emits crawlable rendered-DOM outlinks as JSONL records.
  • search_backend_sqlite - writes a searchable SQLite FTS index database.
  • search_backend_sonic - pushes content into Sonic search; no local archive files declared.
  • claudecodecleanup - writes cleanup/deduplication results as plain text.
  • hashes - writes file hash manifests as JSON.
  • and more via the abx-plugins marketplace...

AI Skill

This repo includes an abx-dl skill for coding agents that need to run the standalone ArchiveBox extractor pipeline without a full ArchiveBox install.


Architecture

abx-dl is built on these components:

  • abx_dl/plugins.py - Plugin discovery from abx-plugins or ABX_PLUGINS_DIR
  • abx_dl/executor.py - Hook execution engine with config propagation
  • abx_dl/config.py - Environment variable configuration
  • abx_dl/cli.py - Rich CLI with live progress display

Related Projects


For more advanced use with collections, parallel downloading, a Web UI + REST API, etc. See: ArchiveBox/ArchiveBox

About

⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...

Topics

Resources

Stars

Watchers

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages