Files
goget/docs/recursive-mirror.md

7.5 KiB
Raw Permalink Blame History

Recursive Downloading & Website Mirroring

goget can recursively download linked pages and mirror entire websites for offline viewing. This document explains the two modes, their configuration, and the underlying architecture.

Quick Comparison

Feature Recursive (--recursive) Mirror (--mirror)
Depth limit By default (--max-depth N) Infinite
Link conversion Optional (--convert-links) Automatic
Page requisites Manual (--page-requisites) Automatic
robots.txt Respected by default Respected by default
Use case Download a subtree of a site Full offline archive

Recursive Mode

Recursive mode follows links from HTML pages and downloads linked resources.

# Basic recursive download, depth 3
goget --url https://example.com/docs/ --recursive --max-depth 3

# Only PDF files
goget --recursive --accept "*.pdf" --url https://example.com/docs/

# Only images and HTML
goget --recursive --accept "image/*,text/html" --url https://example.com/
flowchart TD
    HTML[Parse HTML page]
    Extract["Extract all links\n(a, img, link, script)"]
    FilterByDomain{"Domain match?"}
    FilterByPattern{"Accept pattern\nmatch?"}
    FilterByDepth{"Depth ≤ max?"}
    FilterByParent{"No-parent\ncheck?"}
    ExcludeFilter{"Not in exclude\nlist?"}
    Enqueue["Add to download queue"]
    Skip["Skip"]

    HTML --> Extract --> FilterByDomain
    FilterByDomain -->|Yes| FilterByPattern
    FilterByDomain -->|No, --span-hosts| FilterByPattern
    FilterByDomain -->|No| Skip
    FilterByPattern -->|Yes| FilterByDepth
    FilterByPattern -->|No| Skip
    FilterByDepth -->|Yes| FilterByParent
    FilterByDepth -->|No| Skip
    FilterByParent -->|Pass| ExcludeFilter
    FilterByParent -->|Fail| Skip
    ExcludeFilter -->|Pass| Enqueue
    ExcludeFilter -->|Fail| Skip

Accept Patterns

The --accept flag supports two types:

Type Example Matches
Glob *.pdf Filenames ending in .pdf
Glob *.html,*.css Multiple patterns (comma-separated)
MIME text/html Exact MIME type
MIME image/* Any MIME in the image/ category

Domain Control

Flag Effect
(default) Only follow links within the starting domain
--span-hosts Follow links to any domain
--domains a.com,b.com Restrict to specific domains
--follow-external Follow external links (aliases for --span-hosts)
--no-parent Don't ascendant above the starting URL path
--recursive-parallel <N>

Request Throttling

# Fixed 2-second delay between requests
goget --recursive --wait 2s --url https://example.com/docs/

# Randomize delay (0.5x  1.5x of --wait)
goget --recursive --wait 2s --random-wait --url https://example.com/docs/

Page Requisites

--page-requisites downloads CSS, JavaScript, and images needed to render each HTML page:

goget --recursive --page-requisites --url https://example.com/page.html

This parses <link>, <script>, <img>, <source>, and <video> tags and downloads their src/href targets.

Mirror Mode

| --dry-run | — | List all files that would be downloaded in recursive/mirror mode without saving to disk |

--mirror is a shorthand for --recursive --convert-links --page-requisites --infinite-depth:

# Full site mirror
goget --url https://example.com --mirror --output ./mirror

# With link conversion for offline viewing
goget --url https://example.com --mirror --convert-links --output ./mirror

When --convert-links is enabled, HTML links are rewritten for local offline viewing:

Before: <a href="https://example.com/about/">About</a>
After:  <a href="./about/index.html">About</a>

Before: <img src="https://example.com/img/logo.png">
After:  <img src="./img/logo.png">

The link rewriter handles:

  • Absolute URLs → relative paths
  • Protocol-relative URLs (//example.com/...)
  • Root-relative URLs (/about/)
  • CSS url() references

Robots.txt

goget respects robots.txt by default. Use --no-robots to ignore it:

goget --mirror --no-robots --url https://example.com

Asset Control

# Mirror but skip CSS/JS/images
goget --mirror --no-mirror-assets --url https://example.com

# Only mirror specific file types
goget --mirror --accept "*.html,*.jpg" --url https://example.com

Output Structure

Mirrored sites preserve the URL path structure:

mirror/
├── index.html
├── about/
│   └── index.html
├── blog/
│   ├── index.html
│   └── post-1.html
├── css/
│   └── style.css
└── img/
    └── logo.png

Use --cut-dirs N to strip directory components:

# Original: https://example.com/pub/docs/file.html
# Without cut: ./mirror/pub/docs/file.html
# With --cut-dirs 2: ./mirror/doc/file.html
goget --mirror --cut-dirs 2 --url https://example.com/pub/docs/

Architecture

Crawler (internal/recursive)

type CrawlerConfig struct {
    MaxDepth        int
    FollowExternal  bool
    ExcludePatterns []string
    IncludePatterns []string
    AcceptPatterns  []string
    NoParent        bool
    PageRequisites  bool
    ConvertLinks    bool
    OutputDir       string
    Delay           time.Duration
    RandomWait      bool
    MaxWorkers      int  // concurrent download workers (0 = sequential)
}

The crawler:

  1. Downloads an HTML page
  2. Parses it with golang.org/x/net/html
  3. Extracts all links (a[href], img[src], link[href], script[src])
  4. Filters against accept/reject patterns, domain rules, and depth limits
  5. Enqueues matching URLs for download
  6. Applies configurable delay between requests

Mirror (internal/mirror)

The mirror wraps the crawler with additional configuration:

  • Infinite depth (or configurable via MirrorConfig.MaxDepth)
  • Automatic page requisites
  • Automatic link conversion
  • robots.txt compliance

After the download completes, links are rewritten in-place:

  1. Read each HTML file
  2. Parse with golang.org/x/net/html
  3. Identify all href and src attributes
  4. Convert absolute URLs to relative paths
  5. Write back the modified HTML (backup .orig if --backup-converted)

Timestamping

--timestamping downloads a file only if the server copy is newer than the local copy:

goget --timestamping --url https://example.com/file.zip

goget compares the Last-Modified header against the local file's modification time.

WARC Archiving

For legal and archival compliance, output can be written in WARC format:

goget --warc-file archive.warc --url https://example.com/page.html

Each downloaded resource gets a separate WARC record with:

  • Request metadata (URL, timestamp, headers)
  • Response metadata (status code, content type, headers)
  • Raw response body

Performance Tips

  • Limit depth--max-depth 3 prevents runaway recursion on large sites
  • Filter aggressively — Use --accept "*.pdf,*.html" and --reject "*.zip" to narrow scope
  • Use delay--wait 500ms prevents server overload and IP blocks
  • Limit domains--domains example.com prevents crawling external CDNs
  • Skip assets--no-mirror-assets for text-only archives