257 lines
7.8 KiB
Markdown
257 lines
7.8 KiB
Markdown
# Recursive Downloading & Website Mirroring
|
||
|
||
goget can recursively download linked pages and mirror entire websites for offline viewing. This document explains the two modes, their configuration, and the underlying architecture.
|
||
|
||
## Quick Comparison
|
||
|
||
| Feature | Recursive (`--recursive`) | Mirror (`--mirror`) |
|
||
|---|---|---|
|
||
| Depth limit | By default (`--max-depth N`) | Infinite |
|
||
| Link conversion | Optional (`--convert-links`) | Automatic |
|
||
| Page requisites | Manual (`--page-requisites`) | Automatic |
|
||
| robots.txt | Respected by default | Respected by default |
|
||
| Use case | Download a subtree of a site | Full offline archive |
|
||
|
||
## Recursive Mode
|
||
|
||
Recursive mode follows links from HTML pages and downloads linked resources.
|
||
|
||
```bash
|
||
# Basic recursive download, depth 3
|
||
goget --url https://example.com/docs/ --recursive --max-depth 3
|
||
|
||
# Only PDF files
|
||
goget --recursive --accept "*.pdf" --url https://example.com/docs/
|
||
|
||
# Only images and HTML
|
||
goget --recursive --accept "image/*,text/html" --url https://example.com/
|
||
```
|
||
|
||
### Link Filtering
|
||
|
||
```mermaid
|
||
flowchart TD
|
||
HTML[Parse HTML page]:::accent1
|
||
Extract["Extract all links\n(a, img, link, script)"]:::accent1
|
||
FilterByDomain{"Domain match?"}:::accent7
|
||
FilterByPattern{"Accept pattern\nmatch?"}:::accent7
|
||
FilterByDepth{"Depth ≤ max?"}:::accent7
|
||
FilterByParent{"No-parent\ncheck?"}:::accent7
|
||
ExcludeFilter{"Not in exclude\nlist?"}:::accent7
|
||
Enqueue["Add to download queue"]:::accent1
|
||
Skip["Skip"]:::accent4
|
||
|
||
HTML --> Extract --> FilterByDomain
|
||
FilterByDomain -->|Yes| FilterByPattern
|
||
FilterByDomain -->|No, --span-hosts| FilterByPattern
|
||
FilterByDomain -->|No| Skip
|
||
FilterByPattern -->|Yes| FilterByDepth
|
||
FilterByPattern -->|No| Skip
|
||
FilterByDepth -->|Yes| FilterByParent
|
||
FilterByDepth -->|No| Skip
|
||
FilterByParent -->|Pass| ExcludeFilter
|
||
FilterByParent -->|Fail| Skip
|
||
ExcludeFilter -->|Pass| Enqueue
|
||
ExcludeFilter -->|Fail| Skip
|
||
|
||
classDef accent1 fill:#22C55E,stroke:#16A34A,color:#fff
|
||
classDef accent4 fill:#EF4444,stroke:#DC2626,color:#fff
|
||
classDef accent7 fill:#64748B,stroke:#475569,color:#fff
|
||
```
|
||
|
||
### Accept Patterns
|
||
|
||
The `--accept` flag supports two types:
|
||
|
||
| Type | Example | Matches |
|
||
|---|---|---|
|
||
| **Glob** | `*.pdf` | Filenames ending in `.pdf` |
|
||
| **Glob** | `*.html,*.css` | Multiple patterns (comma-separated) |
|
||
| **MIME** | `text/html` | Exact MIME type |
|
||
| **MIME** | `image/*` | Any MIME in the `image/` category |
|
||
|
||
### Domain Control
|
||
|
||
| Flag | Effect |
|
||
|---|---|
|
||
| (default) | Only follow links within the starting domain |
|
||
| `--span-hosts` | Follow links to any domain |
|
||
| `--domains a.com,b.com` | Restrict to specific domains |
|
||
| `--follow-external` | Follow external links (aliases for `--span-hosts`) |
|
||
| `--no-parent` | Don't ascendant above the starting URL path |
|
||
| `--recursive-parallel` | `<N>` | Number of concurrent file downloads in recursive mode across all protocols (`0` = sequential). HTTP, WebDAV, FTP, and SFTP each get a worker pool bounded by a semaphore; subdirectory recursion propagates the limit. Connection pooling (FTP, SFTP) opens N independent control connections |
|
||
|
||
### Request Throttling
|
||
|
||
```bash
|
||
# Fixed 2-second delay between requests
|
||
goget --recursive --wait 2s --url https://example.com/docs/
|
||
|
||
# Randomize delay (0.5x – 1.5x of --wait)
|
||
goget --recursive --wait 2s --random-wait --url https://example.com/docs/
|
||
```
|
||
|
||
### Page Requisites
|
||
|
||
`--page-requisites` downloads CSS, JavaScript, and images needed to render each HTML page:
|
||
|
||
```bash
|
||
goget --recursive --page-requisites --url https://example.com/page.html
|
||
```
|
||
|
||
This parses `<link>`, `<script>`, `<img>`, `<source>`, and `<video>` tags and downloads their `src`/`href` targets.
|
||
|
||
## Mirror Mode
|
||
|
||
| `--dry-run` | — | List all files that would be downloaded in recursive/mirror mode without saving to disk |
|
||
|
||
`--mirror` is a shorthand for `--recursive --convert-links --page-requisites --infinite-depth`:
|
||
|
||
```bash
|
||
# Full site mirror
|
||
goget --url https://example.com --mirror --output ./mirror
|
||
|
||
# With link conversion for offline viewing
|
||
goget --url https://example.com --mirror --convert-links --output ./mirror
|
||
```
|
||
|
||
### Link Conversion
|
||
|
||
When `--convert-links` is enabled, HTML links are rewritten for local offline viewing:
|
||
|
||
```
|
||
Before: <a href="https://example.com/about/">About</a>
|
||
After: <a href="./about/index.html">About</a>
|
||
|
||
Before: <img src="https://example.com/img/logo.png">
|
||
After: <img src="./img/logo.png">
|
||
```
|
||
|
||
The link rewriter handles:
|
||
- Absolute URLs → relative paths
|
||
- Protocol-relative URLs (`//example.com/...`)
|
||
- Root-relative URLs (`/about/`)
|
||
- CSS `url()` references
|
||
|
||
### Robots.txt
|
||
|
||
goget respects `robots.txt` by default. Use `--no-robots` to ignore it:
|
||
|
||
```bash
|
||
goget --mirror --no-robots --url https://example.com
|
||
```
|
||
|
||
### Asset Control
|
||
|
||
```bash
|
||
# Mirror but skip CSS/JS/images
|
||
goget --mirror --no-mirror-assets --url https://example.com
|
||
|
||
# Only mirror specific file types
|
||
goget --mirror --accept "*.html,*.jpg" --url https://example.com
|
||
```
|
||
|
||
## Output Structure
|
||
|
||
Mirrored sites preserve the URL path structure:
|
||
|
||
```
|
||
mirror/
|
||
├── index.html
|
||
├── about/
|
||
│ └── index.html
|
||
├── blog/
|
||
│ ├── index.html
|
||
│ └── post-1.html
|
||
├── css/
|
||
│ └── style.css
|
||
└── img/
|
||
└── logo.png
|
||
```
|
||
|
||
Use `--cut-dirs N` to strip directory components:
|
||
|
||
```bash
|
||
# Original: https://example.com/pub/docs/file.html
|
||
# Without cut: ./mirror/pub/docs/file.html
|
||
# With --cut-dirs 2: ./mirror/doc/file.html
|
||
goget --mirror --cut-dirs 2 --url https://example.com/pub/docs/
|
||
```
|
||
|
||
## Architecture
|
||
|
||
### Crawler (`internal/recursive`)
|
||
|
||
```go
|
||
type CrawlerConfig struct {
|
||
MaxDepth int
|
||
FollowExternal bool
|
||
ExcludePatterns []string
|
||
IncludePatterns []string
|
||
AcceptPatterns []string
|
||
NoParent bool
|
||
PageRequisites bool
|
||
ConvertLinks bool
|
||
OutputDir string
|
||
Delay time.Duration
|
||
RandomWait bool
|
||
MaxWorkers int // concurrent download workers (0 = sequential)
|
||
}
|
||
```
|
||
|
||
The crawler:
|
||
1. Downloads an HTML page
|
||
2. Parses it with `golang.org/x/net/html`
|
||
3. Extracts all links (`a[href]`, `img[src]`, `link[href]`, `script[src]`)
|
||
4. Filters against accept/reject patterns, domain rules, and depth limits
|
||
5. Enqueues matching URLs for download
|
||
6. Applies configurable delay between requests
|
||
|
||
### Mirror (`internal/mirror`)
|
||
|
||
The mirror wraps the crawler with additional configuration:
|
||
- Infinite depth (or configurable via `MirrorConfig.MaxDepth`)
|
||
- Automatic page requisites
|
||
- Automatic link conversion
|
||
- robots.txt compliance
|
||
|
||
### Link Rewriting (`internal/linkrewrite`)
|
||
|
||
After the download completes, links are rewritten in-place:
|
||
1. Read each HTML file
|
||
2. Parse with `golang.org/x/net/html`
|
||
3. Identify all `href` and `src` attributes
|
||
4. Convert absolute URLs to relative paths
|
||
5. Write back the modified HTML (backup `.orig` if `--backup-converted`)
|
||
|
||
## Timestamping
|
||
|
||
`--timestamping` downloads a file only if the server copy is newer than the local copy:
|
||
|
||
```bash
|
||
goget --timestamping --url https://example.com/file.zip
|
||
```
|
||
|
||
goget compares the `Last-Modified` header against the local file's modification time.
|
||
|
||
## WARC Archiving
|
||
|
||
For legal and archival compliance, output can be written in WARC format:
|
||
|
||
```bash
|
||
goget --warc-file archive.warc --url https://example.com/page.html
|
||
```
|
||
|
||
Each downloaded resource gets a separate WARC record with:
|
||
- Request metadata (URL, timestamp, headers)
|
||
- Response metadata (status code, content type, headers)
|
||
- Raw response body
|
||
|
||
## Performance Tips
|
||
|
||
- **Limit depth** — `--max-depth 3` prevents runaway recursion on large sites
|
||
- **Filter aggressively** — Use `--accept "*.pdf,*.html"` and `--reject "*.zip"` to narrow scope
|
||
- **Use delay** — `--wait 500ms` prevents server overload and IP blocks
|
||
- **Limit domains** — `--domains example.com` prevents crawling external CDNs
|
||
- **Skip assets** — `--no-mirror-assets` for text-only archives
|