docs: describe the sglang endpoint behind caddy
Assisted-by: GLM 5.3 Flash
This commit is contained in:
@@ -22,7 +22,7 @@ own builtins, and none of them needs a module installed beside it.
|
|||||||
automatic updates); `workstation-setup.pl` sets up a Fedora desktop (Brave, the
|
automatic updates); `workstation-setup.pl` sets up a Fedora desktop (Brave, the
|
||||||
official Go toolchain, Rust, GoLand, Flatpak applications, firewall, SELinux).
|
official Go toolchain, Rust, GoLand, Flatpak applications, firewall, SELinux).
|
||||||
- **Deployment**: `sglang-deploy.pl` puts an SGLang inference server on an AMD GPU
|
- **Deployment**: `sglang-deploy.pl` puts an SGLang inference server on an AMD GPU
|
||||||
behind nginx with HTTPS and an API key, runs the engine from the project's ROCm
|
behind Caddy with HTTPS and an API key, runs the engine from the project's ROCm
|
||||||
container image, and downloads the model weights from ModelScope.
|
container image, and downloads the model weights from ModelScope.
|
||||||
- **Maintenance**: `system-optimise.pl` removes old kernels (the running one and one
|
- **Maintenance**: `system-optimise.pl` removes old kernels (the running one and one
|
||||||
fallback always stay), vacuums journals, clears temporary files and core dumps, and
|
fallback always stay), vacuums journals, clears temporary files and core dumps, and
|
||||||
@@ -92,7 +92,7 @@ a change would do without doing it.
|
|||||||
| `security-audit.pl` | 2.0.0 | Security posture with a grade from A to F and an exit code for cron. Linux only |
|
| `security-audit.pl` | 2.0.0 | Security posture with a grade from A to F and an exit code for cron. Linux only |
|
||||||
| `server-setup.pl` | 2.0.0 | Server initial setup for Fedora, CentOS Stream and openEuler |
|
| `server-setup.pl` | 2.0.0 | Server initial setup for Fedora, CentOS Stream and openEuler |
|
||||||
| `workstation-setup.pl` | 2.0.0 | Fedora desktop setup |
|
| `workstation-setup.pl` | 2.0.0 | Fedora desktop setup |
|
||||||
| `sglang-deploy.pl` | 2.0.0 | SGLang inference server on an AMD GPU, behind nginx with HTTPS |
|
| `sglang-deploy.pl` | 2.1.0 | SGLang inference server on an AMD GPU, behind Caddy with HTTPS |
|
||||||
| `system-optimise.pl` | 2.0.0 | System cleanup; refuses rpm-ostree systems |
|
| `system-optimise.pl` | 2.0.0 | System cleanup; refuses rpm-ostree systems |
|
||||||
|
|
||||||
The full flag reference is in [docs/CLI.md](docs/CLI.md).
|
The full flag reference is in [docs/CLI.md](docs/CLI.md).
|
||||||
|
|||||||
+1
-1
@@ -61,7 +61,7 @@ is a defect worth reporting:
|
|||||||
user who runs the script.
|
user who runs the script.
|
||||||
- Missing hardening with no demonstrated impact, such as a script not setting a stricter
|
- Missing hardening with no demonstrated impact, such as a script not setting a stricter
|
||||||
umask than the system's.
|
umask than the system's.
|
||||||
- Flaws in a third-party tool the scripts drive (`dnf`, `podman`, `openssl`, `nginx` and
|
- Flaws in a third-party tool the scripts drive (`dnf`, `podman`, `openssl`, `caddy` and
|
||||||
the rest): report those to that project, and here only when this repository's use of
|
the rest): report those to that project, and here only when this repository's use of
|
||||||
the tool makes the flaw reachable in a way the tool's own documentation does not
|
the tool makes the flaw reachable in a way the tool's own documentation does not
|
||||||
anticipate.
|
anticipate.
|
||||||
|
|||||||
+15
-3
@@ -37,7 +37,7 @@ reports everything as already in place and changes nothing.
|
|||||||
| `system-diag.pl` | Reading CPU, memory, disk, network, GPU, services, security and performance, and grading the result | Change anything on the host; it is read-only |
|
| `system-diag.pl` | Reading CPU, memory, disk, network, GPU, services, security and performance, and grading the result | Change anything on the host; it is read-only |
|
||||||
| `server-setup.pl` | Base packages, firewall (with the SSH rule verified before the service starts), SELinux, Podman, automatic updates | Install the services themselves; that is the operator's, and the deploy scripts' |
|
| `server-setup.pl` | Base packages, firewall (with the SSH rule verified before the service starts), SELinux, Podman, automatic updates | Install the services themselves; that is the operator's, and the deploy scripts' |
|
||||||
| `workstation-setup.pl` | A Fedora desktop: Brave, the official Go toolchain, Rust, GoLand, Flatpak applications, firewall, SELinux | Anything on a server distribution; it targets Fedora Workstation |
|
| `workstation-setup.pl` | A Fedora desktop: Brave, the official Go toolchain, Rust, GoLand, Flatpak applications, firewall, SELinux | Anything on a server distribution; it targets Fedora Workstation |
|
||||||
| `sglang-deploy.pl` | The SGLang deployment: the container engine, the systemd unit, nginx with TLS, the API key, the firewall rule and the SELinux boolean | The host's own setup, which `server-setup.pl` does first |
|
| `sglang-deploy.pl` | The SGLang deployment: the container engine, the systemd unit, Caddy with TLS, the API key, the firewall rule and the SELinux checks | The host's own setup, which `server-setup.pl` does first |
|
||||||
| `system-optimise.pl` | Old kernels (the running one and one fallback always stay), journals, temporary files, core dumps, and the package audit | Touch an rpm-ostree system, which it refuses |
|
| `system-optimise.pl` | Old kernels (the running one and one fallback always stay), journals, temporary files, core dumps, and the package audit | Touch an rpm-ostree system, which it refuses |
|
||||||
|
|
||||||
## Data flow: a forked section collection
|
## Data flow: a forked section collection
|
||||||
@@ -97,8 +97,8 @@ that CentOS Stream and openEuler cannot carry at all.
|
|||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
Client[client] -->|443| Nginx[nginx on the host, TLS]
|
Client[client] -->|443| Caddy[Caddy on the host, TLS]
|
||||||
Nginx -->|loopback, plain HTTP| Engine[SGLang in a Podman container]
|
Caddy -->|loopback, plain HTTP| Engine[SGLang in a Podman container]
|
||||||
Engine -->|device nodes| GPU[/dev/kfd, /dev/dri]
|
Engine -->|device nodes| GPU[/dev/kfd, /dev/dri]
|
||||||
Engine -->|bind mount| Cache[state directory, model cache]
|
Engine -->|bind mount| Cache[state directory, model cache]
|
||||||
Unit[systemd unit] -->|podman run| Engine
|
Unit[systemd unit] -->|podman run| Engine
|
||||||
@@ -114,12 +114,24 @@ instead, and `--image` pins one. Radeon cards need `SGLANG_USE_AITER=false` and
|
|||||||
`SGLANG_ROCM_FUSED_DECODE_MLA=false` in the unit, which the script writes for them and
|
`SGLANG_ROCM_FUSED_DECODE_MLA=false` in the unit, which the script writes for them and
|
||||||
never for an Instinct host.
|
never for an Instinct host.
|
||||||
|
|
||||||
|
Caddy serves the endpoint from a drop-in under `/etc/caddy/Caddyfile.d`, which the
|
||||||
|
distribution's default Caddyfile imports. Fedora carries the caddy package and CentOS
|
||||||
|
Stream gets it from EPEL, whose repository file the script installs first; openEuler
|
||||||
|
packages no caddy at all, so there the official release binary is installed instead,
|
||||||
|
with the unit file the package would have carried. The service runs as the caddy user,
|
||||||
|
so the private key is made group-readable for the caddy group, and on an
|
||||||
|
SELinux-enforcing host the packaged caddy runs unconfined: no boolean is needed. A
|
||||||
|
deployment made by an earlier release of this script carries an nginx configuration,
|
||||||
|
which both a deploy and an uninstall remove.
|
||||||
|
|
||||||
## Dependencies
|
## Dependencies
|
||||||
|
|
||||||
Nothing outside the interpreter, and nothing that has to be installed beyond the tools
|
Nothing outside the interpreter, and nothing that has to be installed beyond the tools
|
||||||
each script's own dependency section installs. The non-obvious ones and their reasons:
|
each script's own dependency section installs. The non-obvious ones and their reasons:
|
||||||
|
|
||||||
- `podman` for `sglang-deploy.pl`, because the engine is a container.
|
- `podman` for `sglang-deploy.pl`, because the engine is a container.
|
||||||
|
- `caddy` for `sglang-deploy.pl`, because the endpoint serves TLS on 443; the
|
||||||
|
release binary and `tar` stand in where no repository carries the package.
|
||||||
- `lspci` for the GPU family, and `rocm-smi` in `system-diag.pl` for AMD memory and
|
- `lspci` for the GPU family, and `rocm-smi` in `system-diag.pl` for AMD memory and
|
||||||
utilisation figures.
|
utilisation figures.
|
||||||
- `sha256sum` in `workstation-setup.pl`, because Perl's builtins have no hash.
|
- `sha256sum` in `workstation-setup.pl`, because Perl's builtins have no hash.
|
||||||
|
|||||||
+5
-5
@@ -111,7 +111,7 @@ verified by checksum; Brave's repository key is verified by fingerprint before i
|
|||||||
```
|
```
|
||||||
Usage: sglang-deploy.pl [options]
|
Usage: sglang-deploy.pl [options]
|
||||||
|
|
||||||
--model ID Hugging Face model ID (menu when omitted)
|
--model ID ModelScope model ID (menu when omitted)
|
||||||
--port N internal engine port (default: 8000, not 443)
|
--port N internal engine port (default: 8000, not 443)
|
||||||
--tensor-parallel N GPUs for tensor parallelism (default: 1, written as
|
--tensor-parallel N GPUs for tensor parallelism (default: 1, written as
|
||||||
the engine's --tp-size)
|
the engine's --tp-size)
|
||||||
@@ -130,19 +130,19 @@ Usage: sglang-deploy.pl [options]
|
|||||||
--api-key KEY API key for the endpoint (default: generate and
|
--api-key KEY API key for the endpoint (default: generate and
|
||||||
store in /etc/sysconfig)
|
store in /etc/sysconfig)
|
||||||
--dry-run preview without making changes
|
--dry-run preview without making changes
|
||||||
--uninstall tear down the service, container, nginx config
|
--uninstall tear down the service, container, caddy drop-in
|
||||||
and certificates
|
and certificates
|
||||||
--help show this help
|
--help show this help
|
||||||
--version show the version
|
--version show the version
|
||||||
```
|
```
|
||||||
|
|
||||||
Needs root. Without `--model` an interactive menu offers GLM 5.3, GLM 5.3 Flash,
|
Needs root. Without `--model` an interactive menu offers GLM 5.3, GLM 5.3 Flash,
|
||||||
DeepSeek V4 Pro, DeepSeek V4 Flash, MiMo V2.5 Pro and MiMo V2.5, plus a free-form
|
Qwen 3.8 Max, Qwen 3.8 Flash and DeepSeek V4.1 Flash, plus a free-form entry.
|
||||||
entry. `--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
|
`--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
|
||||||
written to the unit as the engine's own `--tp-size`, `--context-length` and
|
written to the unit as the engine's own `--tp-size`, `--context-length` and
|
||||||
`--mem-fraction-static`. The API key is shown once when it is generated; a key given
|
`--mem-fraction-static`. The API key is shown once when it is generated; a key given
|
||||||
with `--api-key` is never echoed. `--uninstall` stops and disables the service, removes
|
with `--api-key` is never echoed. `--uninstall` stops and disables the service, removes
|
||||||
the unit, the container, the nginx configuration and the certificates, and keeps the
|
the unit, the container, the caddy drop-in and the certificates, and keeps the
|
||||||
image and the model cache.
|
image and the model cache.
|
||||||
|
|
||||||
## system-optimise.pl
|
## system-optimise.pl
|
||||||
|
|||||||
+3
-3
@@ -84,7 +84,7 @@ deployment an operator cares about, so the pipeline that publishes them never ru
|
|||||||
| Script | What it leaves behind | Where it is documented |
|
| Script | What it leaves behind | Where it is documented |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `server-setup.pl` | Packages, firewalld with the SSH rule already allowed, SELinux enforcing, Podman, unattended updates | [docs/CLI.md](CLI.md) |
|
| `server-setup.pl` | Packages, firewalld with the SSH rule already allowed, SELinux enforcing, Podman, unattended updates | [docs/CLI.md](CLI.md) |
|
||||||
| `sglang-deploy.pl` | A systemd unit running the engine in a Podman container, nginx with TLS in front, an API key file, the firewall rule and the SELinux boolean | [docs/ARCHITECTURE.md](ARCHITECTURE.md) |
|
| `sglang-deploy.pl` | A systemd unit running the engine in a Podman container, Caddy with TLS in front, an API key file, the firewall rule and the SELinux checks | [docs/ARCHITECTURE.md](ARCHITECTURE.md) |
|
||||||
| `workstation-setup.pl` | A Fedora desktop with the toolchains and applications installed | [docs/CLI.md](CLI.md) |
|
| `workstation-setup.pl` | A Fedora desktop with the toolchains and applications installed | [docs/CLI.md](CLI.md) |
|
||||||
|
|
||||||
The SGLang deployment is the only service in the collection, and the host needs no ROCm
|
The SGLang deployment is the only service in the collection, and the host needs no ROCm
|
||||||
@@ -100,6 +100,6 @@ where the script reads them:
|
|||||||
|
|
||||||
- The published copies are unversioned: whatever is on `main` is what is served.
|
- The published copies are unversioned: whatever is on `main` is what is served.
|
||||||
- `sglang-deploy.pl` keeps the engine's API key in `/etc/sysconfig/sglang`, mode 0600,
|
- `sglang-deploy.pl` keeps the engine's API key in `/etc/sysconfig/sglang`, mode 0600,
|
||||||
written by the script and never committed. A Hugging Face token for a gated model goes
|
written by the script and never committed. A ModelScope token for a gated model goes
|
||||||
into that same file as `HF_TOKEN`, which the unit forwards to the container.
|
into that same file as `MODELSCOPE_TOKEN`, which the unit forwards to the container.
|
||||||
- The deploy secrets live in the Gitea repository settings, under Actions, Secrets.
|
- The deploy secrets live in the Gitea repository settings, under Actions, Secrets.
|
||||||
|
|||||||
+9
-7
@@ -61,13 +61,15 @@ only be exercised as root.
|
|||||||
|
|
||||||
`tests/container/rig.pl` runs the real `sglang-deploy.pl` as root inside a container
|
`tests/container/rig.pl` runs the real `sglang-deploy.pl` as root inside a container
|
||||||
against a stub `PATH`: every command the script drives (`dnf`, `rpm`, `podman`,
|
against a stub `PATH`: every command the script drives (`dnf`, `rpm`, `podman`,
|
||||||
`systemctl`, `curl`, `openssl`, `nginx`, `firewall-cmd`, `getsebool`, `lspci`) is a
|
`systemctl`, `curl`, `openssl`, `caddy`, `tar`, `useradd`, `semodule`, `firewall-cmd`,
|
||||||
stub that answers from a fixture, so a run is deterministic and needs no network and no
|
`getsebool`, `lspci`) is a stub that answers from a fixture, so a run is deterministic
|
||||||
GPU. It covers the deploy end to end: the resolved image tag, the unit file, the nginx
|
and needs no network and no GPU. It covers the deploy end to end: the resolved image
|
||||||
configuration, the TLS certificate and its modes, the API key file, the SELinux
|
tag, the unit file, the caddy drop-in and the main Caddyfile's import, the release
|
||||||
boolean, the firewall rule, the idempotent second run, the dry run, the uninstall, the
|
binary install for the hosts without a caddy package, the EPEL bootstrap on CentOS
|
||||||
Radeon and MI300 paths, the offline and unpublished-tag failures, the argument
|
Stream, the legacy nginx clean-up, the TLS certificate and its modes, the API key file,
|
||||||
validation and the non-root refusal.
|
the SELinux decisions, the firewall rule, the idempotent second run, the dry run, the
|
||||||
|
uninstall, the Radeon and MI300 paths, the offline and unpublished-tag failures, the
|
||||||
|
argument validation and the non-root refusal.
|
||||||
|
|
||||||
Prepare the platform image once, since the base images carry no Perl:
|
Prepare the platform image once, since the base images carry no Perl:
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user