docs: describe the sglang endpoint behind caddy

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-29 00:18:15 +02:00
parent 9455c65f10
commit 37e06d090c
6 changed files with 35 additions and 21 deletions
+5 -5
View File
@@ -111,7 +111,7 @@ verified by checksum; Brave's repository key is verified by fingerprint before i
```
Usage: sglang-deploy.pl [options]
--model ID Hugging Face model ID (menu when omitted)
--model ID ModelScope model ID (menu when omitted)
--port N internal engine port (default: 8000, not 443)
--tensor-parallel N GPUs for tensor parallelism (default: 1, written as
the engine's --tp-size)
@@ -130,19 +130,19 @@ Usage: sglang-deploy.pl [options]
--api-key KEY API key for the endpoint (default: generate and
store in /etc/sysconfig)
--dry-run preview without making changes
--uninstall tear down the service, container, nginx config
--uninstall tear down the service, container, caddy drop-in
and certificates
--help show this help
--version show the version
```
Needs root. Without `--model` an interactive menu offers GLM 5.3, GLM 5.3 Flash,
DeepSeek V4 Pro, DeepSeek V4 Flash, MiMo V2.5 Pro and MiMo V2.5, plus a free-form
entry. `--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
Qwen 3.8 Max, Qwen 3.8 Flash and DeepSeek V4.1 Flash, plus a free-form entry.
`--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
written to the unit as the engine's own `--tp-size`, `--context-length` and
`--mem-fraction-static`. The API key is shown once when it is generated; a key given
with `--api-key` is never echoed. `--uninstall` stops and disables the service, removes
the unit, the container, the nginx configuration and the certificates, and keeps the
the unit, the container, the caddy drop-in and the certificates, and keeps the
image and the model cache.
## system-optimise.pl