docs: describe the sglang endpoint behind caddy
Assisted-by: GLM 5.3 Flash
This commit is contained in:
+5
-5
@@ -111,7 +111,7 @@ verified by checksum; Brave's repository key is verified by fingerprint before i
|
||||
```
|
||||
Usage: sglang-deploy.pl [options]
|
||||
|
||||
--model ID Hugging Face model ID (menu when omitted)
|
||||
--model ID ModelScope model ID (menu when omitted)
|
||||
--port N internal engine port (default: 8000, not 443)
|
||||
--tensor-parallel N GPUs for tensor parallelism (default: 1, written as
|
||||
the engine's --tp-size)
|
||||
@@ -130,19 +130,19 @@ Usage: sglang-deploy.pl [options]
|
||||
--api-key KEY API key for the endpoint (default: generate and
|
||||
store in /etc/sysconfig)
|
||||
--dry-run preview without making changes
|
||||
--uninstall tear down the service, container, nginx config
|
||||
--uninstall tear down the service, container, caddy drop-in
|
||||
and certificates
|
||||
--help show this help
|
||||
--version show the version
|
||||
```
|
||||
|
||||
Needs root. Without `--model` an interactive menu offers GLM 5.3, GLM 5.3 Flash,
|
||||
DeepSeek V4 Pro, DeepSeek V4 Flash, MiMo V2.5 Pro and MiMo V2.5, plus a free-form
|
||||
entry. `--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
|
||||
Qwen 3.8 Max, Qwen 3.8 Flash and DeepSeek V4.1 Flash, plus a free-form entry.
|
||||
`--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
|
||||
written to the unit as the engine's own `--tp-size`, `--context-length` and
|
||||
`--mem-fraction-static`. The API key is shown once when it is generated; a key given
|
||||
with `--api-key` is never echoed. `--uninstall` stops and disables the service, removes
|
||||
the unit, the container, the nginx configuration and the certificates, and keeps the
|
||||
the unit, the container, the caddy drop-in and the certificates, and keeps the
|
||||
image and the model cache.
|
||||
|
||||
## system-optimise.pl
|
||||
|
||||
Reference in New Issue
Block a user