feat: initial release of the scripts collection
Deploy / deploy (push) Successful in 17s
Test / test (push) Successful in 53s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-10 04:00:00 +00:00
commit 788cf0571f
27 changed files with 19257 additions and 0 deletions
+128
View File
@@ -0,0 +1,128 @@
# Architecture
How the scripts collection is put together. Every script, helper and arrow below
exists in the repository; nothing is aspirational.
## Overview
There is no shared library. Each script is one file that can be fetched and run on its
own, and the six of them share a skeleton rather than a module: the same command
runner, the same progress reporting, the same state checks before every action, and the
same shape of a run.
```mermaid
flowchart TD
Start[invocation] --> Args[parse_args reads ARGV]
Args --> Version{version or help?}
Version -->|yes| Print[print and exit]
Version -->|no| Root[check_root, detect_os]
Root --> Sections[sections run in order]
Sections --> Each[each section checks state first]
Each --> Done{changed?}
Done -->|no| Report[report what is already in place]
Done -->|yes| Act[act, then verify what was done]
Act --> Report
Report --> Summary[print_summary on stderr]
```
A run is idempotent because of that loop: a section reads the state of the system,
compares it with the desired one, and acts only on the difference, so a second run
reports everything as already in place and changes nothing.
## The seven scripts
| Script | Owns | Deliberately does not |
|---|---|---|
| `network-diag.pl` | Latency, loss, DNS timing, MTU discovery, dual-stack reachability, listening ports, and the grade of the connection | Touch any configuration; it only measures |
| `system-diag.pl` | Reading CPU, memory, disk, network, GPU, services, security and performance, and grading the result | Change anything on the host; it is read-only |
| `server-setup.pl` | Base packages, firewall (with the SSH rule verified before the service starts), SELinux, Podman, automatic updates | Install the services themselves; that is the operator's, and the deploy scripts' |
| `workstation-setup.pl` | A Fedora desktop: Brave, the official Go toolchain, Rust, GoLand, Flatpak applications, firewall, SELinux | Anything on a server distribution; it targets Fedora Workstation |
| `sglang-deploy.pl` | The SGLang deployment: the container engine, the systemd unit, nginx with TLS, the API key, the firewall rule and the SELinux boolean | The host's own setup, which `server-setup.pl` does first |
| `system-optimise.pl` | Old kernels (the running one and one fallback always stay), journals, temporary files, core dumps, and the package audit | Touch an rpm-ostree system, which it refuses |
## Data flow: a forked section collection
`system-diag.pl` is the one script whose shape is not linear. Its sections are
independent and dominated by waiting, so each runs in its own child process, writes its
result into a private scratch directory as a Perl literal, and the parent reads the
results back after reaping. The runner keeps stdout and stderr apart in that same
directory, which is why the file names carry the process id.
```mermaid
sequenceDiagram
participant Parent
participant Child as Section child
participant Disk as Scratch directory
Parent->>Child: fork, one per requested section
Child->>Child: collect one section
Child->>Disk: write the result as a Perl literal
Parent->>Parent: waitpid, report the section as it lands
Parent->>Disk: read_literal for each child
Parent->>Parent: compute the grade and render the report
Parent->>Disk: remove the scratch directory
```
Only the process that created the scratch directory removes it: a forked child inherits
the `END` block, so the cleanup is guarded by a parent pid comparison. A child that
cannot write its result is reported as a missing section, and the grade drops that
component rather than inventing a value for it.
## Conventions every script follows
- **Builtins only.** No module beyond the interpreter is loaded, because the systems
this targets package modules separately. What Perl has no builtin for is written out:
the command runner, RPM's version comparison, the JSON decoders and the encoder, the
IPv4 and IPv6 arithmetic, `sockaddr` packing, a `which`, and argument parsing.
- **External binaries as transport.** `dnf`, `rpm`, `curl`, `openssl`, `podman`,
`systemctl`, `journalctl`, `find`, `stat`, `sha256sum`, `tar`, `gpg`, `lspci` and the
rest are driven as argument lists through `run`, which keeps stdout and stderr apart,
forces `LANG=C` and `LC_ALL=C` for parseable output, and turns a timeout, a missing
binary and a failed exec into exit codes 124, 127 and 126 rather than exceptions.
- **One optional module.** `Time::HiRes` is loaded inside an `eval` in the scripts that
time something. Where it is absent the affected figures are reported as not measured
and the grade drops that component, rather than the script failing.
- **Atomic writes for system configuration.** A file that a boot depends on is written
to a temporary name, flushed, given its mode and owner, and renamed over the target.
Perl has no `fsync`, so that barrier is delegated to `sync` where the binary exists.
- **Errors carry what the tool said.** A failure reproduces the errno, the reason and
the path, so a message can be searched for as it stands.
## The SGLang deployment
`sglang-deploy.pl` is the only script that deploys a service, and its engine runs as a
container. That is a consequence of the material rather than a preference: SGLang
publishes no ROCm wheel, and its AMD install paths are the project's container images
or a source build against a full ROCm toolchain that Fedora carries only partly and
that CentOS Stream and openEuler cannot carry at all.
```mermaid
flowchart LR
Client[client] -->|443| Nginx[nginx on the host, TLS]
Nginx -->|loopback, plain HTTP| Engine[SGLang in a Podman container]
Engine -->|device nodes| GPU[/dev/kfd, /dev/dri]
Engine -->|bind mount| Cache[state directory, model cache]
Unit[systemd unit] -->|podman run| Engine
EnvFile[EnvironmentFile, mode 0600] -->|API key| Unit
```
The host therefore needs the amdgpu kernel driver and its device nodes, never a ROCm
userland. The image tag is resolved from the card's family (mi30x for an MI300 or
MI325, mi35x for an MI350 or MI355) and from the host's ROCm, because the container's
ROCm userland must not be newer than the host's kernel driver. A Radeon 8060S or 8050S
(gfx1151) has no stable tag anywhere; AMD's dated development builds are resolved
instead, and `--image` pins one. Radeon cards need `SGLANG_USE_AITER=false` and
`SGLANG_ROCM_FUSED_DECODE_MLA=false` in the unit, which the script writes for them and
never for an Instinct host.
## Dependencies
Nothing outside the interpreter, and nothing that has to be installed beyond the tools
each script's own dependency section installs. The non-obvious ones and their reasons:
- `podman` for `sglang-deploy.pl`, because the engine is a container.
- `lspci` for the GPU family, and `rocm-smi` in `system-diag.pl` for AMD memory and
utilisation figures.
- `sha256sum` in `workstation-setup.pl`, because Perl's builtins have no hash.
- `sync` for the write barrier after an atomic write.
- `getent` or `host` in `network-diag.pl` as the resolver transport, since `getent` on
FreeBSD implements neither of the `ahosts` databases.
+197
View File
@@ -0,0 +1,197 @@
# Command line
Every script is its own program, and there is no global command. The reference below is
taken from each script's own `--help`; if the two disagree, the program is right and
this file is a defect.
Common to all six: `--version` prints the version and exits, `-h` / `--help` prints the
usage and exits, options are given as `--name value` or `--name=value`, and an unknown
option is refused with the usage and exit code 2.
## network-diag.pl
```
Usage: network-diag.pl [options]
Network diagnostics: latency, DNS, MTU, dual-stack, ports
Options:
--target HOST Target host for tests (default: cloudflare.com)
--protocol FAMILY Address family: auto (dual-stack), v4 (IPv4 only),
v6 (IPv6 only) (default: auto)
--version Show the version and exit
-h, --help Show this help and exit
```
| Flag | Default | Effect |
|---|---|---|
| `--target HOST` | `cloudflare.com` | The host pinged, resolved and probed for MTU |
| `--protocol FAMILY` | `auto` | `auto` measures both families, `v4` and `v6` restrict the run |
Runs without root. Ports belonging to other users are reported as `(no permission)`
rather than silently dropped.
## system-diag.pl
```
Usage: system-diag.pl [options]
System diagnostics: CPU, memory, disk, network, GPU, services, security, performance
Options:
--section LIST Comma-separated sections to run. Available: overview, cpu, memory, disk, network, gpu, services, security, performance, issues [default: all]
--external-ip Also detect external IP addresses via icanhazip.com (sends a request to a third-party service)
--json Output machine-readable JSON to stdout
--version Show the version and exit
-h, --help Show this help and exit
```
| Flag | Default | Effect |
|---|---|---|
| `--section LIST` | all | Runs only the named sections |
| `--external-ip` | off | Sends a request to `icanhazip.com` to report the public address |
| `--json` | off | Writes the JSON document to stdout instead of the report |
Runs without root; some figures need it and are reported as unavailable instead. Colour
is used only on a terminal and is suppressed by `NO_COLOR`. The report goes to stderr,
so `--json` keeps stdout clean.
## server-setup.pl
```
Usage: server-setup.pl [options]
Idempotent server setup for Fedora Server, CentOS Stream and openEuler
Options:
--dry-run Print what would be done without making changes
--skip-update Skip system update
--skip-packages Skip base package installation
--skip-epel Skip EPEL repository setup on CentOS
--skip-firewall Skip firewall setup
--skip-selinux Skip SELinux configuration
--skip-podman Skip Podman installation
--skip-auto-updates Skip automatic updates configuration
--version Show the version and exit
-h, --help Show this help and exit
```
Needs root. The `--skip-*` flags exist per section so a section can be left to another
tool; the sections are ordered update, EPEL, packages, firewall, SELinux, Podman,
automatic updates.
## workstation-setup.pl
```
Usage: workstation-setup.pl [options]
Idempotent workstation setup for Fedora
Options:
--dry-run Print what would be done without making changes
--skip-update Skip system update
--skip-rpm Skip RPM package installation
--skip-flatpak Skip Flatpak apps
--skip-repos Skip adding third-party repos
--skip-remove Skip removing pre-installed apps
--skip-go Skip Go toolchain installation
--skip-rust Skip Rust toolchain installation
--skip-jetbrains Skip JetBrains IDE installation
--skip-firewall Skip firewall setup
--skip-selinux Skip SELinux setup
--version Show the version and exit
-h, --help Show this help and exit
```
Needs root, and targets Fedora only. The Go toolchain and GoLand are downloaded and
verified by checksum; Brave's repository key is verified by fingerprint before import.
## sglang-deploy.pl
```
Usage: sglang-deploy.pl [options]
--model ID Hugging Face model ID (menu when omitted)
--port N internal engine port (default: 8000, not 443)
--tensor-parallel N GPUs for tensor parallelism (default: 1, written as
the engine's --tp-size)
--max-model-len N context length (default: 4096, written as the
engine's --context-length)
--gpu-memory-utilization F static memory fraction (default: 0.90, written as
the engine's --mem-fraction-static)
--state-dir PATH state and model cache directory (default: /opt/sglang)
--service-name NAME systemd service name (default: sglang)
--cert-dir PATH TLS certificate directory (default: /etc/ssl/sglang)
--image TAG engine image (default: resolved from the GPU and
the newest SGLang release; a Radeon card resolves
AMD's newest dated gfx1151 build)
--rocm-flavour NAME ROCm flavour of the image (default: from the host
ROCm; rocm10, rocm724, rocm720 or rocm700)
--api-key KEY API key for the endpoint (default: generate and
store in /etc/sysconfig)
--dry-run preview without making changes
--uninstall tear down the service, container, nginx config
and certificates
--help show this help
--version show the version
```
Needs root. Without `--model` an interactive menu offers GLM 5.3, GLM 5.3 Flash,
DeepSeek V4 Pro, DeepSeek V4 Flash, MiMo V2.5 Pro and MiMo V2.5, plus a free-form
entry. `--tensor-parallel`, `--max-model-len` and `--gpu-memory-utilization` are
written to the unit as the engine's own `--tp-size`, `--context-length` and
`--mem-fraction-static`. The API key is shown once when it is generated; a key given
with `--api-key` is never echoed. `--uninstall` stops and disables the service, removes
the unit, the container, the nginx configuration and the certificates, and keeps the
image and the model cache.
## system-optimise.pl
```
Usage: system-optimise.pl [options]
Idempotent system cleanup and optimisation for Fedora, CentOS Stream and openEuler
Options:
--dry-run Preview without making changes
--skip-dnf Skip DNF cleanup (autoremove, old kernels, cache)
--skip-journal Skip journal vacuum
--skip-tmp Skip temp file cleanup
--skip-cores Skip core dump cleanup
--version Show the version and exit
-h, --help Show this help and exit
```
Needs root. The running kernel and one fallback are always kept. Both dnf generations
are driven: Fedora ships dnf 5, CentOS Stream 10 and openEuler ship dnf 4.
## Exit codes
| Code | Meaning |
|---|---|
| `0` | The run completed, and every step it attempted succeeded |
| `1` | A step failed, a required tool is missing, the hardware does not qualify, or the system is unsupported |
| `2` | The arguments were wrong, or interactive input was needed and stdin is not a terminal |
A run that finishes with failed steps prints them at the end of the summary and exits 1,
so a pipeline notices even when the failure was not fatal.
## Examples
```sh
# Is the connection healthy, and is IPv6 working?
perl network-diag.pl --target example.org --protocol auto
# The same machine's health as JSON, for a monitoring poll
perl system-diag.pl --json --section cpu,memory,disk
# A new server, seen before it is changed
sudo perl server-setup.pl --dry-run
# Cleanup, leaving journals and core dumps alone
sudo perl system-optimise.pl --skip-journal --skip-cores
# Deploy a model and then take it down again
sudo perl sglang-deploy.pl --model deepseek-ai/DeepSeek-V4-Flash-0731
sudo perl sglang-deploy.pl --uninstall
```
+105
View File
@@ -0,0 +1,105 @@
# Deployment
The collection has two senses of deployment: how the scripts themselves are published,
and what they deploy when they run on a host. Both are below.
## Topology: publishing
The scripts are served from `https://petrbalvin.org/scripts/`, which is a directory on a
host reached over SSH. The Gitea instance at `sourcedock.dev` holds the repository and
runs the pipeline; there is no package, no tag and no artefact store, because a script
that is a single file is published by copying it.
```mermaid
flowchart LR
Dev[development branch] -->|merge| Main[main]
Main --> Runner[Gitea runner, alpine]
Runner -->|rsync over ssh| Host[petrbalvin.org]
Host -->|https| Client[client]
Runner -->|SHA256SUMS| Host
```
## Requirements
- The runner image is `alpine`, which carries `ssh`, `sshpass`, `rsync` and no Perl: the
pipeline installs Perl in its first step.
- Five repository secrets: `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PATH`,
`DEPLOY_PASSWORD`, and `DEPLOY_KNOWN_HOSTS`.
- `DEPLOY_KNOWN_HOSTS` holds the host key, from
`ssh-keyscan -t ed25519 <host>`. The pipeline writes it into `~/.ssh/known_hosts` and
connects with `StrictHostKeyChecking=yes`; without the secret the deploy stops before
it starts, deliberately, because the alternative is trusting whatever answers on the
first connection.
## Build
There is nothing to build. The pipeline writes `SHA256SUMS`, one line per script, hashed
with `sha256sum` in sorted filename order so the file is stable between runs, and then
sends the scripts and that file.
## Run
The pipeline runs itself: merging into `main` publishes, and a manual dispatch republishes
the current tree.
```mermaid
sequenceDiagram
participant Dev
participant CI as Gitea runner
participant Host as petrbalvin.org
Dev->>CI: push to main
CI->>CI: write SHA256SUMS over every script
CI->>CI: pin the host key from DEPLOY_KNOWN_HOSTS
CI->>Host: rsync -avz --delete, the scripts and SHA256SUMS
Host-->>CI: exit status of rsync
CI->>CI: stop the run unless it was zero
```
`--delete` is deliberate: the host directory mirrors the repository, so a script removed
here disappears there. Nothing else lives in that directory.
## Upgrade and rollback
A publish is atomic per file as rsync writes it, and the previous copy is overwritten in
place. Rollback is a revert commit on `main`: the pipeline publishes the reverted tree on
the next push, and the host is back to the earlier scripts within the minute the run
takes. There is no rehearsal of that procedure on a staging host; it is the same
pipeline with the same single target.
## Monitoring
Nothing polls the host. Two things are worth watching by hand:
- The pipeline's own run history in Gitea, which fails loudly on a non-zero `rsync` or a
missing secret.
- `https://petrbalvin.org/scripts/SHA256SUMS`, which a client can compare against what it
downloaded. A file that is on the host but missing from `SHA256SUMS` means a publish was
interrupted.
## What the scripts deploy
Two of the six change a host's role, and a third prepares a desktop. Their output is the
deployment an operator cares about, so the pipeline that publishes them never runs them.
| Script | What it leaves behind | Where it is documented |
|---|---|---|
| `server-setup.pl` | Packages, firewalld with the SSH rule already allowed, SELinux enforcing, Podman, unattended updates | [docs/CLI.md](CLI.md) |
| `sglang-deploy.pl` | A systemd unit running the engine in a Podman container, nginx with TLS in front, an API key file, the firewall rule and the SELinux boolean | [docs/ARCHITECTURE.md](ARCHITECTURE.md) |
| `workstation-setup.pl` | A Fedora desktop with the toolchains and applications installed | [docs/CLI.md](CLI.md) |
The SGLang deployment is the only service in the collection, and the host needs no ROCm
installation for it: the engine runs from the project's ROCm image and reaches the GPU
through `/dev/kfd` and `/dev/dri`, which the amdgpu kernel driver provides. The image tag
follows the card and the host's ROCm, which is why the script resolves it rather than
carrying a constant.
## Production configuration
Nothing in this repository holds a secret. The values that differ per host come from
where the script reads them:
- The published copies are unversioned: whatever is on `main` is what is served.
- `sglang-deploy.pl` keeps the engine's API key in `/etc/sysconfig/sglang`, mode 0600,
written by the script and never committed. A Hugging Face token for a gated model goes
into that same file as `HF_TOKEN`, which the unit forwards to the container.
- The deploy secrets live in the Gitea repository settings, under Actions, Secrets.
+146
View File
@@ -0,0 +1,146 @@
# Development
How to work on the scripts collection.
## Prerequisites
- Perl 5.38 or newer, which is what the oldest supported system ships. Measured on the
three test platforms: Fedora 44 carries 5.42.3, CentOS Stream 10 carries 5.40.2, and
openEuler 24.03 LTS carries 5.38.0.
- Podman, for the container rigs that verify a script against a real system.
- Nothing else. There is no build step, no dependency to install and no module to fetch:
the scripts use the interpreter's builtins, and the checks use nothing beyond them
either.
## Setup
```sh
git clone https://sourcedock.dev/petrbalvin/scripts.git
cd scripts
perl -c network-diag.pl
```
## Commands
There is no recipe file: the gates are the commands below, and the test pipeline runs
the same set.
| Command | What it does |
|---|---|
| `perl -c <script>` | Compiles one script. All six in a loop is what the pipeline's first gate does |
| `perl tests/<script>.pl` | One script's checks, each printing `ok` or `FAIL` and exiting non-zero on a failure |
| `perl tests/container/rig.pl` | The container rig, which must be run through Podman as root: see below |
The whole local gate, one line per script:
```sh
for f in *.pl; do perl -c "$f" || exit 1; done
for t in tests/*.pl; do perl "$t" || exit 1; done
```
Two more gates have no command of their own because the pipeline carries them: every
script opens with the shebang `#!/usr/bin/env perl` and the two-line licence header,
and no file in the repository uses a dash as punctuation. Both are checked in
`.gitea/workflows/test.yml`.
## The checks under tests/
One file per script, named after it, run as `perl tests/<script>.pl`. They cover the
pure functions, which is where the arithmetic and the parsing live: RPM's version
comparison (checked against `rpm`'s own implementation, which is the definition of the
ordering), the two JSON decoders and the encoder, IPv4 and IPv6 parsing with the
`sockaddr` packing and the address classification, the ping transcript parser, the
grading, the dnf-automatic renderer, the checksum reader, the writers, and the argument
parsing of every script.
They need no root, no network and no Podman, except where a case is skipped for that
reason and says so: `atomic_write` settles the owner of the file it writes, so it can
only be exercised as root.
## The container rig
`tests/container/rig.pl` runs the real `sglang-deploy.pl` as root inside a container
against a stub `PATH`: every command the script drives (`dnf`, `rpm`, `podman`,
`systemctl`, `curl`, `openssl`, `nginx`, `firewall-cmd`, `getsebool`, `lspci`) is a
stub that answers from a fixture, so a run is deterministic and needs no network and no
GPU. It covers the deploy end to end: the resolved image tag, the unit file, the nginx
configuration, the TLS certificate and its modes, the API key file, the SELinux
boolean, the firewall rule, the idempotent second run, the dry run, the uninstall, the
Radeon and MI300 paths, the offline and unpublished-tag failures, the argument
validation and the non-root refusal.
Prepare the platform image once, since the base images carry no Perl:
```sh
podman run --name sglang-prep registry.fedoraproject.org/fedora:44 dnf install -y perl
podman commit sglang-prep localhost/sglang-rig:fedora
podman rm sglang-prep
```
Then run the scenarios. The repository is mounted read-only at its own path and the
work directory is a scratch directory outside it:
```sh
podman run --rm --privileged \
-v "$PWD:$PWD:ro,z" \
-v /tmp/sglang-rig:/tmp/sglang-rig:Z \
localhost/sglang-rig:fedora perl "$PWD/tests/container/rig.pl"
```
The device nodes the script requires are faked by the rig itself, which is why the
container needs `--privileged`. The one scenario that checks the *missing* driver needs
a run without it:
```sh
podman run --rm \
-v "$PWD:$PWD:ro,z" \
-v /tmp/sglang-rig:/tmp/sglang-rig:Z \
localhost/sglang-rig:fedora perl "$PWD/tests/container/rig.pl" no_driver
```
The same rig was run on all three platforms, which is how the script's behaviour
against their package managers and their `os-release` was confirmed; only the image tag
in the commands changes:
| Platform | Image | Perl |
|---|---|---|
| Fedora 44 | `registry.fedoraproject.org/fedora:44` | 5.42.3 |
| CentOS Stream 10 | `quay.io/centos/centos:stream10` | 5.40.2 |
| openEuler 24.03 LTS | `docker.io/openeuler/openeuler:24.03-lts` | 5.38.0 |
A run leaves its log in `/tmp/sglang-rig/log/`, including `run.out`, the script's own
output, and `stubs.log`, every command it issued with its arguments. That log is the
quickest way to see what a run actually did.
## Running one script for real
The diagnostic scripts need no root and can simply be run:
```sh
perl network-diag.pl
perl system-diag.pl --json
```
The four that change a system need root and, except for
`workstation-setup.pl`, run in a disposable container the same way the rig does. Start
from the platform image, and read the summary before anything else: every one of them
supports `--dry-run`, which changes nothing and prints what it would do.
## Continuous integration
Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | Steps |
|---|---|---|
| Test | push or pull request to `development` | install Perl, then the syntax gate, the licence-header gate, the punctuation gate, and the six check files |
| Deploy | push to `main`, or dispatched by hand | write `SHA256SUMS`, pin the host key from `DEPLOY_KNOWN_HOSTS`, publish over rsync |
The container rig is deliberately not in the pipeline: it needs `--privileged`, and the
runner is a small box shared with the forge. A change that touches a system is expected
to be verified with the rig locally, and the pull request says which platform was used.
## Releases
There are none. Each script carries its own version string, and the deploy pipeline
publishes the scripts on every push to `main`, so a merge is a release and the commit
history is the record of what changed.