1
0
Fork 0

Add strix-fleet omp skill (fleet management + bench runbook)

This commit is contained in:
Rain 2026-08-29 20:28:45 -04:00
parent 3414885fda
commit 5ad11db70c
3 changed files with 340 additions and 0 deletions

View file

@ -0,0 +1,213 @@
---
name: strix-fleet
description: Manage and benchmark the two-host local LLM fleet (miche Strix Halo + byte R9700) behind llama-swap. Use when SSHing to miche/byte, editing llama-swap config, adding/retiring models, serving via the 9292 proxy, running bench_framework tiers, regenerating or pushing the fleet bench page, or debugging spawn/OOM/bench-state failures.
---
# Strix Halo Fleet Management & Bench
Two-host local LLM fleet behind llama-swap. All work happens over SSH from this machine.
Read `GOTCHAS.md` in this skill directory before ANY bench re-run or config change — 17 failure modes that cost hours, each with symptom + fix.
## 1. Hosts, access, services
| Host | SSH | Hardware | User | Config |
|---|---|---|---|---|
| miche | `ssh rain@miche.local` | Ryzen AI MAX+ 395 "Strix Halo", Radeon 8060S gfx1151, 128 GB unified | rain (home `/home/rain`) | `/home/rain/AI/llama-swap/config.yaml` |
| byte | `ssh melon@byte.local` | R9700 + RX 9070XT, 32 GB VRAM | melon | `/home/melon/AI/llama-swap/config.yaml` |
Both run llama-swap as a **systemd USER service** on `0.0.0.0:9292` (OpenAI-compatible:
`/v1/models`, `/v1/chat/completions`, `/v1/completions`, `/health`, `/logs`, `/metrics`,
`/api/models/unload`).
```bash
# after any config.yaml edit (NO --watch-config; restart is mandatory):
ssh rain@miche.local 'systemctl --user restart llama-swap'
# verify:
ssh rain@miche.local 'curl -s http://127.0.0.1:9292/v1/models'
# spawn failures appear ONLY here (llama-swap lists models even when files are gone):
ssh rain@miche.local 'curl -s http://127.0.0.1:9292/logs | tail -50'
# unload a model without waiting for TTL:
ssh rain@miche.local 'curl -s -X POST http://127.0.0.1:9292/api/models/unload'
```
Remote login shells are **zsh**: `echo ===` fails (`=cmd` expansion). Quote delimiters.
### Hermes (on miche)
`~/.hermes/config.yaml` is **agent-write-BLOCKED** — use `hermes config set` for provider changes.
Providers: `custom:llama-swap``http://localhost:9292/v1`, `custom:byte-llama`
`http://byte.local:9292/v1`, both `discover_models: true`, NO static `models:` dict.
**Never re-add a hardcoded `models:` dict** — it pins the catalog and goes stale.
Hermes auto-discovers from `/v1/models` on every model-picker run.
## 2. Fleet inventory
Refresh (single source of truth — never trust the table below for debugging):
```bash
for h in rain@miche.local melon@byte.local; do
echo "== $h"; ssh -o BatchMode=yes $h 'curl -s http://127.0.0.1:9292/v1/models | python3 -c "import json,sys;[print(m[\"id\"]) for m in json.load(sys.stdin)[\"data\"]]"'
done
```
Verified 2026-08-29 (miche: 21 ids, byte: 8 ids).
**miche chat LLMs:**
| Model id | What | Engine | Ctx |
|---|---|---|---|
| `qwen38-27b-ciru-promptforge` | CIRU v3 ActiveFPX PromptForge 27B dense, **FLAGSHIP**, vision via BF16 mmproj | ROCmFPX build-promptforge (`qwen38-ciru-wrapper.sh`) | 262K |
| `qwen38-flash-ciru-strix-iu4` | Qwen3.8 Flash CIRU-STRIX-IU4, 125B MoE Qwen4 arch. **The ONLY Flash-Next.** Aliases: `qwen38-flash-next-iq4xs`, `-rocmfp4-fast`, `-apex`, `-apex-middle` | CIRU runtime v1.1 `baba5e0` (`ciru_strix_server` macro) | 262K |
| `qwen36-35b-heretic-apex` | Heretic 35B APEX MTP | stock Vulkan | 256K |
| `ornith-1.5-apex-i-quality` | Ornith 1.5 Heretic-MTP (SC117 shisa graft), fastest chat 7884 t/s | stock Vulkan | 262K |
| `gemma4-26b-hauhau` | Gemma4 26B A4B | stock Vulkan | 128K |
| `godoter-27b` | 27B | stock Vulkan | — |
| `muse-glimmer-30b-rocmfpx` | 30B | ROCmFPX build-strix-rocmfp4 (`muse-glimmer-rocmfpx-wrapper.sh`) | — |
**miche utilities:** embedding `qwen3-embedding-0.6b`, reranker `qwen3-reranker-0.6b`,
STT `stt`/`stt_parakeet`/`stt_qwenasr`/`qwenasr`/`asr`, TTS `tts_qwen3`,
image `img-klein9b`/`img-moodymix`/`img-qwen-edit`/`img-wai-anima` + `dall-e-2`/`dall-e-3` fronts.
(`music-ace`/`vid-wan2` were listed in an older inventory — NOT in live `/v1/models` 2026-08-29; check live before assuming.)
**byte (8):** `qwen36-35b-heretic-apex`, `ornith-1.5-apex-i-quality`, `qwen38-27b-apex`,
`gemma4-26b-hauhau`, `godoter-27b`, `toriigate-0.5` (vision, 16K), embedding, reranker.
**Retired 2026-08-29, files deleted** (bench results stay on disk, shown in page Retired table):
all other Flash-Next tiers (UD-IQ4_XS, ROCmFP4-FAST, ROCmFP4-STRIX, APEX-MIDDLE), Kairic-Edge,
711 family, Genesis, Ornith-DualView, DwarfStar, DS4 variants, Ling-3.0-flash, Laguna family.
## 3. Config conventions
miche config layout: `macros:` then per-model profiles. Match existing entry style
(byte style differs slightly — variable names differ, follow byte's file).
Key macros (see `GOTCHAS.md` #4, #2, #3 before touching):
- `llama_server`: **stock Vulkan daily driver**`env LD_LIBRARY_PATH=/opt/llama.cpp-vulkan-gfx1151/lib <bin> …` — LD_LIBRARY_PATH lives INSIDE the macro, never global (shadows ROCmFP4 fork RUNPATH).
- `ciru_strix_server`: CIRU runtime v1.1 (IU4 WMMA + NVMe-paged FP8 PLE + MTP depth 3).
- `rocmfpx_server` / `dualview_server` / `heretic_server` / `ds4_server`: wrapper scripts `~/AI/llama-swap/*.sh` setting per-build LD_LIBRARY_PATH.
- `server_common` (`--host 0.0.0.0 --no-webui --timeout 7200`), `llama_common` (`-ngl 999 --no-mmap -fa 1 --jinja`), `rocm_common` (no `--jinja` — ROCmFPX needs `--chat-template-file`), `kv_moe`/`kv_dense`, `sampler_think*`, `tpl_qwen`/`tpl_gemma4`, `mmproj_*`.
- Globals: `healthCheckTimeout: 900`, `globalTTL: 900`, `logToStdout: proxy`, `includeAliasesInList: true`, `performance.every: 15s`.
Profile pattern: `name:`, `description:`, `cmd: >-` composing macros, one `--alias <id>` per
alias, `capabilities: {in: [text, image], out: [text], context: <n>}` block matching
`--ctx-size`. CIRU-STRIX entry additionally: `--ple-sidecar <dir>` (mandatory),
`--ple-cache-mib 4096`, `--slot-save-path /tmp/ciru-strix-slots` (**dir must pre-exist**),
`--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-split 0.10`.
### Add / retire = THREE places
1. llama-swap `config.yaml` entry (+ `mkdir -p` any slot/sidecar dirs) → `systemctl --user restart llama-swap` → verify `/v1/models` + `/logs`.
2. `~/AI/strix-halo-fleet-bench/bench_framework.py` `MODELS` dict (fields: `model_file`, `mmproj`, `template`, `bench_bin`, `has_mtp`, `card_name`, `params`, `size_gb`, `base_arch`, `color`; `bench_bin` ∈ default vulkan | `rocmfpx` | `dualview` | `ciru-hip` | `api_timing`; `api_only: True` for external-API models).
3. `~/AI/strix-halo-fleet-bench/gen_page.py`: `META` (+ `SPECS` as `(total_B, active_B, file_GB)`, `ORDER`).
Retiring = reverse all three + move id from `ORDER` to `ORDER_RETIRED` + **delete model files
immediately** (disk is finite; byte was 99%). Never fabricate model ids — trust live `/v1/models`.
## 4. Bench framework (miche `~/AI/strix-halo-fleet-bench/`)
```bash
cd ~/AI/strix-halo-fleet-bench
python3 bench_framework.py --model <id> --all # all tiers
python3 bench_framework.py --model <id> --tier <1|2|3|4|6>
python3 bench_framework.py --status # completion state
python3 bench_framework.py --model <id> --resume # resume interrupted
python3 bench_framework.py --model <id> --tier 1 --mtp-off
```
- Tier 2 imports `datasets` lazily — run under the venv that has it:
`~/.hermes/hermes-agent/venv/bin/python3` (verify `import datasets` first; `uv cache clean` breaks it → `uv pip install --python ~/.hermes/hermes-agent/venv/bin/python3 datasets`).
- External-API models (`deepseek-v4-*`, `glm-5.2`) carry `api_only: True` in `MODELS`; `--all` then runs tiers 2/3/4/6 only (no CLI flag for this).
- **No concurrent benches — single GPU.** `pgrep -fa bench_framework.py` first; kill zombies `pkill -f bench_framework.py` and verify dead.
- `api_chat()` sends `cache_prompt: False` globally — keep it (CIRU eval methodology).
- state.json is flock'd; `--all` SKIPS tiers marked completed.
Tiers: **T1** speed (llama-bench frontiers pp/tg short/medium/…; `api_timing` for sharded >50 GB and ANY custom-runtime/PLE model — `ple_sidecar` is server-only, llama-bench cannot load it); **T2** knowledge MMLU 200 / ARC 237 / GSM8K 200 / HellaSwag 300 (thinking OFF for MC, ON for GSM8K); **T3** NIAH long-context, cap comparison at ctx ≤ 131072; **T4** HumanEval 164 (thinking ON, 8192 tok, errors = failures); **T6** IFEval (thinking OFF). **Standing rule: thinking ON for T4 + GSM8K** — off collapses scores to ~1030%.
Results on disk: `results/<model>_tier<N>.json` + `results/state.json` (`{"completed":[{model,tier,completed,error}]}`).
Schemas (verified 2026-08-29):
- T1 `{model, tier, mtp, results:[{label, pp, type: pp|tg, avg_ts, stddev_ts, model_n_params, model_size, backend}], timestamp}`; api_timing rows add `method: "api_timing"`.
- T2 `{model, tier, benchmarks: {mmlu, arc_challenge, gsm8k, hellaswag}, timestamp}`.
- T3 `{model, tier, tests: {single_niah[], single_niah_summary[], multi_needle[], variable_tracking, aggregation[]}, timestamp}`.
- T4 `{model, tier, "pass@1", passed, total, details[], timestamp}`**verify `total == 164`**; a fake 100% = silent skip bug.
- T6 `{model, tier, accuracy, passed, total, details[], timestamp}`.
### Re-run protocol (mandatory, gotcha #1)
```bash
rm -f results/<model>_tier*.json
python3 - <<'EOF'
import json
p = "results/state.json"
s = json.load(open(p))
s["completed"] = [e for e in s["completed"] if e["model"] != "<model>"]
json.dump(s, open(p, "w"), indent=1)
EOF
```
Forgetting state.json = instant "ALL TIERS COMPLETE" with zero new data.
### Page
```bash
python3 gen_page.py # reads results/*.json -> fleet_bench.md + fleet_bench.html (real numbers off disk only)
scp fleet_bench.html melon@byte.local:~/Desktop/fleet_bench.html
```
Keep `META` / `SPECS` / `ORDER` (active) / `ORDER_RETIRED` in sync; reconcile `ORDER`
against live `/v1/models` when refreshing (it drifts). External-API models live in
`META_FRONTIER`/`ORDER_FRONTIER`.
## 5. Engine map
| Binary / build | Serves | Access |
|---|---|---|
| `/opt/llama.cpp-vulkan-gfx1151/bin/llama-server` (stock Vulkan, nisavid pkg) | gemma4-26b-hauhau, godoter-27b, qwen36-35b-heretic-apex, ornith-1.5-apex-i-quality, embedding, reranker | `llama_server` macro |
| `~/AI/Qwen3.8-Flash-CIRU-STRIX-IU4-runtime/build-gfx1151` (v1.1 `baba5e0`) | qwen38-flash-ciru-strix-iu4 | `ciru_strix_server` macro |
| `~/AI/ROCmFPX/build-promptforge` | qwen38-27b-ciru-promptforge | `qwen38-ciru-wrapper.sh` |
| `~/AI/ROCmFPX/build-strix-rocmfp4` (PRIMARY fork build: bailingmoe3, dflash, spec checkpoint restore) | muse-glimmer-30b-rocmfpx | `muse-glimmer-rocmfpx-wrapper.sh` |
| `~/AI/ROCmFPX/build-laguna`, `build-dualview`, `build-ciru` | Laguna/DualView/legacy — models retired, builds kept for benches | `dualview-wrapper.sh`, `rocmfpx-wrapper.sh`, `heretic-wrapper.sh`, `ds4-wrapper.sh` |
| `~/AI/llama-swap/.venv/bin/python3 ~/AI/llama-swap/*-server.py` | img-*, stt_*, tts_qwen3 | direct python |
Bench binaries (`bench_bin` in `MODELS`): default `/opt/llama.cpp-vulkan-gfx1151/bin/llama-bench`
(libs `…/lib`); `rocmfpx` `~/AI/ROCmFPX/build-strix-rocmfp4/bin/llama-bench`; `dualview`
`build-dualview`; `ciru-hip` build-ciru; `api_timing` through llama-swap at
`http://localhost:9292/v1` — REQUIRED for sharded >50 GB and PLE-sidecar models.
**Never `git reset/pull` the flagship ROCmFPX tree** (uncommitted patches) — use a worktree.
CIRU runtime build: `ROCM_ROOT=/opt/rocm ./scripts/ciru/build-linux-amd.sh`, tag `v1.1`
(commit `baba5e0`); **never `v1.0.0-h121`** (MTP+cache-prefix state-sync bug → 500
"sequence 0 positions are decreasing" / "failed to process speculative batch"). Verify:
`build-gfx1151/bin/llama-server --version` must show `baba5e0`.
qwen4exp arch merged upstream (PR #27742) base-only, NO MTP; MTP for Flash-Next exists only
in the CIRU runtime (LaurentZuijdwijk fork deleted — no longer needed).
## 6. Verification checklist — every operation
| Operation | Verify |
|---|---|
| Config edit + restart | `systemctl --user is-active llama-swap` + `/v1/models` shows new/changed ids + `/logs` clean (no "exited prematurely") |
| Model add | id in `/v1/models`; warm-load: 1-token chat completes; slot/sidecar dirs exist; `--help` flags valid for that binary |
| Bench re-run | state.json entries removed AND `rm results/<model>_tier*.json` before launch; after: each `results/<model>_tier*.json` fresh mtime, T2 has all 4 benchmarks, T4 `total==164`, T6 `total==541`-ish non-zero |
| New runtime build | `llama-server --version` commit matches intended tag |
| Download | `stat` size vs HF API + README checksums (CIRU: ALL 5) + GGUF header via `scripts/gguf_walk.py`; aria2c can report done truncated |
| Retire | id gone from `/v1/models`, `MODELS`, `ORDER`; present in `ORDER_RETIRED`; files deleted; results/*.json kept |
| Page push | `python3 gen_page.py` exits 0, scp returns, `ssh melon@byte.local ls -la ~/Desktop/fleet_bench.html` newer than generation |
| Bench hygiene | `pgrep -fa bench_framework.py` empty before starting |
## 7. One-command fleet status
`scripts/fleet-status.sh` in this skill dir. Prints: service state, model count + ids, disk,
bench state, running benches, on both hosts.
## 8. "Full bench new model" runbook
1. **Download** GGUF (+mmproj / MTP draft / PLE sidecar). Verify every README checksum, size vs HF API, GGUF header (`scripts/gguf_walk.py`).
2. **Custom runtime?** Build (CIRU: `ROCM_ROOT=/opt/rocm ./scripts/ciru/build-linux-amd.sh`), verify `--version` commit. Check upstream/HF discussions for runtime fixes before rebuilding anything mid-bench (gotcha #16).
3. **llama-swap**: add profile matching file style; `mkdir -p` slot/sidecar dirs; check flag set vs `--help` (`--load-mode mmap` vs `--mmap` differs per fork base, gotcha #2); restart; verify `/v1/models` + `/logs`.
4. **Warm-load** through proxy: 1-token `/v1/chat/completions`, expect full load (healthCheckTimeout 900 covers MoE ~2 min loads).
5. **Framework**: add `MODELS` entry (`bench_bin`: `api_timing` if PLE/sharded); clear any stale state (`--all` skips completed).
6. **Bench**: no other bench running → `python3 bench_framework.py --model <id> --all`.
7. **Validate results**: T4 `total==164`, T2 4 benchmarks, T3 rows non-empty.
8. **Page**: add META (+SPECS, ORDER) → `python3 gen_page.py``scp fleet_bench.html melon@byte.local:~/Desktop/fleet_bench.html` → verify mtime on byte.
9. **Report**: speed TG/PP, MMLU, ARC, GSM8K, HellaSwag, NIAH, HumanEval, IFEval, vs-fleet comparison table (condensed names, params only; A4B = MoE, no suffix = dense — never mislabel), page pushed to byte.
Head-to-head quants: bench both, delete the loser; lighter quant wins if quality holds.
API benches: verify API keys first; source `.env` in Python, not bash.