1
0
Fork 0
gnu-plus-dotfiles/dot_agents/skills/strix-fleet/GOTCHAS.md

97 lines
5.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Fleet gotchas — the ones that cost hours
Each: symptom → cause → fix. Read before bench re-runs, config edits, downloads, rebuilds.
## 1. Bench silently "completes" with no new data
- **Symptom:** `--all` prints "ALL TIERS COMPLETE" instantly, no GPU activity, no new JSON.
- **Cause:** `--all` skips tiers recorded completed in `results/state.json`; tier JSONs still on disk.
- **Fix:** BOTH steps before re-running:
`rm -f results/<model>_tier*.json` AND strip that model's entries from `results/state.json`
(`completed` list; see SKILL.md §4 re-run protocol snippet).
## 2. `error: invalid argument: --load-mode` / "upstream command exited prematurely"
- **Symptom:** model never appears ready; llama-swap `/logs` shows upstream exit right after spawn.
- **Cause:** flag sets differ between fork bases. `--load-mode mmap` exists only on modern
builds; older ROCmFPX PR#98-era builds want `--mmap`. Similar drift for other flags.
- **Fix:** `<binary> --help | grep <flag>` before writing every profile. Never copy a flag
block between builds blindly.
## 3. Missing pre-created directories (`--slot-save-path`, mmproj paths)
- **Symptom:** spawn fails, upstream exits; `/logs` may show path error.
- **Cause:** llama-swap spawns the binary directly — no mkdir. CIRU `--slot-save-path
/tmp/ciru-strix-slots` and any mmproj/sidecar path must pre-exist.
- **Fix:** `mkdir -p` every directory referenced in a `cmd:` before restart.
## 4. Global LD_LIBRARY_PATH poisons the ROCmFP4 fork
- **Symptom:** ROCmFPX/ROCmFP4 binary loads wrong ggml/llama libs, crashes or picks wrong backend.
- **Cause:** the fork binary has its own RUNPATH; a systemd-global LD_LIBRARY_PATH shadows it.
- **Fix:** LD_LIBRARY_PATH belongs INSIDE the macro/profile: `env LD_LIBRARY_PATH=<path> <binary>`.
Never in the systemd unit or shell profile.
## 5. Model listed but broken
- **Symptom:** id present in `/v1/models`, requests hang or 500.
- **Cause:** llama-swap lists models even when files are missing/deleted; failure only surfaces at spawn.
- **Fix:** `curl -s http://127.0.0.1:9292/logs | tail -50` FIRST, before any other debugging.
## 6. `api_chat()` sends `cache_prompt: False`
- Keep it. Matches CIRU README eval methodology; harmless elsewhere. Changing it makes
T2 numbers incomparable to the existing fleet results.
## 7. `uv cache clean` kills the bench venv
- **Symptom:** `ModuleNotFoundError: datasets` mid-T2.
- **Fix:** `uv pip install --python ~/.hermes/hermes-agent/venv/bin/python3 datasets`
(verify with `…venv/bin/python3 -c "import datasets"`).
## 8. Big models OOM T1 at pp32768+
- llama-bench long prompt-processing frontiers blow up VRAM+KV on 100 GB-class models.
- **Fix:** skip the long frontiers, or run llama-bench manually with short frontiers and hand-write
`results/<model>_tier1.json` matching the schema (see SKILL.md §4).
## 9. Truncated downloads
- **Symptom:** GGUF loads garbage / checksum mismatch / corrupt metadata.
- **Cause:** aria2c can report "done" on a truncated file.
- **Fix:** verify `stat` size vs HF API, README checksums (CIRU-STRIX: ALL 5), GGUF header via
`scripts/gguf_walk.py`. Never skip.
## 10. Two benches at once
- Single GPU — concurrent benches OOM/destroy timings.
- **Fix:** `pgrep -fa bench_framework.py` before launch; `pkill -f bench_framework.py` + re-check
(llama-bench children can survive) before relaunch.
## 11. MTP self-draft flags
- Custom ROCmFPX quants with self-draft MTP need `-md <file>` pointing at the SAME gguf.
- MTP benefit shows in live-server timings only, NOT in llama-bench T1 — don't chase it there.
## 12. T4 HumanEval fake 100%
- **Rule:** verify `total == 164` in `results/<model>_tier4.json`. Smaller total or 100% with tiny
details list = silent skip bug; re-run. Errors count as failures by design.
## 13. Thinking toggles
- ON for T4 (HumanEval) + GSM8K (standing user rule). OFF for MMLU/ARC/HellaSwag and T6 IFEval.
- Wrong toggle collapses scores to ~1030% and poisons the comparison table.
## 14. Phantom model ids
- Never fabricate/guess model ids in commands, configs, or reports. Check live `/v1/models`
on the target host whenever an id is in doubt.
## 15. Disk exhaustion
- miche /home ~1.6 T (87 % used 2026-08-29); byte /home 931 G at **99 % — needs attention**.
- Purge retired model files immediately after retirement ("head-to-head, delete losers").
Check `df -h /home` on both hosts before any download.
## 16. CIRU-STRIX bench specifics
- T1 MUST use `bench_bin: api_timing` — `--ple-sidecar` is server-only, llama-bench cannot load it.
- Mid-bench decode errors (positions-decreasing / speculative-batch 500s): check the HF repo
discussions for a runtime fix BEFORE rebuilding anything.
## 17. Hermes config is write-blocked
- `~/.hermes/config.yaml` on miche rejects agent writes. Use `hermes config set …`.
- Provider entries stay discover-only: `discover_models: true`, no static `models:` dict.
## Version pins (verify, don't trust)
- CIRU runtime: tag `v1.1` = commit `baba5e0` (`build-gfx1151/bin/llama-server --version`).
**Never `v1.0.0-h121`** — MTP+cache-prefix state sync bug → 500 "sequence 0 positions are
decreasing" / "failed to process speculative batch".
- qwen4exp arch: upstream PR #27742 is base-model only, NO MTP. Flash-Next MTP lives only in the
CIRU runtime (LaurentZuijdwijk fork deleted 2026-08-29).
- ROCmFPX flagship tree has uncommitted patches — never `git reset`/`pull`; worktree for updates.