summaryrefslogtreecommitdiff
path: root/snippets/hyperstack/.gitignore
diff options
context:
space:
mode:
authorPaul Buetow <paul@buetow.org>2026-03-18 16:50:38 +0200
committerPaul Buetow <paul@buetow.org>2026-03-18 16:50:38 +0200
commitd3821c76ecd18bf6256d7493596c304fff784d29 (patch)
tree4c940bbff57ba48ede1d057c6803aa03635f8bc3 /snippets/hyperstack/.gitignore
parente9f57c66ba76b11e11a715c112e35394386a7831 (diff)
Add extra_vllm_args support; fix nemotron-super to real 120B; add deepseek-r1-32b, qwen3-32b, devstral presets
- hyperstack.rb: add extra_vllm_args array field to preset resolver and vllm_install_script; flags are appended verbatim to the docker run command, enabling per-preset vLLM flags (reasoning parsers, Mistral loader) - hyperstack.rb: show extra_args in dry-run model switch output - hyperstack-vm.toml: fix nemotron-super to use actual NVIDIA Nemotron-3-Super-120B-A12B AWQ (cyankiwi) with trust_remote_code=true; previous preset incorrectly used llama-3.3-70b - hyperstack-vm.toml: add deepseek-r1-32b (--reasoning-parser deepseek_r1, ~18 GB) - hyperstack-vm.toml: add qwen3-32b (--reasoning-parser deepseek_r1, ~18 GB) - hyperstack-vm.toml: add devstral (Mistral tokenizer+config format, ~15 GB); --load_format mistral omitted because AWQ weights are in standard HF safetensors format All 6 new/updated presets end-to-end tested on A100 80GB (vLLM 0.17.1). Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Diffstat (limited to 'snippets/hyperstack/.gitignore')
0 files changed, 0 insertions, 0 deletions