| Age | Commit message (Collapse) | Author |
|
Switch VM1 from n3-H100x1 to n3-H100x2 to run Nemotron-3-Super with
1M token context window via tensor parallelism. The dual-GPU setup
(160 GB total VRAM) provides enough KV cache headroom to override the
model's config.json limit of 262144 tokens.
Key changes:
- flavor_name: n3-H100x1 → n3-H100x2
- tensor_parallel_size: 1 → 2
- max_model_len: 131072 → 1048576 (with VLLM_ALLOW_LONG_MAX_MODEL_LEN=1)
- gpu_memory_utilization: 0.92 → 0.85 (headroom for Mamba cache + sampler warmup)
- Remove --enforce-eager: no longer needed with dual-GPU VRAM budget
- Disable prefix caching: on NemotronH it forces Mamba "all" cache mode
which pre-allocates states for all max_num_seqs and OOMs before the
sampler warmup pass; per-request allocation is cheaper at startup
Add two new vllm config fields to hyperstack.rb:
- extra_docker_env: passes -e KEY=VALUE flags to Docker before the image
name (used for VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 and
PYTORCH_ALLOC_CONF=expandable_segments:True)
- enable_prefix_caching: makes --enable-prefix-caching conditional
(default true for backward compat; false for NemotronH)
Both fields are supported in [vllm] defaults and [vllm.presets.*]
overrides with the same fallback semantics as existing fields.
Update pi/agent/models.json: Nemotron vm1 entry renamed to
"Nemotron 3 Super 120B 1M [vm1]" with contextWindow 1048576.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
|
Implements two LLM-callable tools using DuckDuckGo's free HTML
endpoint — no API key or account required:
- web_search: searches DDG, returns up to 8 results with title/URL/snippet
- web_fetch: fetches a URL, strips script/style/nav blocks, returns up
to 12,000 chars of readable text
Both tools are TypeScript registered via pi.registerTool() and load
automatically from pi/agent/extensions/web-search/ via the ~/.pi symlink.
HTTP requests use Node.js built-ins (node:https) with a 15s timeout and
single-redirect follow.
Also adds an Extensions section to README.md listing all extensions
with a focused web-search subsection.
|
|
- hyperstack-vm.toml: switch [vllm] default from Qwen3-Coder-Next to
openai/gpt-oss-120b (container_name, max_model_len=131072,
tool_call_parser=''); labels already reflected gpt-oss-120b
- pi/agent/models.json: add 'hyperstack' provider pointing at
hyperstack.wg1:11434/v1 with GPT-OSS 120B as primary model and all
preset models registered (alongside hyperstack1/hyperstack2)
- hyperstack.fish: add pi-hyperstack abbreviation for single-VM GPT-OSS 120B
- README.md: update fish abbreviations table, provider table, VM config
table, and Single-VM setup section to reflect the new defaults
|
|
Merged all still-relevant content from vllm-setup.txt into README.md:
- Why vLLM over Ollama section
- Full monitoring commands with engine metrics table
- Troubleshooting table
- VRAM sizing guide
- Performance characteristics table
Dropped LiteLLM, Anthropic API, Claude Code, and OpenCode sections
which are no longer applicable. Removes the vllm-setup.txt file.
|
|
|
|
|