volta
Here are 52 public repositories matching this topic...
Hand-written NVFP4 W4A16 CUDA kernels for Volta
-
Updated
Aug 19, 2026 - Python
PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.
-
Updated
Aug 28, 2026 - C++
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
-
Updated
Jul 10, 2026 - Python
Zsh plugin to seamlessly install and configure volta
-
Updated
May 18, 2021 - Shell
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
-
Updated
Aug 23, 2026 - Shell
📦 A fully automated method for installing Nvidia drivers on Arch Linux
-
Updated
Jun 19, 2026 - Shell
FlashAttention brought back to Tesla V100 — a deep llama.cpp fork: SM 7.0 D256 kernels, SplitKV3, q4_0 KV cache, DFlash2 speculative decoding and multimodal fixes.
-
Updated
Aug 26, 2026 - C++
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
-
Updated
Aug 14, 2026 - Python
Run Qwen3.6-27B on four Tesla V100s at 366 tok/s using hand-written NVFP4 CUDA kernels and chain-MTP speculation.
-
Updated
Aug 28, 2026 - Python
⚡ A Docker image for Volta
-
Updated
Jul 27, 2026 - Dockerfile
-
Updated
Sep 28, 2023 - TypeScript
ZSH Plugin to install and load Volta: JS Toolchains as Code. ⚡
-
Updated
Jun 29, 2021 - Shell
Improve this page
Add a description, image, and links to the volta topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the volta topic, visit your repo's landing page and select "manage topics."