Popular repositories Loading
-
dgx-spark-nvfp4-serving
dgx-spark-nvfp4-serving PublicGuide for serving fine-tuned Qwen3.5-27B (dense, NVFP4) on DGX Spark via native vLLM. Includes critical config fixes for modelopt export_hf_checkpoint() that prevent silent FP32 dequantization.
Python 2
-
design-challenger
design-challenger PublicAdversarial design review agent — spawns writer + challenger Claude Code agents to stress-test specs and implementation plans for any codebase
TypeScript
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
-
mac-studio-mlx-serving
mac-studio-mlx-serving PublicBenchmarks and serving notes for Qwen3.5-27B fine-tune on Apple Silicon via MLX. Companion to dgx-spark-nvfp4-serving.
Python
-
autokernel
autokernel PublicForked from RightNow-AI/autokernel
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
Python
If the problem persists, check the GitHub status page or contact support.