6.1k 查看 · 0 喜欢 · 55 收藏 · 2026-06-24 更新
gguf-quantization
GGUF格式和llama.cpp量化用于高效的CPU/GPU推理。在将模型部署到消费级硬件、Apple Silicon上,或需要2-8位灵活量化且无需GPU要求时使用。
GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.
