Logo
Explore Help
Sign In
youngkingdom/vllm
1
0
Fork 0
You've already forked vllm
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
Files
9da25a88aa35da4b5ad7da545e6189e08c5f52f4
vllm/docs/source/quantization
History
Stas Bekman 98c12cffe5 [Doc] fix the autoAWQ example (#7937)
2024-08-28 12:12:32 +00:00
..
auto_awq.rst
[Doc] fix the autoAWQ example (#7937)
2024-08-28 12:12:32 +00:00
bnb.rst
[bitsandbytes]: support read bnb pre-quantized model (#5753)
2024-07-23 23:45:09 +00:00
fp8_e4m3_kvcache.rst
[Core/Bugfix] Add FP8 K/V Scale and dtype conversion for prefix/prefill Triton Kernel (#7208)
2024-08-12 22:47:41 +00:00
fp8_e5m2_kvcache.rst
[Core/Bugfix] Add FP8 K/V Scale and dtype conversion for prefix/prefill Triton Kernel (#7208)
2024-08-12 22:47:41 +00:00
fp8.rst
[Doc] Add docs for llmcompressor INT8 and FP8 checkpoints (#7444)
2024-08-16 13:59:16 -07:00
int8.rst
[Doc] Add docs for llmcompressor INT8 and FP8 checkpoints (#7444)
2024-08-16 13:59:16 -07:00
supported_hardware.rst
[Doc] Update quantization supported hardware table (#7595)
2024-08-16 13:59:27 -07:00
Powered by Gitea Version: 1.24.2 Page: 148ms Template: 4ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API