youngkingdom/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Woosuk Kwon	3f942acfe1	Fix latency benchmark script (#118 )	2023-05-22 17:03:40 -07:00
Woosuk Kwon	42f1042e1c	Enhance SamplingParams (#96 )	2023-05-11 15:45:30 -07:00
Zhuohan Li	27f1410d06	New weight loader without np copy (#52 )	2023-05-03 15:32:04 +08:00
Zhuohan Li	4858f3bb45	Add an option to launch cacheflow without ray (#51 )	2023-04-30 15:42:17 +08:00
Woosuk Kwon	0f4b32199e	Support various block sizes & Change default block size to 16 (#38 )	2023-04-15 09:03:24 -07:00
Woosuk Kwon	84eee24e20	Collect system stats in scheduler & Add scripts for experiments (#30 )	2023-04-12 15:03:49 -07:00
Siyuan (Ryans) Zhuang	e3cec88aa5	Memcpy kernel for flash attention (#29 ) * optimize * add benchmark * add assert * add test	2023-04-10 18:22:49 -07:00
Woosuk Kwon	ee88a7e5f3	Add an option to use dummy model weights (#33 )	2023-04-08 23:36:12 -07:00
Woosuk Kwon	c267b1a02c	Add query stride to multi_query_cached_kv_attention & Add kernel benchmark script (#27 ) * Add query stride to multi_query_cached_kv_attention * Add kernel benchmark script	2023-04-08 13:36:09 -07:00
Woosuk Kwon	0f40557af6	Implement block copy kernel to optimize beam search (#32 )	2023-04-07 17:45:07 -07:00
Woosuk Kwon	12659a0bd7	Add CUDA graph-based all reduce launcher (#26 )	2023-04-05 11:16:57 -07:00
Zhuohan Li	c45f3c3ab6	Optimize tensor parallel execution speed (#17 )	2023-04-01 00:51:08 +08:00