youngkingdom/vllm - vllm - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Lu Fang	07064cb1d4	[Bugfix] Check chain_speculative_sampling before calling it (#11673 ) Signed-off-by: Lu Fang <lufang@fb.com>	2025-01-02 16:58:56 -08:00
Sachin Varghese	2f1e8e8f54	Update default max_num_batch_tokens for chunked prefill (#11694 )	2025-01-03 00:25:53 +00:00
Nathan Azrak	68d37809b9	[Misc] Minimum requirements for SageMaker compatibility (#11576 )	2025-01-02 15:59:25 -08:00
wchen61	5dba257506	Resolve race conditions in Marlin kernel (#11493 ) Signed-off-by: wchen61 <wchen61@foxmail.com>	2025-01-02 22:58:56 +00:00
bjmsong	187e32997c	[Bugfix] Change kv scaling factor by param json on nvidia gpu (#11688 ) Signed-off-by: bjmsong <bjmsong@126.com> Co-authored-by: bjmsong <bjmsong@126.com>	2025-01-02 21:11:39 +00:00
Woosuk Kwon	b55ed6ef8a	[V1][Minor] Optimize token_ids_cpu copy (#11692 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-02 12:04:58 -07:00
Kathy Yu	2f385183f3	[Bugfix] Free cross attention block table for preempted-for-recompute sequence group. (#10013 ) Signed-off-by: Kathy Yu <feiyangyu@google.com>	2025-01-02 10:28:09 -08:00
Chunyang Wen	84c35c374a	According to vllm.EngineArgs, the name should be distributed_executor_backend (#11689 )	2025-01-02 18:14:16 +00:00
Cyrus Leung	8c38ee7007	[VLM] Merged multi-modal processor for LLaVA-NeXT (#11682 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-01-02 16:39:27 +00:00
Tobias Pitters	b6087a6bee	[mypy] Pass type checking in vllm/inputs (#11680 ) Signed-off-by: Tobias Pitters <tobias.pitters@gmail.com>	2025-01-02 16:18:15 +00:00
Cyrus Leung	23c1b10a4c	[VLM][Bugfix] Multi-modal processor compatible with V1 multi-input (#11674 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2025-01-02 17:00:00 +08:00
Cyrus Leung	a115ac46b5	[VLM] Move supported limits and max tokens to merged multi-modal processor (#11669 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: Isotr0py <2037008807@qq.com> Co-authored-by: Isotr0py <2037008807@qq.com>	2025-01-01 15:44:42 +00:00
Woosuk Kwon	73001445fb	[V1] Implement Cascade Attention (#11635 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2025-01-01 21:56:46 +09:00
Kazuhiro Serizawa	6d70198b17	[Doc] Fix typo (#11666 ) Signed-off-by: Kazuhiro Serizawa <nserihiro@gmail.com>	2025-01-01 08:10:10 +00:00
Lu Fang	f962f426bc	[Misc] Replace space with - in the file names (#11667 ) Signed-off-by: Lu Fang <lufang@fb.com>	2025-01-01 07:39:30 +00:00
Jee Jee Li	11d8a091c6	[Misc] Optimize Qwen2-VL LoRA test (#11663 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>	2025-01-01 14:42:23 +08:00
Cyrus Leung	365801fedd	[VLM] Add max-count checking in data parser for single image models (#11661 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: Roger Wang <ywang@roblox.com> Co-authored-by: Roger Wang <ywang@roblox.com>	2024-12-31 22:15:21 -08:00
Joe Runde	4db72e57f6	[Bugfix][Refactor] Unify model management in frontend (#11660 ) Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>	2025-01-01 02:21:51 +00:00
Yihua Cheng	0c6f998554	[Benchmark] Add benchmark script for CPU offloading (#11533 ) Signed-off-by: ApostaC <yihua98@uchicago.edu> Co-authored-by: KuntaiDu <kuntai@uchicago.edu>	2025-01-01 00:10:55 +00:00
Roger Wang	e7c7c5e822	[V1][VLM] V1 support for selected single-image models. (#11632 ) Signed-off-by: Roger Wang <ywang@roblox.com> Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk> Signed-off-by: Isotr0py <2037008807@qq.com> Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk> Co-authored-by: Isotr0py <2037008807@qq.com>	2024-12-31 21:17:22 +00:00
Chen Zhang	8c3230d8c1	[V1] Simpify vision block hash for prefix caching by removing offset from hash (#11646 )	2024-12-31 08:56:01 +00:00
sakunkun	2c5718809b	[Bugfix] Move the _touch(computed_blocks) call in the allocate_slots method to after the check for allocating new blocks. (#11565 )	2024-12-31 06:29:04 +00:00
John Giorgi	82c49d3260	[Misc][LoRA] Support Rank Stabilized LoRA (RSLoRA) (#6909 ) Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-30 22:15:58 -08:00
Michael Goin	74fa1d123c	[Bugfix] Fix OpenAI parallel sampling when using xgrammar (#11637 ) Signed-off-by: mgoin <michael@neuralmagic.com>	2024-12-31 03:43:54 +00:00
Matthias Vogler	a2a40bcd0d	[Model][LoRA]LoRA support added for MolmoForCausalLM (#11439 ) Signed-off-by: Matthias Vogler <matthias.vogler@joesecurity.org> Signed-off-by: Jee Jee Li <pandaleefree@gmail.com> Co-authored-by: Matthias Vogler <matthias.vogler@joesecurity.org> Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>	2024-12-30 17:33:06 -08:00
Kevin H. Luu	ccb1aabcca	[benchmark] Remove dependency for H100 benchmark step (#11572 )	2024-12-30 12:27:07 -08:00
whyiug	36e7670045	[Bugfix] Validate and concatenate image embeddings in MiniCPMVBaseModel (#11631 )	2024-12-30 18:51:04 +00:00
Robert Shaw	5886aa496e	[V1] [6/N] API Server: Better Shutdown (#11586 )	2024-12-30 15:51:02 +00:00
Cyrus Leung	8d9b6721e7	[VLM] Abstract out multi-modal data parsing in merged processor (#11620 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-30 15:01:35 +00:00
youkaichao	b12e87f942	[platforms] enable platform plugins (#11602 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-30 20:24:45 +08:00
Li, Jiang	5dbf854553	[CI/Build][CPU] Fix CPU CI by lazy importing triton FP8 kernels (#11618 ) Signed-off-by: jiang1.li <jiang1.li@intel.com>	2024-12-30 10:17:04 +00:00
Tyler Michael Smith	970d6d0776	[Build][Kernel] Update CUTLASS to v3.6.0 (#11607 ) Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>	2024-12-30 17:22:13 +08:00
Liangfu Chen	628ec6c17b	[Docker] bump up neuron sdk v2.21 (#11593 ) Signed-off-by: Liangfu Chen <liangfc@amazon.com>	2024-12-30 13:46:14 +08:00
youkaichao	3682e33f9f	[v1] fix compilation cache (#11598 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-30 04:24:12 +00:00
Michael Goin	0aa38d16f5	Remove print statement in DeepseekScalingRotaryEmbedding (#11604 )	2024-12-29 20:16:46 +00:00
Kuntai Du	faef77c0d6	[Misc] KV cache transfer connector registry (#11481 ) Signed-off-by: KuntaiDu <kuntai@uchicago.edu>	2024-12-29 16:08:09 +00:00
youkaichao	dba4d9dec6	[v1][bugfix] fix cudagraph with inplace buffer assignment (#11596 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-29 09:03:49 +00:00
Cyrus Leung	32b4c63f02	[Doc] Convert list tables to MyST (#11594 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-29 15:56:22 +08:00
Robert Shaw	4fb8e329fd	[V1] [5/N] API Server: unify `Detokenizer` and `EngineCore` input (#11545 ) Signed-off-by: rshaw@neuralmagic.com <rshaw@neuralmagic.com>	2024-12-28 20:51:57 +00:00
youkaichao	328841d002	[bugfix] interleaving sliding window for cohere2 model (#11583 ) Signed-off-by: youkaichao <youkaichao@gmail.com>	2024-12-28 16:55:42 +00:00
Cyrus Leung	d427e5cfda	[Doc] Minor documentation fixes (#11580 ) Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>	2024-12-28 21:53:59 +08:00
Woosuk Kwon	42bb201fd6	[V1][Minor] Set pin_memory=False for token_ids_cpu tensor (#11581 ) Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-12-28 13:33:12 +00:00
hj-wei	59d6bb4c86	[Hardware][AMD]: Replace HIPCC version with more precise ROCm version (#11515 ) Signed-off-by: hjwei <hjwei_xd@163.com>	2024-12-28 11:17:35 +00:00
Roger Wang	b7dcc003dc	[Model] Remove hardcoded image tokens ids from Pixtral (#11582 ) Signed-off-by: Roger Wang <ywang@roblox.com>	2024-12-28 10:54:23 +00:00
Isotr0py	d34be24bb1	[Model] Support InternLM2 Reward models (#11571 ) Signed-off-by: Isotr0py <2037008807@qq.com> Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>	2024-12-28 06:14:10 +00:00
Rajveer Bachkaniwala	b5cbe8eeb3	[Bugfix] Last token measurement fix (#11376 ) Signed-off-by: rajveerb <46040700+rajveerb@users.noreply.github.com> Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>	2024-12-28 11:34:46 +08:00
Robert Shaw	df04dffade	[V1] [4/N] API Server: ZMQ/MP Utilities (#11541 )	2024-12-28 01:45:08 +00:00
Chen Zhang	a60731247f	[Doc] Update mllama example based on official doc (#11567 ) Signed-off-by: Chen Zhang <zhangch99@outlook.com>	2024-12-28 00:31:10 +00:00
Selali	ac79799403	[Bugfix] Fix for ROCM compressed tensor support (#11561 )	2024-12-27 20:12:11 +00:00
Isotr0py	dde1fa18c9	[Misc] Improve BNB loader to handle mixture of sharded and merged weights with same suffix (#11566 ) Signed-off-by: Isotr0py <2037008807@qq.com>	2024-12-27 19:45:13 +00:00

1 2 3 4 5 ...

3987 Commits