Merge QKV into one linear layer #15

zhuohan123 · 2023-03-30T17:04:35Z

@WoosukKwon Feel free to merge this after your review.

WoosukKwon

Thanks for your effort. Please check my comments.

cacheflow/models/opt.py

cacheflow/models/llama.py

WoosukKwon · 2023-04-02T05:19:15Z

The performance regression problem in this PR is fixed in #20 . I will merge the two PRs together when PR #20 is approved.

…no-model-executor-opt [CPU] Avoid copy result and force allocation

This PR updates our grpc_server to add TGIS-style logs similar to https://github.com/IBM/text-generation-inference/blob/main/router/src/grpc_server.rs#L504-L512 This also disables the vllm per-request logging so that we don't double-log each request The timing info collected here is pretty rough, it doesn't plumb into the LLMEngine, it just times the generators to get the total time spent in the engine. We could do better, but this is a start. Example logs: ``` INFO 04-09 21:51:01 logs.py:43] generate_stream{input=[b'This is the story of Obama ridin...'] prefix_id= input_chars=[70] params=sampling { } stopping { max_new_tokens: 200 min_new_tokens: 16 } response { } decoding { } tokenization_time=0.45ms queue_and_inference_time=1096.67ms time_per_token=5.48ms total_time=1097.12ms input_toks=16}: Streaming response generated 200 tokens before NOT_FINISHED, output 848 chars: b' California. The story is told i...' INFO 04-09 21:51:08 logs.py:43] generate{input=[b'Lorem ipsum dolor sit amet, cons...', b'foooood man where is it'] prefix_id= input_chars=[469] params=sampling { } stopping { max_new_tokens: 20 min_new_tokens: 16 } response { } decoding { } tokenization_time=2.03ms queue_and_inference_time=122.23ms time_per_token=6.11ms total_time=124.26ms input_toks=124}: Sub-request 0 from batch of 2 generated 20 tokens before MAX_TOKENS, output 25 chars: b'?\\n\\n<!--\\n<!--\\n<!--\\n<!--\\n<!' INFO 04-09 21:51:08 logs.py:43] generate{input=[b'Lorem ipsum dolor sit amet, cons...', b'foooood man where is it'] prefix_id= input_chars=[469] params=sampling { } stopping { max_new_tokens: 20 min_new_tokens: 16 } response { } decoding { } tokenization_time=2.07ms queue_and_inference_time=122.22ms time_per_token=6.11ms total_time=124.29ms input_toks=7}: Sub-request 1 from batch of 2 generated 20 tokens before MAX_TOKENS, output 70 chars: b"?\\nI don't know.\\nI don't know.\\nI ..." ``` --------- Signed-off-by: Joe Runde <[email protected]> Signed-off-by: Joe Runde <[email protected]>

Correctly calculating the same value for the required cache blocks num for all torchrun processes

…-wenxh/fp8-on-a100-v5-pr Revert "0612 kernel of FP8 on A100"

zhuohan123 added 2 commits March 30, 2023 16:49

Merge QKV for OPT

28df307

merge qkv for llama

2e417f5

zhuohan123 requested a review from WoosukKwon March 30, 2023 17:04

WoosukKwon reviewed Mar 30, 2023

View reviewed changes

cacheflow/models/opt.py Outdated Show resolved Hide resolved

cacheflow/models/llama.py Outdated Show resolved Hide resolved

cacheflow/models/llama.py Outdated Show resolved Hide resolved

zhuohan123 added 2 commits March 31, 2023 07:24

fix the code according to woosuk's comment

06f23ff

Merge branch 'main' into qkv_combined

da0fdd2

WoosukKwon mentioned this pull request Apr 2, 2023

Optimize data movement #20

Merged

WoosukKwon self-requested a review April 2, 2023 07:23

WoosukKwon approved these changes Apr 2, 2023

View reviewed changes

WoosukKwon merged commit 1f01a18 into main Apr 2, 2023

zhuohan123 deleted the qkv_combined branch June 18, 2023 07:22

bigPYJ1151 added a commit to bigPYJ1151/vllm that referenced this pull request Sep 12, 2023

:Add page attention CPU support. (vllm-project#15)

64fafb6

shanshanpt mentioned this pull request Nov 17, 2023

Run long conetxt error : CUDA error: an illegal memory access was encountered #1700

Closed

junior-zsy mentioned this pull request Nov 20, 2023

Error with 32k Long Text in chatglm2-6b-32k Model #1725

Closed

hongxiayang pushed a commit to hongxiayang/vllm that referenced this pull request Feb 13, 2024

Merge QKV into one linear layer (vllm-project#15)

c8f2711

slyalin pushed a commit to slyalin/vllm that referenced this pull request Mar 26, 2024

Merge pull request vllm-project#15 from luo-cheng2021/luocheng/openvi…

f3a397f

…no-model-executor-opt [CPU] Avoid copy result and force allocation

fxmarty pushed a commit to fxmarty/vllm-public that referenced this pull request May 31, 2024

Merge pull request vllm-project#15 from ROCm/torchrun_cache_init_fix

4b39609

Correctly calculating the same value for the required cache blocks num for all torchrun processes

ykim362 pushed a commit to ykim362/vllm that referenced this pull request Jun 17, 2024

Merge pull request vllm-project#15 from xiaoxiawu-microsoft/revert-14…

cbcb64f

…-wenxh/fp8-on-a100-v5-pr Revert "0612 kernel of FP8 on A100"

alixiaodi mentioned this pull request Aug 2, 2024

[Bug]: #7072

Closed

SpaceHunterInf mentioned this pull request Sep 30, 2024

[Bug]: Bus error (core dumped) #8974

Closed

1 task

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Merge QKV into one linear layer #15

Merge QKV into one linear layer #15

zhuohan123 commented Mar 30, 2023 •

edited

Loading

WoosukKwon left a comment

WoosukKwon commented Apr 2, 2023

Merge QKV into one linear layer #15

Merge QKV into one linear layer #15

Conversation

zhuohan123 commented Mar 30, 2023 • edited Loading

WoosukKwon left a comment

Choose a reason for hiding this comment

WoosukKwon commented Apr 2, 2023

zhuohan123 commented Mar 30, 2023 •

edited

Loading