jmaczan/tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Language:C++
Total stars: 525
Stars trend:
#cplusplus
#ai, #attention, #batching, #course, #cpp, #cuda, #hpc, #inference, #llm, #llminference, #pagedattention, #tinyvllm, #vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Language:C++
Total stars: 525
Stars trend:
30 May 2026
5am ▉ +7
6am ▌ +4
7am ▌ +4
8am ▌ +4
9am ▎ +2
10am ▎ +2
11am ▏ +1
12pm ▍ +3#cplusplus
#ai, #attention, #batching, #course, #cpp, #cuda, #hpc, #inference, #llm, #llminference, #pagedattention, #tinyvllm, #vllm
❤1