dipampaul17/KVSplit
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
Language:Python
Total stars: 144
Stars trend:
#python
#applesilicon, #generativeai, #kvcache, #llamacpp, #llm, #m1, #m2, #m3, #memoryoptimization, #metal, #optimization, #quantization
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.
Language:Python
Total stars: 144
Stars trend:
16 May 2025
7pm ▏ +1
8pm █████▌ +44
9pm ████▊ +38
10pm ███▋ +29
11pm ██▎ +18#python
#applesilicon, #generativeai, #kvcache, #llamacpp, #llm, #m1, #m2, #m3, #memoryoptimization, #metal, #optimization, #quantization
MemTensor/MemOS
MemOS (Preview) | Intelligence Begins with Memory
Language:Python
Total stars: 162
Stars trend:
#python
#agent, #kvcache, #languagemodel, #llm, #lora, #memcube, #memory, #memos, #neo4j, #tree
MemOS (Preview) | Intelligence Begins with Memory
Language:Python
Total stars: 162
Stars trend:
6 Jul 2025
11pm ▏ +1
7 Jul 2025
12am +0
1am +0
2am ▉ +7
3am █▎ +10
4am ▊ +6
5am ████▍ +35
6am ███▍ +27#python
#agent, #kvcache, #languagemodel, #llm, #lora, #memcube, #memory, #memos, #neo4j, #tree
bojieli/ai-infra-book
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
Language:Python
Total stars: 2403
#python
#accelerator, #aiinfra, #aiinfrastructure, #book, #datacenternetwork, #deepseek, #distributedsystems, #gpu, #kvcache, #llm, #llminference, #llmserving, #llmtraining, #mixtureofexperts, #performanceengineering
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
Language:Python
Total stars: 2403
#python
#accelerator, #aiinfra, #aiinfrastructure, #book, #datacenternetwork, #deepseek, #distributedsystems, #gpu, #kvcache, #llm, #llminference, #llmserving, #llmtraining, #mixtureofexperts, #performanceengineering