open-compass/MixtralKit
A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
Language: Python
#llm #mistral #moe
Stars: 251 Issues: 6 Forks: 21
https://github.com/open-compass/MixtralKit
A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
Language: Python
#llm #mistral #moe
Stars: 251 Issues: 6 Forks: 21
https://github.com/open-compass/MixtralKit
GitHub
GitHub - open-compass/MixtralKit: A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI - open-compass/MixtralKit
MoonshotAI/MoBA
MoBA: Mixture of Block Attention for Long-Context LLMs
Language: Python
#flash_attention #llm #llm_serving #llm_training #moe #pytorch #transformer
Stars: 521 Issues: 2 Forks: 16
https://github.com/MoonshotAI/MoBA
MoBA: Mixture of Block Attention for Long-Context LLMs
Language: Python
#flash_attention #llm #llm_serving #llm_training #moe #pytorch #transformer
Stars: 521 Issues: 2 Forks: 16
https://github.com/MoonshotAI/MoBA
GitHub
GitHub - MoonshotAI/MoBA: MoBA: Mixture of Block Attention for Long-Context LLMs
MoBA: Mixture of Block Attention for Long-Context LLMs - MoonshotAI/MoBA
FareedKhan-dev/kimi-k3-in-c
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Language: C
#avx2 #c99 #cpu_inference #deep_learning #from_scratch #inference_engine #kimi_k3 #linear_attention #llm #llm_inference #machine_learning #memory_efficient #mixture_of_experts #moe #mxfp4 #quantization #simd #systems_programming #transformer #zero_dependencies
Stars: 1398 Issues: 5 Forks: 220
https://github.com/FareedKhan-dev/kimi-k3-in-c
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Language: C
#avx2 #c99 #cpu_inference #deep_learning #from_scratch #inference_engine #kimi_k3 #linear_attention #llm #llm_inference #machine_learning #memory_efficient #mixture_of_experts #moe #mxfp4 #quantization #simd #systems_programming #transformer #zero_dependencies
Stars: 1398 Issues: 5 Forks: 220
https://github.com/FareedKhan-dev/kimi-k3-in-c
GitHub
GitHub - FareedKhan-dev/kimi-k3-in-c: A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable…
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU. - FareedKhan-dev/kimi-k3-in-c