#python #deep_learning #inference #openai #quantization #speech_recognition #speech_to_text #transformer #whisper
Faster-Whisper is a fast version of OpenAI's Whisper that transcribes audio up to 4x quicker with the same accuracy, using less memory on CPU or GPU—benchmarks show it beats original Whisper (e.g., 1m03s vs 2m23s for 13-min audio on GPU). Install via `pip install faster-whisper`, no FFmpeg needed, and use simple Python code like `WhisperModel("large-v3").transcribe("audio.mp3")` for segments with timestamps. You benefit by getting quick, efficient speech-to-text for real-time apps, saving time and resources on long files or batches.
https://github.com/SYSTRAN/faster-whisper
Faster-Whisper is a fast version of OpenAI's Whisper that transcribes audio up to 4x quicker with the same accuracy, using less memory on CPU or GPU—benchmarks show it beats original Whisper (e.g., 1m03s vs 2m23s for 13-min audio on GPU). Install via `pip install faster-whisper`, no FFmpeg needed, and use simple Python code like `WhisperModel("large-v3").transcribe("audio.mp3")` for segments with timestamps. You benefit by getting quick, efficient speech-to-text for real-time apps, saving time and resources on long files or batches.
https://github.com/SYSTRAN/faster-whisper
GitHub
GitHub - SYSTRAN/faster-whisper: Faster Whisper transcription with CTranslate2
Faster Whisper transcription with CTranslate2. Contribute to SYSTRAN/faster-whisper development by creating an account on GitHub.
❤1
#python #deepseek #demo #easy #embedding #flask #gpt #huggingface_transformers #llm #mcp #multimodal #openai #qwen #rag #sentence_transformers #ui #vllm #vlm
UltraRAG is a lightweight framework that makes building retrieval-augmented generation (RAG) systems simple and fast. It uses a low-code approach where you write just dozens of lines of YAML configuration instead of complex code to create sophisticated AI workflows with conditional logic and loops. The framework includes a visual development environment where you can drag-and-drop to build pipelines, adjust parameters in real-time, and instantly convert your logic into interactive chat applications. This means you can deploy powerful AI systems that ground answers in your own data—reducing hallucinations and improving accuracy—without needing extensive coding expertise or lengthy development cycles.
https://github.com/OpenBMB/UltraRAG
UltraRAG is a lightweight framework that makes building retrieval-augmented generation (RAG) systems simple and fast. It uses a low-code approach where you write just dozens of lines of YAML configuration instead of complex code to create sophisticated AI workflows with conditional logic and loops. The framework includes a visual development environment where you can drag-and-drop to build pipelines, adjust parameters in real-time, and instantly convert your logic into interactive chat applications. This means you can deploy powerful AI systems that ground answers in your own data—reducing hallucinations and improving accuracy—without needing extensive coding expertise or lengthy development cycles.
https://github.com/OpenBMB/UltraRAG
GitHub
GitHub - OpenBMB/UltraRAG: A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines - OpenBMB/UltraRAG
#python #agents #ai #ai_engineer #ai_engineering #copilot #data_science #data_scientist #generative_ai #gpt #machine_learning #ml_engineer #ml_engineering #openai
AI Data Science Team is a free Python library with AI agents that speed up your data work 10X by handling loading, cleaning, visualization, EDA, feature engineering, modeling, and SQL tasks. Its flagship AI Pipeline Studio app creates visual, reproducible pipelines you can run with Streamlit after easy install (Python 3.10+, OpenAI or Ollama). This saves you hours on repetitive jobs, boosts accuracy, and lets you focus on insights and business results.
https://github.com/business-science/ai-data-science-team
AI Data Science Team is a free Python library with AI agents that speed up your data work 10X by handling loading, cleaning, visualization, EDA, feature engineering, modeling, and SQL tasks. Its flagship AI Pipeline Studio app creates visual, reproducible pipelines you can run with Streamlit after easy install (Python 3.10+, OpenAI or Ollama). This saves you hours on repetitive jobs, boosts accuracy, and lets you focus on insights and business results.
https://github.com/business-science/ai-data-science-team
GitHub
GitHub - business-science/ai-data-science-team: An AI-powered data science team of agents to help you perform common data science…
An AI-powered data science team of agents to help you perform common data science tasks 10X faster. - business-science/ai-data-science-team
#python #ai #claude #gemini #llama #llm #openai
You can access powerful AI language models for free or with trial credits through multiple legitimate platforms. Services like OpenRouter, Google AI Studio, Groq, and Mistral offer free tiers with varying request limits, while others like Fireworks, Baseten, and Inference.net provide trial credits ranging from $1 to $30. These platforms support diverse models including Llama, Gemma, Qwen, and DeepSeek, enabling you to build and test AI applications without upfront costs. The benefit is clear: you can prototype, develop, and deploy AI-powered features while managing your budget effectively, with options to scale up as your needs grow.
https://github.com/cheahjs/free-llm-api-resources
You can access powerful AI language models for free or with trial credits through multiple legitimate platforms. Services like OpenRouter, Google AI Studio, Groq, and Mistral offer free tiers with varying request limits, while others like Fireworks, Baseten, and Inference.net provide trial credits ranging from $1 to $30. These platforms support diverse models including Llama, Gemma, Qwen, and DeepSeek, enabling you to build and test AI applications without upfront costs. The benefit is clear: you can prototype, develop, and deploy AI-powered features while managing your budget effectively, with options to scale up as your needs grow.
https://github.com/cheahjs/free-llm-api-resources
GitHub
GitHub - cheahjs/free-llm-api-resources: A list of free LLM inference resources accessible via API.
A list of free LLM inference resources accessible via API. - cheahjs/free-llm-api-resources
#go #ai_agents #ai_security_tool #anthropic #autonomous_agents #golang #gpt #graphql #multi_agent_system #offensive_security #open_source #openai #penetration_testing #penetration_testing_tools #react #security_automation #security_testing #security_tools #self_hosted
PentAGI is an AI-powered tool that automates penetration testing with smart agents using 20+ pro tools like nmap and metasploit in a safe Docker sandbox. It researches vulnerabilities, executes attacks, stores knowledge for reuse, and creates detailed reports via a simple web UI. Quick setup needs Docker, an LLM API key (OpenAI/Anthropic), and `docker compose up -d`. This saves you hours of manual work, speeds up secure testing, cuts errors, and helps find issues faster for better protection.
https://github.com/vxcontrol/pentagi
PentAGI is an AI-powered tool that automates penetration testing with smart agents using 20+ pro tools like nmap and metasploit in a safe Docker sandbox. It researches vulnerabilities, executes attacks, stores knowledge for reuse, and creates detailed reports via a simple web UI. Quick setup needs Docker, an LLM API key (OpenAI/Anthropic), and `docker compose up -d`. This saves you hours of manual work, speeds up secure testing, cuts errors, and helps find issues faster for better protection.
https://github.com/vxcontrol/pentagi
GitHub
GitHub - vxcontrol/pentagi: Fully autonomous AI Agents system capable of performing complex penetration testing tasks
Fully autonomous AI Agents system capable of performing complex penetration testing tasks - vxcontrol/pentagi
👍2❤1
#rust #ai_gateway #ai_gateway_support #envoy #envoyproxy #gateway #generative_ai #llm_gateway #llm_inference #llm_proxy #llm_routing #llmops #llms #openai #prompt #proxy #proxy_server #routing
Plano is an AI-native proxy server that handles key tasks for agentic apps like routing between agents, smart LLM model selection, safety guardrails, and automatic traces for observability. Define agents in simple YAML, write basic HTTP code in any language, and start Plano to run multi-agent systems without custom plumbing or framework lock-in. You benefit by building and shipping reliable agents to production much faster, focusing on core logic while gaining safety, low latency, and easy scaling.
https://github.com/katanemo/plano
Plano is an AI-native proxy server that handles key tasks for agentic apps like routing between agents, smart LLM model selection, safety guardrails, and automatic traces for observability. Define agents in simple YAML, write basic HTTP code in any language, and start Plano to run multi-agent systems without custom plumbing or framework lock-in. You benefit by building and shipping reliable agents to production much faster, focusing on core logic while gaining safety, low latency, and easy scaling.
https://github.com/katanemo/plano
GitHub
GitHub - katanemo/plano: Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability,…
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic. - katanemo/p...
❤2
#javascript #ai #algorithm #artificial_intelligence #chatgpt #claude #cursor #deep_learning #deepseek #gemini #generative_ai #gpt #llm #mcp #openai #python #rag #vibe_coding #vibecoding #vue #vuepress
鱼皮的 AI知识库 offers a free Vibe Coding tutorial for beginners, teaching AI-powered programming with natural language prompts to build and monetize apps fast—no coding skills needed. It covers tools, projects, tips, and paths like making your first work in 10 minutes, plus AI guides on DeepSeek, Cursor, and more. You benefit by quickly creating profitable products, breaking tech barriers, and enjoying AI perks to improve life and work. Start at ai.codefather.cn/vibe.
https://github.com/liyupi/ai-guide
鱼皮的 AI知识库 offers a free Vibe Coding tutorial for beginners, teaching AI-powered programming with natural language prompts to build and monetize apps fast—no coding skills needed. It covers tools, projects, tips, and paths like making your first work in 10 minutes, plus AI guides on DeepSeek, Cursor, and more. You benefit by quickly creating profitable products, breaking tech barriers, and enjoying AI perks to improve life and work. Start at ai.codefather.cn/vibe.
https://github.com/liyupi/ai-guide
鱼皮AI导航
🌟 AI 编程零基础入门教程 Vibe Coding - 鱼皮的 AI 知识库(免费) - 鱼皮AI导航
大家好,我是程序员鱼皮。
如今 Vibe Coding 已经火遍全网,不仅是程序员,连设计师、产品运营、甚至完全不懂技术的人都开始用 Vibe Coding 实现自己的想法,用 AI 做出了自己的产。鱼皮AI导航收录全球AI工具网站应用,专业学习资源资讯知识库,AI学习与交流社区。
如今 Vibe Coding 已经火遍全网,不仅是程序员,连设计师、产品运营、甚至完全不懂技术的人都开始用 Vibe Coding 实现自己的想法,用 AI 做出了自己的产。鱼皮AI导航收录全球AI工具网站应用,专业学习资源资讯知识库,AI学习与交流社区。
❤4
#python #agentic_ai #agentic_coding #ai_coding_agent #ai_plugins #anthropic_claude #claude_ai #claude_ai_skills #claude_code #claude_code_plugins #claude_code_skills #claude_skills #claudecode_subagents #developer_tools #devtools #mcp_tools #openai_codex #prompt_engineering
Claude Code Skills offers 169 free, ready-to-use plugins that turn AI coding agents like Claude Code, OpenAI Codex, and OpenClaw into experts in engineering, marketing, product, compliance, and more. Install easily via simple commands to add skills like security auditing, test automation, or C-level advice, with 160+ Python tools included. This saves you time by automating complex tasks, boosting code quality, and handling grunt work so you focus on creative problem-solving and faster results.
https://github.com/alirezarezvani/claude-skills
Claude Code Skills offers 169 free, ready-to-use plugins that turn AI coding agents like Claude Code, OpenAI Codex, and OpenClaw into experts in engineering, marketing, product, compliance, and more. Install easily via simple commands to add skills like security auditing, test automation, or C-level advice, with 160+ Python tools included. This saves you time by automating complex tasks, boosting code quality, and handling grunt work so you focus on creative problem-solving and faster results.
https://github.com/alirezarezvani/claude-skills
GitHub
GitHub - alirezarezvani/claude-skills: 345 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 330+ skills…
345 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 330+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8 mor...
#python #agent #agents #ai #anthropic #claudecode #llm #llms #openai
Open SWE is a free, open-source framework to build internal coding agents like those at Stripe, Ramp, and Coinbase. Trigger it via Slack, Linear, or GitHub (@openswe) to research codebases, plan tasks, code, test, review, and auto-open PRs in secure cloud sandboxes—running parallel jobs without your machine's resources. Customize models, tools, and workflows easily. You benefit by automating routine coding, slashing review cycles and production time by 30-50%, freeing you to focus on high-value work while ensuring safe, high-quality changes.
https://github.com/langchain-ai/open-swe
Open SWE is a free, open-source framework to build internal coding agents like those at Stripe, Ramp, and Coinbase. Trigger it via Slack, Linear, or GitHub (@openswe) to research codebases, plan tasks, code, test, review, and auto-open PRs in secure cloud sandboxes—running parallel jobs without your machine's resources. Customize models, tools, and workflows easily. You benefit by automating routine coding, slashing review cycles and production time by 30-50%, freeing you to focus on high-value work while ensuring safe, high-quality changes.
https://github.com/langchain-ai/open-swe
GitHub
GitHub - langchain-ai/open-swe: An Open-Source Asynchronous Coding Agent
An Open-Source Asynchronous Coding Agent. Contribute to langchain-ai/open-swe development by creating an account on GitHub.
❤4
#typescript #agent #agentic_rag #ai_coding #claude_code #code_generation #code_search #cursor #embedding #gemini_cli #mcp #merkle_tree #nodejs #openai #rag #semantic_search #typescript #vector_database #vibe_coding #voyage_ai #vscode_extension
Claude Context is a plugin that adds semantic code search to Claude Code and other AI tools, using your full codebase as context via a vector database like Zilliz Cloud. It finds relevant code instantly with natural language queries, indexes efficiently (only changed files), and cuts token use by ~40% for the same quality. You save costs on large projects, get precise results without loading whole files, and code faster with deep, relevant context across millions of lines. Setup needs free Zilliz/OpenAI keys and Node.js 20+; works with VS Code, Cursor, and more.
https://github.com/zilliztech/claude-context
Claude Context is a plugin that adds semantic code search to Claude Code and other AI tools, using your full codebase as context via a vector database like Zilliz Cloud. It finds relevant code instantly with natural language queries, indexes efficiently (only changed files), and cuts token use by ~40% for the same quality. You save costs on large projects, get precise results without loading whole files, and code faster with deep, relevant context across millions of lines. Setup needs free Zilliz/OpenAI keys and Node.js 20+; works with VS Code, Cursor, and more.
https://github.com/zilliztech/claude-context
GitHub
GitHub - zilliztech/claude-context: Code search MCP for Claude Code. Make entire codebase the context for any coding agent.
Code search MCP for Claude Code. Make entire codebase the context for any coding agent. - zilliztech/claude-context
#typescript #analytics #autogen #evaluation #langchain #large_language_models #llama_index #llm #llm_evaluation #llm_observability #llmops #monitoring #observability #open_source #openai #playground #prompt_engineering #prompt_management #self_hosted #ycombinator
Langfuse is a free, open-source platform to build, monitor, evaluate, and debug AI apps using large language models (LLMs). It offers tracing for app logic, prompt management, evaluations, datasets, a playground, and easy integrations like OpenAI, LangChain, and LlamaIndex. Deploy it on Langfuse Cloud (free tier) or self-host with Docker in minutes. This helps you quickly spot issues, improve prompts without slowing apps, test reliably, and speed up development—saving time and boosting AI performance.
https://github.com/langfuse/langfuse
Langfuse is a free, open-source platform to build, monitor, evaluate, and debug AI apps using large language models (LLMs). It offers tracing for app logic, prompt management, evaluations, datasets, a playground, and easy integrations like OpenAI, LangChain, and LlamaIndex. Deploy it on Langfuse Cloud (free tier) or self-host with Docker in minutes. This helps you quickly spot issues, improve prompts without slowing apps, test reliably, and speed up development—saving time and boosting AI performance.
https://github.com/langfuse/langfuse
GitHub
GitHub - langfuse/langfuse: 🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground…
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23 ...
❤1
#go #api #claude_api #deepseek #deepseek_api #docker #freeapi #go #openai_api #proxy #proxy_server #react #vercel #vercel_deployment #zeabur
DS2API turns DeepSeek web chat into APIs compatible with OpenAI, Claude, and Gemini, using Go backend and React web UI for easy management. It supports multi-account rotation, concurrency queues, tool calling, and models like deepseek-chat/reasoner with aliases (e.g., gpt-5). Deploy simply via release binaries, Docker, Vercel, or source—edit config.json with your keys/accounts and run. You benefit by accessing DeepSeek affordably through familiar SDKs, saving costs and simplifying integration for apps or testing.
https://github.com/CJackHwang/ds2api
DS2API turns DeepSeek web chat into APIs compatible with OpenAI, Claude, and Gemini, using Go backend and React web UI for easy management. It supports multi-account rotation, concurrency queues, tool calling, and models like deepseek-chat/reasoner with aliases (e.g., gpt-5). Deploy simply via release binaries, Docker, Vercel, or source—edit config.json with your keys/accounts and run. You benefit by accessing DeepSeek affordably through familiar SDKs, saving costs and simplifying integration for apps or testing.
https://github.com/CJackHwang/ds2api
GitHub
GitHub - CJackHwang/ds2api: DeepSeek-Compatible Middleware Interface: A technical exploration project in Go, focusing on high-concurrency…
DeepSeek-Compatible Middleware Interface: A technical exploration project in Go, focusing on high-concurrency protocol adaptation. It serves as a reference implementation for converting diverse web...
❤2
#rust #ai #claude #cli #coding_agent #llm #mcp #openai #rust #terminal #tui
jcode is a fast, low-RAM coding agent for Linux, macOS, and Windows that boosts your skills with multi-session workflows, smart memory recall, swarm collaboration, side panels for diagrams/files, and logins for models like Claude or OpenAI. Install easily via `curl -fsSL https://raw.githubusercontent.com/1jehuang/jcode/master/scripts/install.sh | bash`, then run `jcode`. It uses far less memory (27MB vs. 300MB+ for rivals) and starts in 14ms, letting you handle many agents smoothly without slowdowns or high costs—perfect for efficient, scalable coding.
https://github.com/1jehuang/jcode
jcode is a fast, low-RAM coding agent for Linux, macOS, and Windows that boosts your skills with multi-session workflows, smart memory recall, swarm collaboration, side panels for diagrams/files, and logins for models like Claude or OpenAI. Install easily via `curl -fsSL https://raw.githubusercontent.com/1jehuang/jcode/master/scripts/install.sh | bash`, then run `jcode`. It uses far less memory (27MB vs. 300MB+ for rivals) and starts in 14ms, letting you handle many agents smoothly without slowdowns or high costs—perfect for efficient, scalable coding.
https://github.com/1jehuang/jcode
GitHub
GitHub - 1jehuang/jcode: The most RAM efficient harness
The most RAM efficient harness. Contribute to 1jehuang/jcode development by creating an account on GitHub.
#python #academia #anthropic #arxiv #brave #deep_research #encryption #home_automation #homeserver #local #local_deep_research #local_llm #mistral #ollama #openai #pubmed #research #research_tool #retrieval_augmented_generation #searxng #self_hosted
Local Deep Research is a free, open-source AI tool you run locally for private, deep research on any topic. It auto-searches the web, academic papers (arXiv, PubMed), and your documents using LLMs like Ollama or GPT, then synthesizes cited reports in minutes. Install easily via Docker or pip, build an encrypted knowledge base from downloads, and get 95% accuracy. Benefits: total privacy (no tracking), zero cost for local models, customizable strategies, and compounding knowledge—saving hours on complex queries while owning your data.
https://github.com/LearningCircuit/local-deep-research
Local Deep Research is a free, open-source AI tool you run locally for private, deep research on any topic. It auto-searches the web, academic papers (arXiv, PubMed), and your documents using LLMs like Ollama or GPT, then synthesizes cited reports in minutes. Install easily via Docker or pip, build an encrypted knowledge base from downloads, and get 95% accuracy. Benefits: total privacy (no tracking), zero cost for local models, customizable strategies, and compounding knowledge—saving hours on complex queries while owning your data.
https://github.com/LearningCircuit/local-deep-research
GitHub
GitHub - LearningCircuit/local-deep-research: ~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs…
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local &...
#javascript #ai_agents #ai_gateway #anthropic #chatgpt #claude #claude_code #cline #codex #copilot #cursor #deepseek #free_ai #gemini #gemini_cli #llm #llm_gateway #openai #openai_proxy #qwen #token_saver
# 9Router//localhost:20128/v1`. You get unlimited coding with zero cost, no downtime, and intelligent fallback routing—perfect for developers who want maximum value from AI without subscription limits or surprise bills.
https://github.com/decolua/9router
# 9Router//localhost:20128/v1`. You get unlimited coding with zero cost, no downtime, and intelligent fallback routing—perfect for developers who want maximum value from AI without subscription limits or surprise bills.
https://github.com/decolua/9router
GitHub
GitHub - decolua/9router: Unlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini…
Unlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40% tokens, never hit limits. - decolua/9r...
❤3
#javascript #agent #ai #coding #course #deepseek #gemini #genai #gpt #llm #low_code #mcp #nextjs #no_code #openai #programming #tutorial #vibe_coding #vibecoding #vscode #workflow
# Easy-Vibe: Learn to Build Apps by Speaking
Easy-Vibe is a learning platform that teaches you to create real applications using AI by simply describing what you want. It offers beginner-friendly guides, step-by-step visual tutorials, and interactive coding simulations that make learning feel like having a private tutor. The platform covers everything from your first project to advanced full-stack development, with animated explanations of AI principles and game-like learning for complex topics like data retrieval systems. You benefit by gaining practical skills to turn ideas into working products quickly, whether you're a complete beginner, student, or developer wanting to master AI-assisted coding in the modern era.
https://github.com/datawhalechina/easy-vibe
# Easy-Vibe: Learn to Build Apps by Speaking
Easy-Vibe is a learning platform that teaches you to create real applications using AI by simply describing what you want. It offers beginner-friendly guides, step-by-step visual tutorials, and interactive coding simulations that make learning feel like having a private tutor. The platform covers everything from your first project to advanced full-stack development, with animated explanations of AI principles and game-like learning for complex topics like data retrieval systems. You benefit by gaining practical skills to turn ideas into working products quickly, whether you're a complete beginner, student, or developer wanting to master AI-assisted coding in the modern era.
https://github.com/datawhalechina/easy-vibe
GitHub
GitHub - datawhalechina/easy-vibe: 💻 vibe coding 2026 | Your First Modern Coding Course Beginners to Master Step by Step.
💻 vibe coding 2026 | Your First Modern Coding Course Beginners to Master Step by Step. - datawhalechina/easy-vibe
👎1🎃1
#python #apple_silicon #inference_server #llm #macos #mlx #openai_api
# oMLX: Run AI Models Faster on Your Mac
oMLX is a tool that lets you run large language models directly on your Mac with Apple Silicon chips. It uses smart memory management—keeping frequently used models in RAM and storing less-used ones on your SSD—so everything runs smoothly without slowdowns. You control everything from a simple menu bar app or web dashboard. It works with many popular AI models and connects easily to coding tools like Claude Code. The benefit is you get powerful AI capabilities locally on your Mac without needing cloud services, saving money and keeping your data private while enjoying fast, responsive performance.
https://github.com/jundot/omlx
# oMLX: Run AI Models Faster on Your Mac
oMLX is a tool that lets you run large language models directly on your Mac with Apple Silicon chips. It uses smart memory management—keeping frequently used models in RAM and storing less-used ones on your SSD—so everything runs smoothly without slowdowns. You control everything from a simple menu bar app or web dashboard. It works with many popular AI models and connects easily to coding tools like Claude Code. The benefit is you get powerful AI capabilities locally on your Mac without needing cloud services, saving money and keeping your data private while enjoying fast, responsive performance.
https://github.com/jundot/omlx
GitHub
GitHub - jundot/omlx: LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu…
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar - jundot/omlx