#javascript #ai_art_generator #ai_image_generation #ai_video_generation #creative_tools #flux_1 #generative_ai #higgsfield #higgsfield_ai #higgsfield_alternative #image_to_video #javascript #kling_ai #midjourney_alternative #muapi #open_source #sora_alternative #text_to_video #uncensored #unrestricted #wan_video
Open Generative AI is a free, open-source tool for creating AI images, videos, lip syncs, and cinema with 200+ models like Flux and Kling. Unlike paid apps with filters and limits, it has no content blocks, subscriptions, or data sharing—you get full creative freedom, self-hosting, and local generation on your device. Try it online at dev.muapi.ai/open-generative-ai or download desktop apps for Mac, Windows, or Linux. This saves money, protects privacy, and lets you make unrestricted art fast.
https://github.com/Anil-matcha/Open-Generative-AI
Open Generative AI is a free, open-source tool for creating AI images, videos, lip syncs, and cinema with 200+ models like Flux and Kling. Unlike paid apps with filters and limits, it has no content blocks, subscriptions, or data sharing—you get full creative freedom, self-hosting, and local generation on your device. Try it online at dev.muapi.ai/open-generative-ai or download desktop apps for Mac, Windows, or Linux. This saves money, protects privacy, and lets you make unrestricted art fast.
https://github.com/Anil-matcha/Open-Generative-AI
Muapi
Muapi | AI Image & Video API Platform
Build advanced AI workflows with our high-performance Image and Video generation APIs.
❤3👍1
#swift #cpp #csharp #go #ios #java #lightweight #nodejs #on_device #python #rust #swift #text_to_speech #tts #web
Supertonic is a fast, lightweight text-to-speech system that runs directly on your device without needing the internet or cloud services. It supports 31 languages and works across phones, computers, browsers, and other platforms. The system is small enough to run on devices like Raspberry Pi while staying accurate and quick. You get complete privacy since everything happens locally on your device, and you can use it for free with no network dependency. It handles complex text like phone numbers and currency amounts better than many larger systems.
https://github.com/supertone-inc/supertonic
Supertonic is a fast, lightweight text-to-speech system that runs directly on your device without needing the internet or cloud services. It supports 31 languages and works across phones, computers, browsers, and other platforms. The system is small enough to run on devices like Raspberry Pi while staying accurate and quick. You get complete privacy since everything happens locally on your device, and you can use it for free with no network dependency. It handles complex text like phone numbers and currency amounts better than many larger systems.
https://github.com/supertone-inc/supertonic
GitHub
GitHub - supertone-inc/supertonic: Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX. - supertone-inc/supertonic
#python #ai #ai_agents #conversational_ai #fastapi #llm #nextjs #open_source #outbound_calls #pipecat #python #self_hosted #speech_to_text #telephony #text_to_speech #voice #voice_agents #voice_ai #voice_assistant #voip #webrtc
Dograh AI is an open-source, self-hostable tool for building voice agents with a drag-and-drop workflow. You can start fast, run it on your own server, use your own LLM, TTS, and STT services, and avoid vendor lock-in. The benefit to you is more control, more privacy, and a working voice bot in minutes without needing API keys.
https://github.com/dograh-hq/dograh
Dograh AI is an open-source, self-hostable tool for building voice agents with a drag-and-drop workflow. You can start fast, run it on your own server, use your own LLM, TTS, and STT services, and avoid vendor lock-in. The benefit to you is more control, more privacy, and a working voice bot in minutes without needing API keys.
https://github.com/dograh-hq/dograh
GitHub
GitHub - dograh-hq/dograh: Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech…
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support. -...
#python #ai_agents #amd #comfyui #docker #llama_cpp #llm #local_ai #n8n #nvidia #open_webui #rag #self_hosted #speech_to_text #strix_halo #text_to_speech #workflow_automation
Dream Server lets you run AI on your own machine instead of renting it from a cloud service. It works on Linux, Windows, and macOS, and it can set up chat, voice, agents, search, image tools, and privacy tools with one command. The main benefit is more control: your data stays with you, costs can be lower, and you can keep using AI even without a cloud account.
https://github.com/Light-Heart-Labs/DreamServer
Dream Server lets you run AI on your own machine instead of renting it from a cloud service. It works on Linux, Windows, and macOS, and it can set up chat, voice, agents, search, image tools, and privacy tools with one command. The main benefit is more control: your data stays with you, costs can be lower, and you can keep using AI even without a cloud account.
https://github.com/Light-Heart-Labs/DreamServer
GitHub
GitHub - Osmantic/ODS: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG…
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. - Osmantic/ODS
❤1
#python #audio #audio_tokenizer #llm #multimodal #text_to_speech #voice_cloning
MOSS-TTS is an open-source family of speech and sound models for natural, high-quality audio generation, including voice cloning, multi-speaker dialogue, real-time speech, and sound effects. It supports 31 languages in v1.5, better voice stability, and pause control, and it also offers a lightweight Nano version that can run on 4 CPU cores. The benefit to you is simple: you can create realistic speech or sound for apps, demos, or products with strong quality, flexible control, and multiple ways to run it.
https://github.com/OpenMOSS/MOSS-TTS
MOSS-TTS is an open-source family of speech and sound models for natural, high-quality audio generation, including voice cloning, multi-speaker dialogue, real-time speech, and sound effects. It supports 31 languages in v1.5, better voice stability, and pause control, and it also offers a lightweight Nano version that can run on 4 CPU cores. The benefit to you is simple: you can create realistic speech or sound for apps, demos, or products with strong quality, flexible control, and multiple ways to run it.
https://github.com/OpenMOSS/MOSS-TTS
GitHub
GitHub - OpenMOSS/MOSS-TTS: MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS…
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenario...
#rust #document_ocr #document_processing #ocr #ocr_recognition #pdf #pdf_parser #text_extraction
LiteParse is a fast, local PDF parser that extracts text with bounding boxes, can use OCR, and works in Rust, Python, Node.js, and the browser. It also makes screenshots and can handle files like DOCX, XLSX, PPTX, and images after conversion. Benefit: you can turn documents into clean text or JSON on your own machine, which helps with private, quick, and structured document processing.
https://github.com/run-llama/liteparse
LiteParse is a fast, local PDF parser that extracts text with bounding boxes, can use OCR, and works in Rust, Python, Node.js, and the browser. It also makes screenshots and can handle files like DOCX, XLSX, PPTX, and images after conversion. Benefit: you can turn documents into clean text or JSON on your own machine, which helps with private, quick, and structured document processing.
https://github.com/run-llama/liteparse
GitHub
GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser
A fast, helpful, and open-source document parser. Contribute to run-llama/liteparse development by creating an account on GitHub.
❤1
#python #agent #agentic_ai #ai #claude #copilot #cursor #elevenlabs #ffmpeg #flux #image_generation #open_source #openai #python #remotion #stable_diffusion #text_to_speech #text_to_video #video_generation #video_production
OpenMontage turns a plain idea or even a reference video into a full video production workflow, handling research, script writing, asset creation, editing, captions, and final rendering. Your benefit is faster video creation with lower cost, more control, and fewer surprises, because it can use free/open footage or AI tools, estimate cost first, and check quality before showing you the result.
https://github.com/calesthio/OpenMontage
OpenMontage turns a plain idea or even a reference video into a full video production workflow, handling research, script writing, asset creation, editing, captions, and final rendering. Your benefit is faster video creation with lower cost, more control, and fewer surprises, because it can use free/open footage or AI tools, estimate cost first, and check quality before showing you the result.
https://github.com/calesthio/OpenMontage
GitHub
GitHub - calesthio/OpenMontage: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools…
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full v...
❤1
#javascript #3mf #agents #ai_agents #build123d #cad #dxf #glb #mechanical_engineering #opencascade #robotics #sdf #srdf #step #stl #stp #text_to_cad #urdf
CAD Skills is a free, open-source library that lets AI agents like Codex and Claude Code create, edit, and verify CAD models, robot files, and 3D print code from simple text or images . You benefit by automating complex design tasks—such as generating STEP parts, writing robot descriptions (URDF/SRDF), slicing G-code, and previewing files locally—without manual modeling, saving time and reducing errors in hardware and robotics projects .
https://github.com/earthtojake/text-to-cad
CAD Skills is a free, open-source library that lets AI agents like Codex and Claude Code create, edit, and verify CAD models, robot files, and 3D print code from simple text or images . You benefit by automating complex design tasks—such as generating STEP parts, writing robot descriptions (URDF/SRDF), slicing G-code, and previewing files locally—without manual modeling, saving time and reducing errors in hardware and robotics projects .
https://github.com/earthtojake/text-to-cad
GitHub
GitHub - earthtojake/text-to-cad: A library of agent skills for CAD, CAE and CAM
A library of agent skills for CAD, CAE and CAM. Contribute to earthtojake/text-to-cad development by creating an account on GitHub.
#python #audiobook #faster_whisper #gradio #karaoke #podcasts #speech_recognition #speech_synthesis #speech_to_text #subtitles #text_to_speech #transcription #translator #tts #voice_cloning #voice_conversion #webui #whisper #whisperx #yt_dlp
Voice-Pro is a free, open-source Windows app that lets you download YouTube videos, separate voices, turn speech into text, translate into 100+ languages, and create new speech or cloned voices. It gives you one tool for subtitles, dubbing, and voice work, so you can save time and make multilingual content more easily.
https://github.com/abus-aikorea/voice-pro
Voice-Pro is a free, open-source Windows app that lets you download YouTube videos, separate voices, turn speech into text, translate into 100+ languages, and create new speech or cloned voices. It gives you one tool for subtitles, dubbing, and voice work, so you can save time and make multilingual content more easily.
https://github.com/abus-aikorea/voice-pro
GitHub
GitHub - abus-aikorea/voice-pro: Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice…
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs ...
👍2
#rust #markdown #nodejs #ocr_routing #pdf #pdf_classification #pdf_extraction #pdf_parser #python #rust #text_extraction
pdf-inspector is a fast Rust tool that checks whether a PDF has real text or is scanned, then extracts text and turns it into clean Markdown without OCR for text PDFs. It works in Python, Node.js, Rust, and browsers, and it is useful because it can save time and cost by skipping OCR, while still giving you searchable, copyable, well-structured output for reports, papers, invoices, and legal files.
https://github.com/firecrawl/pdf-inspector
pdf-inspector is a fast Rust tool that checks whether a PDF has real text or is scanned, then extracts text and turns it into clean Markdown without OCR for text PDFs. It works in Python, Node.js, Rust, and browsers, and it is useful because it can save time and cost by skipping OCR, while still giving you searchable, copyable, well-structured output for reports, papers, invoices, and legal files.
https://github.com/firecrawl/pdf-inspector
GitHub
GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects…
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions. - firecrawl/pdf-inspector
❤1