lucidrains/soundstorm-pytorch
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch
Language: Python
#artificial_intelligence #attention_mechanism #audio_generation #deep_learning #non_autoregressive #transformers
Stars: 181 Issues: 0 Forks: 6
https://github.com/lucidrains/soundstorm-pytorch
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch
Language: Python
#artificial_intelligence #attention_mechanism #audio_generation #deep_learning #non_autoregressive #transformers
Stars: 181 Issues: 0 Forks: 6
https://github.com/lucidrains/soundstorm-pytorch
GitHub
GitHub - lucidrains/soundstorm-pytorch: Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind…
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch - lucidrains/soundstorm-pytorch
😁1
OFA-Sys/ONE-PEACE
A general representation modal across vision, audio, language modalities.
Language: Python
#audio_language #foundation_models #multimodal #representation_learning #vision_language
Stars: 185 Issues: 2 Forks: 5
https://github.com/OFA-Sys/ONE-PEACE
A general representation modal across vision, audio, language modalities.
Language: Python
#audio_language #foundation_models #multimodal #representation_learning #vision_language
Stars: 185 Issues: 2 Forks: 5
https://github.com/OFA-Sys/ONE-PEACE
GitHub
GitHub - OFA-Sys/ONE-PEACE: A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring…
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities - OFA-Sys/ONE-PEACE
VASTDynamics/Vaporizer2
Vaporizer2 hybrid wavetable additive / subtractive VST / AU / AAX synthesizer / sampler workstation plugin
Language: C++
#aax #audio #audiounit_plugins #cpp #daw #music #plugin #sampler #synthesizer #vst #vst3 #vst3_plugin #wavetable
Stars: 186 Issues: 5 Forks: 9
https://github.com/VASTDynamics/Vaporizer2
Vaporizer2 hybrid wavetable additive / subtractive VST / AU / AAX synthesizer / sampler workstation plugin
Language: C++
#aax #audio #audiounit_plugins #cpp #daw #music #plugin #sampler #synthesizer #vst #vst3 #vst3_plugin #wavetable
Stars: 186 Issues: 5 Forks: 9
https://github.com/VASTDynamics/Vaporizer2
GitHub
GitHub - VASTDynamics/Vaporizer2: Vaporizer2 hybrid wavetable additive / subtractive VST / AU / AAX synthesizer / sampler workstation…
Vaporizer2 hybrid wavetable additive / subtractive VST / AU / AAX synthesizer / sampler workstation plugin - VASTDynamics/Vaporizer2
🔥1
huggingface/distil-whisper
#audio #speech_recognition #whisper
Stars: 261 Issues: 2 Forks: 9
https://github.com/huggingface/distil-whisper
#audio #speech_recognition #whisper
Stars: 261 Issues: 2 Forks: 9
https://github.com/huggingface/distil-whisper
GitHub
GitHub - huggingface/distil-whisper: Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word…
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate. - huggingface/distil-whisper
ZiqiaoPeng/SyncTalk
This is the official source for our paper "SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis"
#audio_driven_talking_face #talking_face #talking_face_generation #talking_head
Stars: 180 Issues: 5 Forks: 2
https://github.com/ZiqiaoPeng/SyncTalk
This is the official source for our paper "SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis"
#audio_driven_talking_face #talking_face #talking_face_generation #talking_head
Stars: 180 Issues: 5 Forks: 2
https://github.com/ZiqiaoPeng/SyncTalk
GitHub
GitHub - ZiqiaoPeng/SyncTalk: [CVPR 2024] This is the official source for our paper "SyncTalk: The Devil is in the Synchronization…
[CVPR 2024] This is the official source for our paper "SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis" - ZiqiaoPeng/SyncTalk
TuneNN/TuneNN
A transformer-based network model for pitch detection
Language: Python
#audio #machine_learning #music #pitch_detection #pitch_estimation
Stars: 142 Issues: 0 Forks: 3
https://github.com/TuneNN/TuneNN
A transformer-based network model for pitch detection
Language: Python
#audio #machine_learning #music #pitch_detection #pitch_estimation
Stars: 142 Issues: 0 Forks: 3
https://github.com/TuneNN/TuneNN
GitHub
GitHub - TuneNN/TuneNN: A transformer-based network model for pitch detection
A transformer-based network model for pitch detection - TuneNN/TuneNN
👍1
ali-vilab/dreamtalk
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
Language: Python
#audio_visual_learning #face_animation #talking_head #video_generation
Stars: 217 Issues: 7 Forks: 20
https://github.com/ali-vilab/dreamtalk
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models
Language: Python
#audio_visual_learning #face_animation #talking_head #video_generation
Stars: 217 Issues: 7 Forks: 20
https://github.com/ali-vilab/dreamtalk
GitHub
GitHub - ali-vilab/dreamtalk: Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion…
Official implementations for paper: DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic Models - ali-vilab/dreamtalk
Lessica/TrollRecorder
WIP: A simple audio recorder for TrollStore.
Language: Objective-C++
#audio_recorder #ios #jailbreak #trollstore #tweak
Stars: 282 Issues: 1 Forks: 10
https://github.com/Lessica/TrollRecorder
WIP: A simple audio recorder for TrollStore.
Language: Objective-C++
#audio_recorder #ios #jailbreak #trollstore #tweak
Stars: 282 Issues: 1 Forks: 10
https://github.com/Lessica/TrollRecorder
GitHub
GitHub - Lessica/TrollRecorder: (i18n/CLI) Not the first, but the best phone call recorder with TrollStore.
(i18n/CLI) Not the first, but the best phone call recorder with TrollStore. - Lessica/TrollRecorder
👍5
jishengpeng/WavTokenizer
SOTA discrete acoustic codec models with 40 tokens per second for audio language modeling
Language: Python
#acoustic #audio_representation #codec #dac #encodec #gpt4o #music_representation_learning #semantic #soundstream #speech_language_model #speech_representation #text_to_speech
Stars: 332 Issues: 6 Forks: 20
https://github.com/jishengpeng/WavTokenizer
SOTA discrete acoustic codec models with 40 tokens per second for audio language modeling
Language: Python
#acoustic #audio_representation #codec #dac #encodec #gpt4o #music_representation_learning #semantic #soundstream #speech_language_model #speech_representation #text_to_speech
Stars: 332 Issues: 6 Forks: 20
https://github.com/jishengpeng/WavTokenizer
GitHub
GitHub - jishengpeng/WavTokenizer: [ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language…
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling - GitHub - jishengpeng/WavTokenizer: [ICLR 2025] SOTA discrete acoustic codec models with 4...
antgroup/echomimic_v2
EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
Language: Python
#audio_driven_portrait_animations #audio_driven_talking_face #human_animation #talking_face_generation #talking_head
Stars: 307 Issues: 5 Forks: 28
https://github.com/antgroup/echomimic_v2
EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
Language: Python
#audio_driven_portrait_animations #audio_driven_talking_face #human_animation #talking_face_generation #talking_head
Stars: 307 Issues: 5 Forks: 28
https://github.com/antgroup/echomimic_v2
GitHub
GitHub - antgroup/echomimic_v2: [CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
[CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation - antgroup/echomimic_v2
Tencent/HunyuanCustom
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Language: Python
#audio_driven #diffusion_models #image_to_video #image_to_video_generation #video_editing #video_generation
Stars: 360 Issues: 4 Forks: 14
https://github.com/Tencent/HunyuanCustom
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
Language: Python
#audio_driven #diffusion_models #image_to_video #image_to_video_generation #video_editing #video_generation
Stars: 360 Issues: 4 Forks: 14
https://github.com/Tencent/HunyuanCustom
GitHub
GitHub - Tencent-Hunyuan/HunyuanCustom: HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation - Tencent-Hunyuan/HunyuanCustom
❤1
wildminder/ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
Language: Python
#ai_voice #audio #comfyui_node #t2s #text_to_speech #tts #voice_cloning #voice_generation
Stars: 198 Issues: 2 Forks: 21
https://github.com/wildminder/ComfyUI-VoxCPM
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
Language: Python
#ai_voice #audio #comfyui_node #t2s #text_to_speech #tts #voice_cloning #voice_generation
Stars: 198 Issues: 2 Forks: 21
https://github.com/wildminder/ComfyUI-VoxCPM
GitHub
GitHub - wildminder/ComfyUI-VoxCPM: ComfyUI node for highly expressive speech and realistic zero-shot voice cloning
ComfyUI node for highly expressive speech and realistic zero-shot voice cloning - wildminder/ComfyUI-VoxCPM
❤2
FunAudioLLM/Fun-ASR
Fun-ASR is an end-to-end speech recognition large model launched by Tongyi Lab.
Language: Python
#audio #audio_language_model #audio_understanding #fun_asr #multimodal_large_language_models #pytorch #speaker_diarization #speech_recognition
Stars: 264 Issues: 4 Forks: 8
https://github.com/FunAudioLLM/Fun-ASR
Fun-ASR is an end-to-end speech recognition large model launched by Tongyi Lab.
Language: Python
#audio #audio_language_model #audio_understanding #fun_asr #multimodal_large_language_models #pytorch #speaker_diarization #speech_recognition
Stars: 264 Issues: 4 Forks: 8
https://github.com/FunAudioLLM/Fun-ASR
GitHub
GitHub - QwenAudio/Fun-ASR: Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with…
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes. - QwenAudio/Fun-ASR
ronitsingh10/FineTune
FineTune, a macOS menu bar app to control volume for each app independently, route apps to different output devices, and apply EQ
Language: Swift
#audio #audio_utility #macos #macos_app #menu_bar #menubar #menubar_app #swift #swiftui #utility
Stars: 782 Issues: 19 Forks: 29
https://github.com/ronitsingh10/FineTune
FineTune, a macOS menu bar app to control volume for each app independently, route apps to different output devices, and apply EQ
Language: Swift
#audio #audio_utility #macos #macos_app #menu_bar #menubar #menubar_app #swift #swiftui #utility
Stars: 782 Issues: 19 Forks: 29
https://github.com/ronitsingh10/FineTune
GitHub
GitHub - ronitsingh10/FineTune: FineTune, a macOS menu bar app for per-app volume control, multi-device output, audio routing,…
FineTune, a macOS menu bar app for per-app volume control, multi-device output, audio routing, and 10-band EQ. Free and open-source alternative to SoundSource. - ronitsingh10/FineTune
OpenMOSS/MOSS-TTS-Nano
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.
Language: Python
#audio_tokenizer #chinese #english #multi_modality #multilingual #realtime #streaming_audio #tts #voice_clone
Stars: 855 Issues: 13 Forks: 78
https://github.com/OpenMOSS/MOSS-TTS-Nano
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.
Language: Python
#audio_tokenizer #chinese #english #multi_modality #multilingual #realtime #streaming_audio #tts #voice_clone
Stars: 855 Issues: 13 Forks: 78
https://github.com/OpenMOSS/MOSS-TTS-Nano
GitHub
GitHub - OpenMOSS/MOSS-TTS-Nano: MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the…
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run direc...
❤1
wassgha/rescript
🎬 Open source, transcript-based video/audio editor that lives in the browser.
Language: TypeScript
#audio #media #media_editing #media_editor #transcribe #transcription #video #video_editing #video_editor #video_processing
Stars: 478 Issues: 4 Forks: 54
https://github.com/wassgha/rescript
🎬 Open source, transcript-based video/audio editor that lives in the browser.
Language: TypeScript
#audio #media #media_editing #media_editor #transcribe #transcription #video #video_editing #video_editor #video_processing
Stars: 478 Issues: 4 Forks: 54
https://github.com/wassgha/rescript
GitHub
GitHub - wassgha/rescript: 🎬 Open source, transcript-based video/audio editor that lives in the browser.
🎬 Open source, transcript-based video/audio editor that lives in the browser. - wassgha/rescript
❤2🔥1👏1
sohaibdevv/youtube-music
A lightweight, ad‑free client for streaming music from YouTube Music. No subscription required. Supports background playback, search, and custom playlists via the reverse‑engineered API.
Language: TypeScript
#ad_free #audio_streaming #background_playback #desktop_app #free #music #music_player #no_ads #open_source #player #playlist_manager #streaming #windows #youtube #youtube_music #yt_music
Stars: 836 Issues: 0 Forks: 0
https://github.com/sohaibdevv/youtube-music
A lightweight, ad‑free client for streaming music from YouTube Music. No subscription required. Supports background playback, search, and custom playlists via the reverse‑engineered API.
Language: TypeScript
#ad_free #audio_streaming #background_playback #desktop_app #free #music #music_player #no_ads #open_source #player #playlist_manager #streaming #windows #youtube #youtube_music #yt_music
Stars: 836 Issues: 0 Forks: 0
https://github.com/sohaibdevv/youtube-music
GitHub
GitHub - sohaibdevv/youtube-music: A lightweight, ad‑free client for streaming music from YouTube Music. No subscription required.…
A lightweight, ad‑free client for streaming music from YouTube Music. No subscription required. Supports background playback, search, and custom playlists via the reverse‑engineered API. - sohaibde...
💩1
crmne/fastpotify
Spotify, native and fast. One lightweight Rust app for your whole library, local playback, and Spotify Connect on Linux, macOS, and Windows.
Language: Rust
#audio #cross_platform #desktop_app #egui #gui #librespot #linux #macos #mpris #music #music_player #rust #spotify #spotify_client #spotify_connect #windows
Stars: 514 Issues: 20 Forks: 27
https://github.com/crmne/fastpotify
Spotify, native and fast. One lightweight Rust app for your whole library, local playback, and Spotify Connect on Linux, macOS, and Windows.
Language: Rust
#audio #cross_platform #desktop_app #egui #gui #librespot #linux #macos #mpris #music #music_player #rust #spotify #spotify_client #spotify_connect #windows
Stars: 514 Issues: 20 Forks: 27
https://github.com/crmne/fastpotify
GitHub
GitHub - crmne/fastpotify: Spotify, native and fast. One lightweight Rust app for your whole library, local playback, and Spotify…
Spotify, native and fast. One lightweight Rust app for your whole library, local playback, and Spotify Connect on Linux, macOS, and Windows. - crmne/fastpotify
👍1
jub0t/Concat
Free & Open-Source CapCut replacement.
Language: TypeScript
#audio_processor #auto_caption #automation #capcut #capcut_alternative #content_creation #cross_platform #desktop_app #ffmpeg #free_video_editor #non_linear_editor #offline_first #open_source_video_editor #rust_lang #tauri_app #video_editing #video_editing_software #video_processing_tool #whisper_cpp
Stars: 762 Issues: 7 Forks: 57
https://github.com/jub0t/Concat
Free & Open-Source CapCut replacement.
Language: TypeScript
#audio_processor #auto_caption #automation #capcut #capcut_alternative #content_creation #cross_platform #desktop_app #ffmpeg #free_video_editor #non_linear_editor #offline_first #open_source_video_editor #rust_lang #tauri_app #video_editing #video_editing_software #video_processing_tool #whisper_cpp
Stars: 762 Issues: 7 Forks: 57
https://github.com/jub0t/Concat
GitHub
GitHub - jub0t/Concat: Free & Open-Source CapCut replacement.
Free & Open-Source CapCut replacement. Contribute to jub0t/Concat development by creating an account on GitHub.
🎉1