Collection
Multimodal AI
The Multimodal AI collection tracks 13 curated open-source projects, 13 of them live on TrendingRepo right now — ranked by a cross-source momentum score blending GitHub star velocity with mentions on Hacker News, X, Bluesky, Product Hunt and Dev.to.
Live · top 13 repos · sorted by momentum across 24H
LIVE · 48m| # | Repository | Stars | 24h | 7d | 30d | Trend | Mentions | Actions |
|---|---|---|---|---|---|---|---|---|
| 01 | livekit/agents A framework for building realtime voice AI agents 🤖🎙️📹 | 11.5K | +11+0.1% | +75+0.7% | +369+3.3% | |||
| 02 | pipecat-ai/pipecat Open Source framework for voice and multimodal conversational AI | 13.7K | +4+0.0% | +141+1.0% | +683+5.2% | |||
| 03 | ggml-org/whisper.cpp Port of OpenAI's Whisper model in C/C++ | 52.3K | +30+0.1% | +430+0.8% | +1.3K+2.5% | |||
| 04 | openai/whisper Robust Speech Recognition via Large-Scale Weak Supervision | 105.6K | +114+0.1% | +394+0.4% | +2.1K+2.0% | |||
| 05 | myshell-ai/OpenVoice Instant voice cloning by MIT and MyShell. Audio foundation model. | 37K | +8+0.0% | +52+0.1% | +264+0.7% | |||
| 06 | openai/CLIP CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image | 34.1K | +4+0.0% | +45+0.1% | +240+0.7% | |||
| 07 | haotian-liu/LLaVA [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond. | 24.9K | +2+0.0% | +19+0.1% | +87+0.3% | |||
| 08 | coqui-ai/TTS 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production | 45.8K | +11+0.0% | +35+0.1% | +207+0.5% | |||
| 09 | suno-ai/bark 🔊 Text-Prompted Generative Audio Model | 39.2K | +1+0.0% | +19+0.0% | +100+0.3% | |||
| 10 | salesforce/BLIP PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation | 5.7K | — | +2+0.0% | +11+0.2% | |||
| 11 | lucidrains/imagen-pytorch Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch | 8.4K | — | -1-0.0% | +5+0.1% | |||
| 12 | facebookresearch/ImageBind ImageBind One Embedding Space to Bind Them All | 9.1K | — | — | +15+0.2% | |||
| 13 | microsoft/SpeechT5 Unified-Modal Speech-Text Pre-Training for Spoken Language Processing | 1.4K | — | — | +3+0.2% |