Models AI Projects
Daily ranking page for Models open-source AI repositories.
Models tracks 1545 repositories with 5768220 total GitHub stars.
- cactus-compute/needle - 14MB foundation model for tiny devices; phones, wearables, smart home, and robots. (7987 stars, Python, Models)
- vixhal-baraiya/microgpt-c - The most atomic way to train and inference a GPT in pure, dependency-free C (818 stars, C, Models)
- Tongyi-MAI/MAI-UI - Qwen-UI-Agent: Towards Next-Generation Real-World Centric Foundation GUI Agent (1992 stars, Jupyter Notebook, Models)
- shiyu-coder/Kronos - Kronos: A Foundation Model for the Language of Financial Markets (37630 stars, Python, Models)
- openai/whisper - Robust Speech Recognition via Large-Scale Weak Supervision (107663 stars, Python, Models)
- OpenBMB/VoxCPM - VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning (35912 stars, Python, Models)
- kairos-agi/kairos - Official code for world model Kairos (2499 stars, Python, Models)
- AutoArk/GPA - [AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model! (2518 stars, Python, Models)
- index-tts/index-tts - An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System (23234 stars, Python, Models)
- ultralytics/ultralytics - Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking (60792 stars, Python, Models)
- google-research/timesfm - TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting. (28050 stars, Python, Models)
- fishaudio/fish-speech - SOTA Open Source TTS (32287 stars, Python, Models)
- MoonshotAI/Kimi-K3 - Open Frontier Intelligence (8545 stars, Unknown, Models)
- BytedTsinghua-SIA/CUDA-Agent - CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation (1233 stars, Python, Models)
- QwenLM/Qwen3.8 - Qwen3.8 is the large language model series developed by Qwen team, Alibaba Group. (3879 stars, Unknown, Models)
- facebookresearch/segment-anything - The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and exampl... (54726 stars, Jupyter Notebook, Models)
- facebookresearch/sam3 - The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained mode... (11407 stars, Python, Models)
- snakers4/silero-vad - Silero VAD: pre-trained enterprise-grade Voice Activity Detector (10019 stars, Python, Models)
- deepseek-ai/DeepSeek-V3 - (104352 stars, Python, Models)
- CompVis/stable-diffusion - A latent text-to-image diffusion model (73321 stars, Jupyter Notebook, Models)
- modelscope/FunASR - Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MC... (19939 stars, Python, Models)
- k2-fsa/sherpa-onnx - Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Intern... (14266 stars, C++, Models)
- ByteDance-Seed/Depth-Anything-3 - Depth Anything 3 (6172 stars, Python, Models)
- google-deepmind/weathernext - (7571 stars, Python, Models)
- deepinsight/insightface - State-of-the-art 2D and 3D Face Analysis Project (29536 stars, Python, Models)
- Wan-Video/Wan2.2 - Wan: Open and Advanced Large-Scale Video Generative Models (17210 stars, Python, Models)
- QwenAudio/SenseVoice - Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection. (9116 stars, C, Models)
- xinntao/Real-ESRGAN - Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration. (36527 stars, Python, Models)
- NVIDIA/cosmos - NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, sma... (11567 stars, Jupyter Notebook, Models)
- facebookresearch/vggt-omega - [CVPR 2026 Oral] VGGT Omega (4070 stars, Python, Models)
- NVlabs/LongLive - Long Video Gen Infrastructure (2557 stars, Python, Models)
- jingyaogong/minimind-o - 🎙️ A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing! (2353 stars, Python, Models)
- OpenBMB/MiniCPM-V - A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone (26199 stars, Python, Models)
- QwenAudio/CosyVoice - Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability. (22829 stars, Python, Models)
- QwenLM/Qwen3-TTS - Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech genera... (13025 stars, Python, Models)
- NVIDIA/Isaac-GR00T - NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots. (7868 stars, Python, Models)
- MeiGen-AI/InfiniteTalk - Unlimited-length talking video generation that supports image-to-video and video-to-video generation (7663 stars, Python, Models)
- NVlabs/GR00T-WholeBodyControl - Welcome to GR00T Whole-Body Control (WBC)! This is a unified platform for developing and deploying advanced humanoid controllers. This includes: Decoupl... (3380 stars, Python, Models)
- QwenLM/Qwen - The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud. (21614 stars, Python, Models)
- supertone-inc/supertonic - Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX. (13703 stars, Swift, Models)
- Const-me/Whisper - High-performance GPGPU inference of OpenAI's Whisper automatic speech recognition (ASR) model (10638 stars, C++, Models)
- OpenBMB/MiniCPM - MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful. (10211 stars, Jupyter Notebook, Models)
- Lightricks/LTX-2 - Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model. (9166 stars, Python, Models)
- DepthAnything/Depth-Anything-V2 - [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation (8681 stars, Python, Models)
- PriorLabs/TabPFN - ⚡ TabPFN: Foundation Model for Tabular Data ⚡ (7821 stars, Python, Models)
- zai-org/GLM-5 - GLM-5: From Vibe Coding to Agentic Engineering (7009 stars, Unknown, Models)
- google-research/tabfm - TabFM (Tabular Foundation Model) is a pretrained tabular foundation model developed by Google Research for tabular data regression and classification. (2522 stars, Python, Models)
- sapientinc/HRM-Text - HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning. (1865 stars, Python, Models)
- kyutai-labs/moshi - Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec. (10895 stars, Python, Models)
- NVlabs/Eagle - Eagle: Frontier Vision-Language Models with Data-Centric Strategies (3423 stars, Python, Models)
- pnnbao97/VieNeu-TTS - Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text... (2363 stars, Python, Models)
- yuantianyuan01/FastWAM - Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination? (1328 stars, Python, Models)
- bytedance/Bernini - Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer. (1269 stars, Python, Models)
- facebookresearch/sam2 - The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints,... (19728 stars, Jupyter Notebook, Models)
- google-deepmind/alphafold3 - AlphaFold 3 inference pipeline. (8477 stars, Python, Models)
- LiheYoung/Depth-Anything - [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation (8182 stars, Python, Models)
- myshell-ai/MeloTTS - High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean. (7587 stars, Python, Models)
- maziarraissi/PINNs - Physics Informed Deep Learning: Data-driven Solutions and Discovery of Nonlinear Partial Differential Equations (6098 stars, Python, Models)
- OpenMOSS/MOSS-TTS-Nano - MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed... (4200 stars, Python, Models)
- limix-ldm-ai/LimiX - LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence https://arxiv.org/abs/2509.03505 (3995 stars, Python, Models)