Public · Protected · Private
Top Models by Domain & Category -huggingface
Type: Public  |  Created: 2026-09-01  |  Frozen: No
« Previous Public Blog Next Public Blog »
Comments
  • A comprehensive blog catalog showcasing the highest-performing and widely adopted machine learning models across 3D, Vision, Audio, Reasoning, Code, Multimodal, and Science.


    • 10 Curated AI Categories:
    1. 3D & Spatial AI (TripoSR, InstantMesh, Shap-E, TRELLIS, etc.)
    2. Text-to-Image & Generative Art (FLUX.1-dev, SD 3.5 Large, SDXL, ControlNet, etc.)
    3. LLMs & Reasoning (DeepSeek-R1, Llama-3.3-70B, Qwen2.5-72B, Mistral-Small, etc.)
    4. Vision-Language & Multimodal (Qwen2.5-VL, InternVL2.5, PaliGemma, CLIP, etc.)
    5. Computer Vision & Segmentation (SAM 2, Depth Anything V2, Grounding DINO, etc.)
    6. Audio, Speech & Music (Whisper Large V3, MusicGen, XTTS-v2, Bark, etc.)
    7. Code Generation (Qwen2.5-Coder, DeepSeek-Coder-V2, Codestral, StarCoder2, etc.)
    8. Embeddings, Search & RAG (BGE-M3, BGE-Reranker, MiniLM-L6-v2, Nomic-Embed, etc.)
    9. Science, Mathematics & Biology (Qwen-Math, ESM-2, DeepSeek-Math, LeRobot, etc.)
    10. Video Generation (LTX-Video, Stable Video Diffusion, CogVideoX, etc.)
    • Direct Hugging Face Links & Task Badges: Each card links directly to its official Hugging Face repository and specifies the pipeline task tag.


    2026-09-01 11:40
  • 3D & Spatial Generative AI

    • stabilityai/TripoSR image-to-3d
    • Ultra-fast single-image to 3D textured mesh feedforward reconstruction model in under 0.5s.
    • TencentARC/InstantMesh image-to-3d
    • Sparse-view multi-view reconstruction framework generating textured 3D assets.
    • openai/shap-e text-to-3d
    • Conditional diffusion architecture outputting implicit neural representations and 3D meshes.
    • OpenGVLab/TRELLIS text-to-3d
    • Generates high-density structured meshes, radiance fields, and PBR materials.
    • openai/point-e text-to-3d
    • Point cloud synthesis system for rapid 3D object prototyping.
    • VAST-AI/Tripo3D image-to-3d
    • Generates production-ready quad-dominant mesh topologies for CAD and DCC.
    • cvlab-columbia/Zero123 image-to-3d
    • Novel view synthesis model generating 3D views from a single reference image.

    Text-to-Image & Generative Art

    Large Language Models (LLMs) & Reasoning

    Vision-Language & Multimodal

    Computer Vision & Segmentation

    Audio, Speech & Music Synthesis

    • openai/whisper-large-v3 automatic-speech-recognition
    • Gold-standard sequence-to-sequence multilingual transcription and translation model.
    • facebook/musicgen-large text-to-audio
    • Autoregressive audio transformer generating stereo musical compositions from text.
    • coqui/XTTS-v2 text-to-speech
    • Multilingual voice cloning speech synthesis network using 3-second reference samples.
    • suno/bark text-to-speech
    • Generates natural expressive speech, background audio, laughter, and emotional tone.

    Code Generation

    Vector Search & Embeddings

    • BAAI/bge-m3 feature-extraction
    • Multi-lingual embedding generator supporting dense, sparse lexical, and multi-vector search.
    • BAAI/bge-reranker-v2-m3 text-classification
    • Cross-encoder model designed to re-score retrieval precision in advanced RAG architectures.
    • sentence-transformers/all-MiniLM-L6-v2 sentence-similarity
    • Standard lightweight model converting sentences into compact 384-dimensional dense vectors.

    Science & Mathematics

    Video Generation & Motion Synthesis


    2026-09-01 11:45