Stars
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
Ideogram 4: Open image model at the forefront of design
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT), with only 8B DiT parameters, it reaches…
This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
[ICML 2026] | Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
[ECCV 2026 Oral] Implementation of "Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length"
🚀 An awesome list of curated Nano Banana pro prompts and examples. Your go-to resource for mastering prompt engineering and exploring the creative potential of the Nano banana pro(Nano banana 2) AI…
Official Implementation of "MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation"
FIBO is a SOTA, first open-source, JSON-native text-to-image model built for controllable, predictable, and legally safe image generation.
Native Multimodal Models are World Learners
Generate long Sora 2 videos that exceed OpenAI's native 12-second limit
[ICLR 2026] Official Repo for Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
[SIGGRAPH Asia 2025] WorldExplorer: Towards Generating Fully Navigable 3D Scenes
[SIGGRAPH-ASIA 2025] Official implementation of "VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models"
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
Lynx: Towards High-Fidelity Personalized Video Generation
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
A minimal implementation of DeepMind's Genie world model
Repo for Qwen Image Finetune
[3DV 2026] SpatialGen: Layout-guided 3D Indoor Scene Generation
[CVPR 25] Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation
