-
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
Paper • 2401.09985 • Published • 18 -
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
Paper • 2401.09962 • Published • 9 -
Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution
Paper • 2401.10404 • Published • 10 -
ActAnywhere: Subject-Aware Video Background Generation
Paper • 2401.10822 • Published • 13
Collections
Discover the best community collections!
Collections including paper arXiv:2405.16537
-
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Paper • 2311.04934 • Published • 34 -
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
Paper • 2405.16537 • Published • 17 -
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Paper • 2405.17428 • Published • 19
-
One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning
Paper • 2306.07967 • Published • 25 -
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
Paper • 2306.07954 • Published • 111 -
TryOnDiffusion: A Tale of Two UNets
Paper • 2306.08276 • Published • 74 -
Seeing the World through Your Eyes
Paper • 2306.09348 • Published • 33
-
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
Paper • 2405.16537 • Published • 17 -
ReVideo: Remake a Video with Motion and Content Control
Paper • 2405.13865 • Published • 25 -
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
Paper • 2406.16863 • Published • 11 -
Portrait Video Editing Empowered by Multimodal Generative Priors
Paper • 2409.13591 • Published • 17
-
APISR: Anime Production Inspired Real-World Anime Super-Resolution
Paper • 2403.01598 • Published • 3 -
AnimateDiff-Lightning: Cross-Model Diffusion Distillation
Paper • 2403.12706 • Published • 18 -
1.78k
DALLE 3 XL v2
🔥Generate images from text prompts
-
287
CosXL
💻Generate and edit images using text prompts
-
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
Paper • 2401.09985 • Published • 18 -
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
Paper • 2401.09962 • Published • 9 -
Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution
Paper • 2401.10404 • Published • 10 -
ActAnywhere: Subject-Aware Video Background Generation
Paper • 2401.10822 • Published • 13
-
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
Paper • 2405.16537 • Published • 17 -
ReVideo: Remake a Video with Motion and Content Control
Paper • 2405.13865 • Published • 25 -
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
Paper • 2406.16863 • Published • 11 -
Portrait Video Editing Empowered by Multimodal Generative Priors
Paper • 2409.13591 • Published • 17
-
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Paper • 2311.04934 • Published • 34 -
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models
Paper • 2405.16537 • Published • 17 -
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Paper • 2405.17428 • Published • 19
-
APISR: Anime Production Inspired Real-World Anime Super-Resolution
Paper • 2403.01598 • Published • 3 -
AnimateDiff-Lightning: Cross-Model Diffusion Distillation
Paper • 2403.12706 • Published • 18 -
1.78k
DALLE 3 XL v2
🔥Generate images from text prompts
-
287
CosXL
💻Generate and edit images using text prompts
-
One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning
Paper • 2306.07967 • Published • 25 -
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
Paper • 2306.07954 • Published • 111 -
TryOnDiffusion: A Tale of Two UNets
Paper • 2306.08276 • Published • 74 -
Seeing the World through Your Eyes
Paper • 2306.09348 • Published • 33