MoMA
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
About
we present MoMA: an open-vocabulary, training-free personalized image model that boasts flexible zero-shot capabilities. As foundational text-to-image models rapidly evolve, the demand for robust image-to-image translation grows. Addressing this need, MoMA specializes in subject-driven personalized image generation. Utilizing an open-source, Multimodal Large Language Model (MLLM), we train MoMA to serve a dual role as both a feature extractor and a generator. This approach effectively synergizes reference image and text prompt information to produce valuable image features, facilitating an image diffusion model. To better leverage the generated features, we further introduce a novel self-attention shortcut method that efficiently transfers image features to an image diffusion model, improving the resemblance of the target object in generated images. Remarkably, as a tuning-free plug-and-play module, our model requires only a single reference image and outperforms existing methods in generating images with high detail fidelity, enhanced identity-preservation, and prompt faithfulness. We commit to making our work open-source, thereby providing universal access to these advancements.
Open Source Health
- Stars
- 234
- Forks
- 17
- License
- Not stated
- Last commit
- 2 years ago
Resources & Links
Related Categories
Vendor
ByteDance
Publisher of UI TARS Desktop, Trae Agent and danmu.js
Quick Links
Open Source
More by ByteDance
Related Products
Krea
Krea is the world's most powerful creative AI suite.
Top category match
AI Image Enlarger
Image Enlarger & Upscaler Online
Top category match
Kreator.ai
AI Creative Platform for Images, Video, Ads, and Brand Content
Top category match
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
