Multimodal SDXL Fine-Tuning for Controllable Ulos Motif Generation
September 17, 2026
A framework combining Stable Diffusion XL (via LoRA) with LLaMA 1.5-7B enables controllable pattern generation through text, image, and semantic map conditioning. Testing shows that combining text, image, and semantic maps yields the highest performance with an FID of 270 and CLIP scores between 0.65 and 0.70.
HOW THIS AFFECTS YOU
●
researcherYou can explore how multimodal conditioning combinations impact generative performance in specialized domains.
●
designerYou can achieve higher spatial and semantic control over culturally specific pattern generation.