You are tasked with building a system that generates realistic images based on both textual descriptions and a semantic segmentation map. The segmentation map provides spatial information about the objects present in the scene. Which of the following generative architectures is MOST appropriate for this multimodal task?
When using prompt engineering with text-to-image models, which of the following techniques are most effective in improving the fidelity and relevance of generated images to the input text?
During data analysis for a multimodal A1 project involving image and text data, you discover that the image dataset contains a large number of blurry or low-resolution images. The text data, however, is relatively clean and well-structured. What is the BEST approach to mitigate the impact of the noisy image data on the overall model performance?
You're tasked with building a system that can generate realistic images from text descriptions and, conversely, generate accurate text descriptions from images. You decide to use a GAN (Generative Adversarial Network) architecture, but need to handle both modalities effectively. What GAN variant would be MOST suitable for this bi-directional multimodal task?
You are working on a project that involves analyzing customer reviews which contains the following dataset: 1. customer_id(categorical) 2. customer_review(text) 3. product_image(image) 4. video_of_product_usage(video) What is the best way to handle and address the problem of skewness across each modailities?