You are training a multimodal model with text and audio inputs. You notice that the audio modality dominates the training process, and the text modality is not contributing significantly to the final performance. Which of the following strategies can you use to address this modality imbalance? (Select TWO)
You have trained a text-to-image diffusion model. During inference, you notice that the generated images often lack fine-grained details and appear blurry. Which of the following techniques could you apply to improve the image quality without retraining the model?
Which technique is commonly used to speed up AI model training and inference on hardware accelerators?
You are tasked with generating realistic images of human faces using a GAN. However, you notice that the generated images often contain artifacts, such as distorted facial features or unrealistic textures. Which of the following techniques would be most effective in improving the realism and quality of the generated faces?
When training a Variational Autoencoder (VAE) for generating new data points, which of the following objectives does the VAE optimize?