Explain the role of Tensor Cores and mixed-precision training (e.g., using FP16 or bfloat16) in accelerating the training of large generative AI models.
Consider the following Python code snippet utilizing the Hugging Face Transformers library for multimodal processing. The objective is to perform visual question answering (VQA). Assume 'image' is a PIL Image object and 'question' is a string. However, the code is incomplete. Choose the options to complete the code.
In ML applications, which machine learning algorithm is commonly used for creating new data based on existing data?
Consider a scenario where you are developing a multimodal A1 system to translate sign language videos into text. The system utilizes a CNN for processing video frames and an RNN for generating the text sequence. During evaluation, you observe that the system struggles to accurately translate signs that involve complex hand movements or subtle facial expressions. What are the MOST effective strategies to improve performance in this specific scenario? (Select TWO)
You are tasked with building a system that can generate captions for images. You want to use a transformer-based model. During inference, you notice that the model tends to generate repetitive captions. Which of the following decoding strategies could you use to mitigate this issue?