| Exam Code/Number: | NCA-GENMJoin the discussion |
| Exam Name: | NVIDIA Generative AI Multimodal |
| Certification: | NVIDIA |
| Question Number: | 58 |
| Publish Date: | Oct 06, 2026 |
|
Rating
100%
|
|
For building a zero-shot image classification pipeline, what could be a crucial step in the process?
How does CLIP understand the content of both text and images?
What does 'modality alignment' refer to?
What advantage does multimodal learning have over unimodal learning?
You are developing a ML model for image classification. You have a dataset with 10,000 images of cats, dogs and birds. Which of the following ML models would be the most appropriate choice for this task?