Introduction to visual and multimodal AI
- Overview of visual AI and its role in modern applications
- Understanding multimodal models and their capabilities
- Key use cases for image and document intelligence
- Challenges in processing and interpreting visual data
Working with Azure AI services for visual data
- Introduction to Azure AI vision and document intelligence services
- Processing images and extracting metadata
- Analysing documents for structured and unstructured information
- Integrating visual services into application workflows
Multimodal models for images and documents
- Combining visual and language inputs in AI systems
- Understanding embeddings for multimodal data
- Enhancing context through cross-modal reasoning
- Use cases for multimodal generative AI applications
Structured information extraction from visual inputs
- Techniques for extracting entities and key-value pairs from documents
- Processing forms, receipts, and structured layouts
- Handling unstructured visual data sources
- Improving accuracy through preprocessing and validation
Grounding AI responses in visual data
- Principles of grounding in multimodal AI systems
- Linking model outputs to visual context
- Reducing hallucinations through grounded responses
- Designing systems that maintain traceability to source data
Agent-based orchestration for visual workflows
- Introduction to agent-based AI patterns
- Orchestrating tools and services for visual data processing
- Designing workflows that combine reasoning and action
- Integrating APIs and external tools into AI pipelines
Designing decision-making workflows using visual data
- Building end-to-end pipelines for insight extraction
- Connecting visual analysis to business decision processes
- Automating actions based on extracted insights
- Monitoring and improving workflow performance over time
Exams and assessments
There are no formal exams included in this course. Learners will complete interactive knowledge checks and practical exercises to reinforce their understanding of multimodal AI concepts and visual data processing techniques.
Hands-on learning
This course includes:
- Guided labs using Azure AI services for image and document analysis
- Practical exercises in multimodal model integration
- Scenario-based tasks focused on real-world visual data challenges
- Instructor-led discussions on designing production-ready AI workflows