Multimodal AI is already in the tools you use daily and most people have not fully processed what that means
McKinsey's multimodal AI explainer https://www.mckinsey.com/featured-insights/mckinsey-explainers/what-is-multimodal-ai covers the shift from text-only AI to systems that can process and combine text, images, audio and video and it is worth reading specifically because the shift has already happened rather than being a future development.
GPT-4o, Gemini, Claude, Grok all handle multiple input types. That is multimodal. Most users interact primarily with text because that is the default interface, but the underlying capability is already there and increasingly being used in workflows that involve documents with charts, screenshots of interfaces, audio from meetings, and video content.
The business applications the McKinsey piece covers, marketing, fraud detection, healthcare, call centres, user testing, are the categories where multimodal capability changes what is possible rather than just making existing things faster.
The question the article raises implicitly is worth making explicit: if AI can now process text, images, audio and video in combination, which of your current workflows that touch multiple media types are candidates for AI assistance that were not candidates when AI was text-only?
For people who have been using multimodal features of AI tools: which input modality beyond text has been most useful to you in practice, images, audio, or video?