Meta Develops Multisensory AI Model That Combines Six Types of Data
Meta has launched ImageBind, an open-source AI model that connects text, audio, visual, movement, thermal, and depth data. The model is only a research project so far, but it shows how future AI models can create immersive, multisensory experiences. The embedding of multiple types of data into a single multidimensional index is the core of the research and generates AI tools, like DALL-E, Stable Diffusion, and Midjourney. The model, ImageBind, is the first that combines several types of data into the same embedding space. Future AI systems are likely to incorporate other sensory input, such as touch, speech, smell, and brain fMRI signals. This research brings the machines one step closer to human ability to learn holistically from different types of information.