Home /paper

MM-REACT Prompting ChatGPT for Multimodal Reasoning and Action

Notes on MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action by Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, Lijuan Wang.

The paper, first released in March 2023, proposes MM-REACT, a system that integrates ChatGPT with a collection of computer vision models to solve complicated visual understanding tasks.

Unlike existing vision-language models, which require joint fine-tuning, MM-REACT uses ChatGPT's reasoning abilities to select and invoke specific computer vision models, making it a flexible, training-free approach. As the name suggests, it builds on the ReAct: Synergizing Reasoning and Acting in Language Models pattern of interleaving reasoning with actions.

The paper highlights the capabilities of MM-REACT in various scenarios, such as visual maths and text reasoning, multi-image understanding and video summarisation, demonstrating its potential to address complex visual intelligence problems.

Title and authors of MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Grid of MM-REACT examples across nine capabilities, including visual maths, explaining a meme, locating a frisbee, following a bread recipe, totalling receipts, reading a bar chart, recognising brands and celebrities, and breaking a video into steps.

Flowchart of MM-REACT. ChatGPT responds to the user and, if it produces a thought and action request, the named vision expert (such as image captioning, OCR or Bing search) is run and its output is fed back as an observation. Otherwise it responds to the user.

Example conversation where MM-REACT answers questions about a floor plan image, including the number of bedrooms, room dimensions, kitchen appliances and a final summary.

Example conversation where the user uploads four receipts (a flight, an Uber ride, groceries and a restaurant) and MM-REACT answers how much was spent on groceries, dining out, travel and taxes.

Step-by-step MM-REACT execution on a photo of two basketball players: ChatGPT calls captioning and face detection, then celebrity recognition to identify Kobe Bryant and Paul Pierce, then Bing search to answer a follow-up question about championship rings. Caption for Figure 3 explaining that numbered circles show the order of model calls, and grey text shows the internal thoughts and expert outputs hidden from the user.