Poster

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Jiahui Gao ⋅ Renjie Pi ⋅ Jipeng Zhang ⋅ Jiacheng Ye ⋅ Wanjun Zhong ⋅ Yufei Wang ⋅ Lanqing HONG ⋅ Jianhua Han ⋅ Hang Xu ⋅ Zhenguo Li ⋅ Lingpeng Kong

2025 Poster

[ Poster] [ OpenReview]

Abstract

Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, current work has been largely focused on text-based mathematical problems, with limited investigation in problems involving multi-modal geometric information. Addressing this gap, we aim to enable LLMs to solve geometric problems by understanding image input. We first identify the limitations of current Multimodal Large Language Models (MLLMs) in this area: they struggle to accurately comprehend basic geometric elements and their relationships. To address these challenges, we leverage the inherent attribute of logical structure compactness in geometric figures, utilizing text-only Large Language Models (LLMs) to curate a comprehensive multimodal geometry dataset. This dataset, named Geo170k, contains more than 170K geometric image-caption and question-answer pairs. Utilizing the Geo170k dataset, we introduce G-LLaVA, a model that demonstrates exceptional performance in solving geometric problems. It significantly outperforms GPT4-V on the geometry task of MathVista benchmark with only 7B parameters.

Video

Chat is not available.