MineClaim: A Benchmark for LLM-Based ESG Claim Verification and Greenwashing Detection in Australian Mining Reports
Lujia Yang, Yiheng Lu, Zherui Wang, Yi Ding, Mingchen Ju, Yifu Tang, Zhengyi Yang
The 1st International Workshop on Natural Resources Survey, Monitoring, and Assessment in the Big Data Era (NRSMA)
RAIDS Lab Authors
Details
Research Area
Tags
Abstract
Environmental, Social and Governance (ESG) reports often contain positive sustainability claims, but it is not always clear whether these claims are supported by evidence disclosed in the same report. Large language models (LLMs) offer a promising way to assist this process, because they can analyse disclosure text and compare stated claims with supporting evidence. However, their effectiveness for evidence-grounded ESG claim assessment remains under-explored. This paper presents a controlled study of ESG claim support assessment in the Australian mining sector. We construct a benchmark from 40 Australian mining ESG reports, containing 90 claim-evidence instances balanced across three labels. We evaluate GPT-5.2, DeepSeek-V4-Flash, and Qwen-3.5-Plus under different prompt and evidence settings. Our results show that LLMs can support evidence-based ESG claim assessment, but current performance remains far from fully reliable. Even the best-performing setting just reach 71% in macro-F1, indicating a clear gap for practical auditing or greenwashing screening. Report context improves judgement, but model behaviour differs. GPT benefits most when extracted metrics are combined with their surrounding context, while DeepSeek and Qwen perform better with full context alone. These findings suggest that LLMs may help auditors and analysts identify potentially unsupported ESG claims, but further work is needed to improve the accuracy and robustness of ESG greenwashing detection.

