Data

My Climate Copilot User Study Annotations

Commonwealth Scientific and Industrial Research Organisation
Nguyen, Vincent ; Karimi, Sarvnaz ; Hallgren, Willow ; Harkin, Ashley ; Prakash, Mahesh
Viewed: [[ro.stat.viewed]] Cited: [[ro.stat.cited]] Accessed: [[ro.stat.accessed]]
ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Adc&rfr_id=info%3Asid%2FANDS&rft_id=info:doi10.25919/x5wq-n705&rft.title=My Climate Copilot User Study Annotations&rft.identifier=https://doi.org/10.25919/x5wq-n705&rft.publisher=Commonwealth Scientific and Industrial Research Organisation&rft.description=A collection of 357 annotations from 13 climate experts for Climate Adaptation NLP. It contains evaluations of Large Language Model (LLM) capability on climate questions as well as the raw responses from LLMs.This collection is useful for benchmarking models on their generation performance to expert climate adaptation questions and also measuring their evaluation performance in comparison to experts.Lineage: 3 Language Models (LLMs) were used to generate responses to 50 questions posed by climate experts. LLMs generated responses using climate data/scientific literature, and without, amounting to 300 unique responses. 13 Experts were asked to annotate responses from the different with a set of 7 criteria (annotation guidelines included), which were created by experts. 57 responses were annotated twice by a different expert to measure inter-annotator agreement. Evaluation criteria can be found in the user testing guide PDF. Each sub-criterion is labelled in lexicographic order (i.e., structure_a refers to the first sub-criterion in the user testing guide for the structure criteria). A score of 1 indicates that an expert believes the response meets that sub-criteria, while a score of 0 indicates the response failed to meet that sub-criteria.&rft.creator=Nguyen, Vincent &rft.creator=Karimi, Sarvnaz &rft.creator=Hallgren, Willow &rft.creator=Harkin, Ashley &rft.creator=Prakash, Mahesh &rft.date=2025&rft.edition=v1&rft_rights=Creative Commons Attribution Noncommercial-Share Alike 4.0 Licence https://creativecommons.org/licenses/by-nc-sa/4.0/&rft_rights=Data is accessible online and may be reused in accordance with licence conditions&rft_rights=All Rights (including copyright) CSIRO, Bureau of Meteorology 2025.&rft_subject=climate&rft_subject=climate nlp&rft_subject=question answering&rft_subject=nlp&rft_subject=expert evaluation&rft_subject=large language model&rft_subject=llm&rft_subject=generation&rft_subject=llm as judge&rft_subject=Climate change impacts and adaptation not elsewhere classified&rft_subject=Climate change impacts and adaptation&rft_subject=ENVIRONMENTAL SCIENCES&rft_subject=Natural language processing&rft_subject=Artificial intelligence&rft_subject=INFORMATION AND COMPUTING SCIENCES&rft.type=dataset&rft.language=English Access the data

Licence & Rights:

Non-Commercial Licence view details
CC-BY-NC-SA

Creative Commons Attribution Noncommercial-Share Alike 4.0 Licence
https://creativecommons.org/licenses/by-nc-sa/4.0/

Data is accessible online and may be reused in accordance with licence conditions

All Rights (including copyright) CSIRO, Bureau of Meteorology 2025.

Access:

Open view details

Accessible for free

Contact Information



Full description

A collection of 357 annotations from 13 climate experts for Climate Adaptation NLP. It contains evaluations of Large Language Model (LLM) capability on climate questions as well as the raw responses from LLMs.

This collection is useful for benchmarking models on their generation performance to expert climate adaptation questions and also measuring their evaluation performance in comparison to experts.
Lineage: 3 Language Models (LLMs) were used to generate responses to 50 questions posed by climate experts. LLMs generated responses using climate data/scientific literature, and without, amounting to 300 unique responses.

13 Experts were asked to annotate responses from the different with a set of 7 criteria (annotation guidelines included), which were created by experts. 57 responses were annotated twice by a different expert to measure inter-annotator agreement.

Evaluation criteria can be found in the user testing guide PDF. Each sub-criterion is labelled in lexicographic order (i.e., structure_a refers to the first sub-criterion in the user testing guide for the structure criteria). A score of 1 indicates that an expert believes the response meets that sub-criteria, while a score of 0 indicates the response failed to meet that sub-criteria.

Available: 2025-05-20

Data time period: 2024-12-01 to 2025-01-31

This dataset is part of a larger collection

Click to explore relationships graph
ACN 633 798 857