---
title: How we’re experimenting with AI to benchmark climate ambition in contracts
date: 2026-10-05T14:15:16Z
modified: 2026-10-05T14:43:52Z
permalink: "https://chancerylaneproject.org/news/how-were-experimenting-with-ai-to-benchmark-climate-ambition-in-contracts/"
type: post
status: publish
excerpt: This blog outines how we’re experimenting with AI to benchmark climate ambition in contracts, what we learned from testing early prototypes, and how legal engineering and user research are shaping what we build next.
wpid: 11628
categories:
  - Legal Design
featured_image: "https://chancerylaneproject.org/wp-content/uploads/2026/10/ruthson-zimmerman-FVwG5OzPuzo-unsplash-scaled.jpg"
timestamp: 2026-10-05T14:43:52Z
tags:
  - Legal Design
---

_This is the second in a series on what we learned while building and testing AI tools for climate-aligned contracting. Read blog 1_ [_here._](https://chancerylaneproject.org/wp-content/uploads/wp-mfa-exports/post/the-chancery-lane-project-opens-its-climate-legal-knowledge-to-ai-with-new-api-and-model-context-protocol.md)

Contracts can turn climate commitments into concrete obligations, but climate ambitions are not always reflected in enforceable, legal language. For five years, TCLP has created and curated practical legal knowledge, including [model clauses](https://chancerylaneproject.org/clauses/), to help organisations use contracts to drive positive climate outcomes.

Earlier this year, we set out to explore whether AI-enabled benchmarking tools could make that knowledge easier to apply by helping people identify gaps in how climate ambition is reflected in their contracts. This post is about how legal engineering and user research worked together to test whether that idea holds up, and what we learned about where to take it next.

## **What we wanted to find out**

We had three questions going in. Would people want a tool that gave them a benchmarked assessment of their contract? Would they know what to do with the results? And where in their real-life workflow would something like this be useful?

## **What we built**

We created and published three prototypes on [TCLP Labs](https://labs.chancerylaneproject.org/). Each tool takes a contract, runs a gap analysis against a recognised climate standard (for example, the Global Reporting Initiative, Science Based Targets Initiative), and shows where the language falls short. Within each standard, the tool scores a contract against a set of ambition markers, which are the specific things we decided a “good” climate-aligned contract should include.

![](https://chancerylaneproject.org/wp-content/uploads/2026/10/contracting-workflow.jpg)

IMAGE: _Where the three prototypes work in the contracting process. Two sit at drafting and review, checking language that has already been written. The third sits at tender and specification, where requirements are set before anything reaches a lawyer._
This was early prototype testing to understand the viability and the deeper user needs around a potential solution. The climate ambition markers, the dimensions each contract gets scored against, were written by hand rather than calculated. Interface controls and calls to action were not functional. This is standard practice in user research. It let us find out whether people wanted a benchmark score before we spent months building the engine to produce one.

## **Who we spoke to**

We ran eight semi-structured sessions with 17 people across six countries. Lawyers, procurement leads, sustainability professionals and legal technologists. Some knew TCLP well. Some participants were new to climate-aligned contracting.

We showed users the prototype, asked them to work through it, and asked follow-up questions. We focused on understanding how well the tool’s gap analysis helped users assess their own contracts, rather than testing whether the interface worked. We wanted to know whether the idea underneath was useful and usable, and what else we needed to address before anyone could consider using it with full confidence.

We tested on standard form contracts. Participants in earlier research had told us plainly that they would not upload client-confidential material to an experimental site, so we designed with that in mind from the start.

## **The role of legal**

The technical build was fairly straightforward because we were not building functional applications that needed to chunk documents and run detailed comparisons. We could therefore focus on high-fidelity prototypes, and on making sure the content and user experience we did present felt authentic.

More time went on deciding what counts as good. The legal engineers did this work while the prototypes were being built, and it meant answering questions like these:

- Does “The supplier aims to improve sustainability practices…” count as a commitment, or is it an aspiration nobody has to meet?
- If a contract sets a decarbonisation target but no consequence for missing it, what should that score?
- If a clause covers [scope 1 and 2 emissions](https://chancerylaneproject.org/wp-content/uploads/wp-mfa-exports/glossary-term/scope-1-2-and-3-emissions-and-total-emissions.md) but says nothing about the supply chain, where most emissions sit, is that a partial pass or a fail?

These are core questions that matter and need answering before a tool can run a gap analysis against anything.

![](https://chancerylaneproject.org/wp-content/uploads/2026/10/image-5-1024x467.png)

_IMAGE: interface screenshot showing the user’s contract being analysed against TCLP’s own model clauses rather than aspirational wording that can’t be enforced_## **Our key learnings from the research**

The research set out to answer three questions: whether people wanted a benchmarked assessment at all, whether they would know what to do with one, and where it would sit in their work. The first was straightforward because all users saw the value in the design proposition. The other two produced most of what we learned.

**On knowing what to do with the results**: a gap analysis on its own was not enough. The prototype showed people where their contract language fell short, which is the first pass we were testing. What it did not do was tell them what to change or why it mattered. Several participants wanted a specific next action attached to each gap, something they could take into a meeting or hand to a colleague.

They also needed to see the reasoning behind a score, and this was about accountability rather than doubt. Participants have to justify decisions to clients, suppliers and colleagues.

_One in-house lawyer in Kenya described what they were hoping for rather than what they had seen: “It’s accurate. I believe you will have a database of the laws so that when it’s quoting, it’s quoting something that is existing, not hallucinating.”_ An academic researcher in Estonia was more specific and wanted a dropdown showing the logic behind a score, so anyone looking at 72% could see the working rather than having to assume.

One legal technologist went further. Showing him the original clause next to our recommended one was not much help. He wanted the smallest set of substantive changes that would bring his contract in line with the standard, without discarding language his client had spent months negotiating.

**On where it sits in their work**: we did not get a unified answer here. Some users, mostly lawyers, placed the tool at review, once a draft was in front of them. Others saw it earlier in their workflow, as a way to check a position before drafting started. And some saw it as something to take to clients rather than use internally, to open a conversation about climate terms with people who had not asked for one.

## **Three questions for the next round**

For the next round, we want to bring legal engineering and user research together around three questions:

- **What does a useful next action look like at each gap, and what format does it need to take?** People told us the analysis needs to do more than identify gaps. Participants described a table they could print and take into a meeting, a Word document that fits into an existing drafting workflow, and ready-made wording for a covering note. A lawyer might want redlines or a precedent; a sustainability lead might need something they can take to procurement. We need to understand which outputs matter most, and to whom.
- **How much of the reasoning should we show, and where?** People need to understand how a score was reached before they will act on it. That means making the ambition markers behind each score clearer, while working out how much of that reasoning needs to be visible on screen and how much can sit one click away.
- **Can one tool serve three different points in the process?** We heard three different views about where the tool fits in the contracting process, and they point towards quite different products. The next round needs to establish whether one flexible tool can serve these different needs, or whether they call for more than one tool.

**Some Limitations of the research**We conducted user research with seventeen people across six countries, and while this was enough to see patterns and find problems, it is not enough to tell us how common any of them are.

Nobody used the tools on a real contract of their own, because we did not ask participants to upload confidential material into an experimental tool. We tested against two standards out of many. And every session was a one-off, so we only saw first impressions. We do not yet know how people would use the tool once it became familiar.

## **Get involved**

In the next iteration, we want to dig deeper into the benchmarking score itself and interrogate whether it is useful in the form we have it. That means more detailed sessions, a wider range of contract types, and finding a way to get closer to real work without asking anyone to hand over confidential material.

If you work with contracts and climate obligations, in any role, we would like to hear from you. We are particularly interested in speaking to people outside legal teams.

[Contact us](mailto:http://contact@chancerylaneproject.org.)

## Topics

**Categories:** [Legal Design](https://chancerylaneproject.org/wp-content/uploads/wp-mfa-exports/taxonomy/category/legal-design.md)