First program

AI advice and human judgment

Under what conditions does AI advice help adults handle ordinary disagreements with better judgment and more constructive behavior?

Why this question

The first answer is easy to test. We want to know what happens next.

Existing experimental work already makes this a credible research area. A March 2026 article in Science on sycophancy examined people’s judgments, their prosocial intentions and their preference for agreeable AI.

A new study has to contribute more than a replication of a familiar response benchmark. We start by mapping that evidence and locating a tractable question it leaves unresolved.

The chain we follow

From one piece of advice to how people treat one another.

This is an example research hypothesis, not an established finding. Each link needs its own evidence: a response audit cannot establish the whole chain, and a change observed in one person does not establish a change in public trust.

The chain of consequences we study: Advice in a dispute, then How it’s read, then Repair or escalation, then Patterns over time, then Cooperation and trust.

Advice in a disputeHow it’s readRepair or escalationPatterns over timeCooperation and trustEach link in the chain needs its own evidence.
Advice in a disputeHow it’s readRepair or escalationPatterns over timeCooperation and trust

Each link in the chain needs its own evidence.

Where a study could add something

Five open directions.

  • Durability
  • Subsequent behavior
  • Product differences
  • Usage practices
  • Effects on other people

We assess what a design can actually identify before promising all five.

How a study could work

A possible design.

Compare

Specified AI advice against a real alternative

A few specified AI advice experiences, compared with an appropriate non-AI alternative.

Measure over time

Helpfulness now, judgment later

Immediate perceived helpfulness, measured separately from later judgment, repair attempts and dependence.

Start carefully

Consenting adults, low stakes first

Low-stakes situations to begin, with sample sizes set through power and feasibility work.

Random assignment can identify the effects of assigned conditions, while differences in how much people choose to use AI may remain confounded. Short-term effects will not be described as changes in stable character.

What every report will answer

Six questions, answered in plain language.

  1. What was tested?
  2. With whom?
  3. Against what alternative?
  4. What changed?
  5. How certain are we?
  6. What choice does this inform?

Reports include null and beneficial findings, subgroup limits and version dates. We do not issue a universal moral score for a company or a model.