CASE STUDY / 05HUMAN–AI · CONFIDENCE · WEB EXPERIMENTS

Decision
Intelligence

Two linked research programs that help people decide when to rely on AI, when to trust themselves, and how to preserve agency under uncertainty.

RoleAI Software Engineer / Research Collaborator
OrganizationsHKUST + ECNU
Evidence795 participants · 6 studies
OutcomeCHI 2023 · CHI 2024
A human and AI jointly evaluating confidence and uncertainty before a decision
01 / THE PROBLEM

The right question is not “Should I trust AI?” but “Who is more likely to be right here?”

AI advice can improve a decision, but it can also trigger over-reliance or cause people to reject useful recommendations. AI confidence describes only one side of the partnership: appropriate reliance also depends on the person’s own capability and whether their self-confidence matches their actual performance.

Over-relianceA person follows incorrect AI advice despite having the stronger initial judgment.
Under-relianceA person rejects correct AI advice when the system is more likely to be right.
Mis-calibrated confidenceA person’s certainty does not reliably reflect their probability of being correct.

Decision scenario

01Judge independently

Review one case and make an unaided prediction.

→
02Estimate confidence

Report certainty or build a model of personal capability.

→
03See AI advice

Inspect the model’s prediction and calibrated confidence.

→
04Make the final call

Keep or revise the answer with both parties in view.

02 / RESEARCH PROGRAM

Two complementary ways to support appropriate reliance.

The work progressed from estimating task-specific human capability to improving the quality of the confidence people bring into an AI-assisted decision.

TRACK A · CHI 2023Model both decision makers

Estimate human correctness likelihood from a person’s decisions and editable rules, compare it with calibrated AI confidence, then adapt how advice is presented.

Question: who is more likely to be correct on this instance?
TRACK B · CHI 2024Calibrate the human first

Test reflection, betting, and feedback mechanisms that help people align self-confidence with actual accuracy before they encounter AI advice.

Question: can better self-knowledge improve reliance?

Unifying insight: decision support should represent the capability of both the human and the AI—and account for how accurately the person perceives their own capability.

03 / EXPERIMENT PLATFORM

A reusable web system connecting task interfaces, capability models, and behavioral evidence.

PARTICIPANT CLIENTBrowser-based decision tasks

Renders case information, confidence inputs, AI advice, adaptive interventions, and final decisions while preserving randomized experimental conditions.

HTTPStask state + responses
APPLICATION LAYERFlask study services

Initializes sessions, assigns conditions, serves interaction logic, applies participant-edited rules, and records the complete decision trajectory.

MODEL + DATApredictions + likelihoods
ANALYTICS LAYERHuman and AI capability

Uses scikit-learn models, calibrated confidence, participant rule models, and structured logs for behavioral and statistical analysis.

PythonFlaskscikit-learnNumPy / PandasJavaScriptProlific

How data moves through a study

CONTROLLED STUDY FLOWFrom an independent judgment to a measurable reliance decision
01 / INPUTTask instance

A participant reviews a structured income-prediction case.

→
02 / BASELINEInitial answer

The unaided judgment and self-confidence establish the human baseline.

→
03 / SUPPORTCondition logic

The server returns AI advice, capability cues, or a calibration intervention.

→
04 / OUTCOMEFinal answer

Advice taking, answer changes, confidence, and task performance are captured.

→
05 / ANALYSISBehavioral record

Trial-level logs support reliance, workload, trust, and accuracy analyses.

04 / TRACK A · CHI 2023

Estimate human correctness likelihood, then compare it with AI confidence.

A person first completes 20 unaided decisions. The system fits a compact decision tree, converts it into editable if–then rules, and retrieves similar historical cases to estimate local human correctness likelihood. This estimate is compared with calibrated AI confidence at each new decision.

01 / OBSERVECollect unaided decisions

Use labeled examples to capture individual strengths, weaknesses, and decision patterns.

02 / MAKE LEGIBLEBuild editable rules

Translate a depth-limited decision tree into rules the participant can inspect, add, edit, or remove.

03 / ESTIMATECompute local human CL

Apply the personal rule model to similar cases and estimate correctness likelihood for the current instance.

04 / SUPPORTCompare human and AI

Use Direct Display, Adaptive Workflow, or Adaptive Recommendation to communicate the capability difference.

REAL CHI 2023 INTERFACES

Human capability became inspectable—and actionable.

These published interfaces show the editable human model and the five experimental conditions implemented for the study.

Interactive decision tree and rule-set editors from the CHI 2023 study
Capability-model interface — participants inspect and revise a decision tree or rule set.
Five human-AI decision-support conditions from the CHI 2023 study
Experiment client — Human Only, AI Confidence, and three strategies using human and AI correctness likelihood.
05 / TRACK B · CHI 2024

Calibrate self-confidence before asking people to rely on AI.

The second program tested whether improving human self-confidence calibration changes downstream AI reliance. Three mechanisms represented different design philosophies:

REFLECTIONThink the Opposite

Ask people to identify evidence for a different answer and articulate why their initial prediction could be wrong.

INCENTIVEThinking in Bets

Translate subjective certainty into a wager, making the strength of a prediction concrete.

FEEDBACKCalibration Status

Show whether confidence matched correctness in real time and summarize calibration patterns after a task block.

Think the Opposite, Thinking in Bets, and confidence-calibration feedback interfaces from the CHI 2024 study
Published CHI 2024 interfaces — two elicitation mechanisms plus real-time and post-hoc calibration feedback.
06 / VALIDATION

Six controlled studies separated model quality, interface effects, and reliance behavior.

795participants across two publications and six studies
CHI 2023 · 343 PARTICIPANTSCan human capability be estimated and used effectively?

Two preliminary studies evaluated the human-model interface and correctness-likelihood estimation; a 293-person between-subjects study compared five decision-support conditions. The capability-aware strategies promoted more appropriate trust than showing AI confidence alone, while revealing a tradeoff with interaction complexity and mental demand.

CHI 2024 · 452 PARTICIPANTSDoes calibrating self-confidence improve reliance?

Three studies first characterized confidence and reliance, then compared calibration mechanisms, and finally tested feedback with AI advice. Reflection and feedback improved calibration; feedback reduced under-reliance and helped participants follow the higher-confidence party more often, while over-reliance remained harder to correct.

07 / IMPLICATIONS & LIMITS

Make capability visible, but do not turn uncertainty into false certainty.

  • DESIGN FOR BOTH SIDES
    Represent human and AI capability together.

    AI confidence alone cannot tell a person whether they are the stronger decision maker on a specific case.

  • CALIBRATE BEFORE ADVICE
    Support accurate self-knowledge upstream.

    Timely feedback can help people interpret later AI advice, but reflective interventions must justify their added cognitive load.

  • KEEP MODELS LEGIBLE
    Let people inspect how the system represents them.

    Editable rules improve agency, yet a static, simplified human model can still miss edge cases and changing capability.

  • GENERALIZE CAREFULLY
    Treat these results as controlled evidence, not a universal policy.

    The studies used a low-stakes income-prediction task; higher-stakes domains require domain experts, uncertainty-aware estimates, and further real-world validation.

08 / MY SCOPE & COLLABORATION

A focused engineering and conceptual-design contribution within a multi-institution research team.

  • Contributed as second author on the CHI 2023 paper and third author on the CHI 2024 paper—not as the overall project lead.
  • Developed and iterated web-based human–AI decision interfaces used in controlled online studies.
  • Contributed to interaction and conceptual design around human capability, AI confidence, and self-confidence calibration.
  • Supported study execution, analysis, and the translation of behavioral findings into system-design implications.
2Peer-reviewed CHI papers
4Partner institutions across the program

Collaborators across HKUST, East China Normal University, Purdue University, and Southeast University contributed research framing, experimental design, analysis, and supervision.

Shuai MaFirst author · HKUST
Ying LeiResearch collaborator · ECNU
Xinru Wang & Ming YinResearch collaborators · Purdue University
Chengbo Zheng & Chuhan ShiResearch collaborators · HKUST / Southeast University
Xiaojuan MaFaculty supervisor · HKUST
BACK TO CASE STUDY / 01WatchGuardian →