Estimate human correctness likelihood from a person’s decisions and editable rules, compare it with calibrated AI confidence, then adapt how advice is presented.
Question: who is more likely to be correct on this instance?Decision
Intelligence
Two linked research programs that help people decide when to rely on AI, when to trust themselves, and how to preserve agency under uncertainty.
The right question is not “Should I trust AI?” but “Who is more likely to be right here?”
AI advice can improve a decision, but it can also trigger over-reliance or cause people to reject useful recommendations. AI confidence describes only one side of the partnership: appropriate reliance also depends on the person’s own capability and whether their self-confidence matches their actual performance.
Decision scenario
Review one case and make an unaided prediction.
Report certainty or build a model of personal capability.
Inspect the model’s prediction and calibrated confidence.
Keep or revise the answer with both parties in view.
Two complementary ways to support appropriate reliance.
The work progressed from estimating task-specific human capability to improving the quality of the confidence people bring into an AI-assisted decision.
Test reflection, betting, and feedback mechanisms that help people align self-confidence with actual accuracy before they encounter AI advice.
Question: can better self-knowledge improve reliance?Unifying insight: decision support should represent the capability of both the human and the AI—and account for how accurately the person perceives their own capability.
A reusable web system connecting task interfaces, capability models, and behavioral evidence.
Renders case information, confidence inputs, AI advice, adaptive interventions, and final decisions while preserving randomized experimental conditions.
Initializes sessions, assigns conditions, serves interaction logic, applies participant-edited rules, and records the complete decision trajectory.
Uses scikit-learn models, calibrated confidence, participant rule models, and structured logs for behavioral and statistical analysis.
How data moves through a study
A participant reviews a structured income-prediction case.
The unaided judgment and self-confidence establish the human baseline.
The server returns AI advice, capability cues, or a calibration intervention.
Advice taking, answer changes, confidence, and task performance are captured.
Trial-level logs support reliance, workload, trust, and accuracy analyses.
Estimate human correctness likelihood, then compare it with AI confidence.
A person first completes 20 unaided decisions. The system fits a compact decision tree, converts it into editable if–then rules, and retrieves similar historical cases to estimate local human correctness likelihood. This estimate is compared with calibrated AI confidence at each new decision.
Use labeled examples to capture individual strengths, weaknesses, and decision patterns.
Translate a depth-limited decision tree into rules the participant can inspect, add, edit, or remove.
Apply the personal rule model to similar cases and estimate correctness likelihood for the current instance.
Use Direct Display, Adaptive Workflow, or Adaptive Recommendation to communicate the capability difference.
Human capability became inspectable—and actionable.
These published interfaces show the editable human model and the five experimental conditions implemented for the study.


Calibrate self-confidence before asking people to rely on AI.
The second program tested whether improving human self-confidence calibration changes downstream AI reliance. Three mechanisms represented different design philosophies:
Ask people to identify evidence for a different answer and articulate why their initial prediction could be wrong.
Translate subjective certainty into a wager, making the strength of a prediction concrete.
Show whether confidence matched correctness in real time and summarize calibration patterns after a task block.

Six controlled studies separated model quality, interface effects, and reliance behavior.
Two preliminary studies evaluated the human-model interface and correctness-likelihood estimation; a 293-person between-subjects study compared five decision-support conditions. The capability-aware strategies promoted more appropriate trust than showing AI confidence alone, while revealing a tradeoff with interaction complexity and mental demand.
Three studies first characterized confidence and reliance, then compared calibration mechanisms, and finally tested feedback with AI advice. Reflection and feedback improved calibration; feedback reduced under-reliance and helped participants follow the higher-confidence party more often, while over-reliance remained harder to correct.
Make capability visible, but do not turn uncertainty into false certainty.
- DESIGN FOR BOTH SIDESRepresent human and AI capability together.
AI confidence alone cannot tell a person whether they are the stronger decision maker on a specific case.
- CALIBRATE BEFORE ADVICESupport accurate self-knowledge upstream.
Timely feedback can help people interpret later AI advice, but reflective interventions must justify their added cognitive load.
- KEEP MODELS LEGIBLELet people inspect how the system represents them.
Editable rules improve agency, yet a static, simplified human model can still miss edge cases and changing capability.
- GENERALIZE CAREFULLYTreat these results as controlled evidence, not a universal policy.
The studies used a low-stakes income-prediction task; higher-stakes domains require domain experts, uncertainty-aware estimates, and further real-world validation.
A focused engineering and conceptual-design contribution within a multi-institution research team.
- Contributed as second author on the CHI 2023 paper and third author on the CHI 2024 paper—not as the overall project lead.
- Developed and iterated web-based human–AI decision interfaces used in controlled online studies.
- Contributed to interaction and conceptual design around human capability, AI confidence, and self-confidence calibration.
- Supported study execution, analysis, and the translation of behavioral findings into system-design implications.
Collaborators across HKUST, East China Normal University, Purdue University, and Southeast University contributed research framing, experimental design, analysis, and supervision.