Designing trust in AI interfaces
A review of what calibrates reliance, and what does not
Research Essay. Independent research. Not peer reviewed.
Abstract
Product teams building AI features generally treat trust as a quantity to increase. The empirical literature on trust in automation treats it as a quantity to calibrate, and has done since at least 2004.
This note reviews that literature alongside more recent human-computer interaction studies of AI-assisted decision making, and reports an uncomfortable pattern. The mechanisms most widely deployed to build trust, meaning explanations, reasoning traces and stated accuracy figures, have been tested and largely fail to improve calibration, with several studies finding they increase acceptance of incorrect output. The mechanisms with the best support are less popular precisely because they add friction or reduce satisfaction.
A disanalogy between the two literatures is also identified. Earlier automation research studied aids whose errors were, in principle, detectable by an attentive operator. Generative systems produce errors that are stylistically indistinguishable from correct output, which removes the cue much of the older work assumed.
The method is literature and secondary-source analysis. No user research was conducted.
Research question
Which interface mechanisms are empirically supported as improving the calibration of a user's reliance on an AI system, and which widely deployed mechanisms have been tested and found not to?
The framing of the question is deliberate. It does not ask how to increase trust, because the literature's consistent position is that increasing trust is not the objective and that both excessive and insufficient trust are failures with distinct costs.
Background
Trust in automation has been studied for four decades, mostly in aviation, process control and clinical decision support. The construct that emerged is not a single quantity.
Lee and See's formulation separates three properties. Calibration is the correspondence between a person's trust and the system's actual capability. Resolution is the degree to which trust discriminates between situations, so that a user trusts the system where it is strong and doubts it where it is weak. Specificity is whether trust attaches to particular functions or to the system as a whole.
A product can raise trust while worsening all three. A user who trusts a system uniformly and highly has poor resolution and no specificity, and is therefore reliably wrong in the subset of cases where the system fails.
Parasuraman and Riley's taxonomy names the two failure modes. Misuse is reliance where reliance is not warranted. Disuse is rejection of a system that would have helped. Both are expensive, and the design literature attends almost exclusively to the second, because the second is the one that shows up in adoption metrics.
Methodology
Literature and secondary-source analysis, meaning synthesis of published research, specifications and documentation.
The corpus was assembled from two bodies of work: the trust-in-automation tradition in human factors, and studies of AI-assisted decision making published at human-computer interaction and fairness venues since 2019. Practitioner guidance is included where it is widely used in industry, and is labelled as guidance rather than as evidence.
No systematic search protocol was followed and no inclusion criteria were pre-registered. Papers were selected because they are frequently cited on this question or because they report results that contradict common practice. Selection bias is a genuine weakness and is discussed in the limitations.
No original user research was conducted for this note. Where the author's build experience appears, it is declared as practitioner observation from a single-operator sample.
Evidence
Lee and See, 2004, "Trust in Automation: Designing for Appropriate Reliance" (Human Factors 46(1)). The framework paper. It sets the objective as appropriate rather than maximal trust, and identifies three bases on which trust is formed.
Performance: what the system does, its observed reliability and competence. Process: how it works, whether its method is appropriate to the situation and whether its behaviour is understandable. Purpose: why it exists, whether the designer's intent aligns with the user's.
The three are not interchangeable, and the paper's practical argument is that trust formed on one basis does not transfer to a situation the other two would have flagged.
Muir and Moray, 1996 (Ergonomics). Experimental studies in a process control simulation. Trust tracked perceived competence and predictability, and, importantly, trust fell sharply after a fault and recovered slowly afterwards, over many trials.
The asymmetry is the finding to design around. Trust is expensive to build and cheap to destroy, which means the design objective at the point of failure is limiting the depth of the fall rather than preventing the fault.
Parasuraman and Riley, 1997. Use, misuse, disuse and abuse. The taxonomy that makes the two-sided failure explicit, and the source of the observation that operators frequently rely on automation in ways its designers did not anticipate.
Hoff and Bashir, 2015 (Human Factors 57(3)). A review integrating a large body of empirical work into three categories of influence: dispositional, meaning stable individual differences; situational, meaning features of the context and task; and learned, meaning experience with the specific system. The organising contribution matters for design because only the third category is under a product team's control at the interface level.
Dzindolet et al., 2003, "The role of trust in automation reliance" (International Journal of Human-Computer Studies). Participants who saw an automated aid err without explanation lost trust in it sharply. Participants given information about why the aid might err retained higher trust and relied on it more appropriately.
This is the most directly actionable older finding in the corpus. Pre-announcing the failure mode protects trust when the failure arrives.
Yin, Wortman Vaughan and Wallach, 2019, "Understanding the Effect of Accuracy on Trust in Machine Learning Models" (CHI). Stated accuracy influenced trust, but the effect was moderated by observed accuracy during use: when what people saw contradicted what they were told, what they saw won.
The implication is that accuracy claims in marketing or onboarding copy are weak instruments, and that they are overwritten by the first few interactions.
Zhang, Liao and Bellamy, 2020 (FAT*). A study of confidence scores and local feature-importance explanations in an AI-assisted prediction task. Confidence information produced some improvement in trust calibration. Local explanations did not reliably improve calibration, and there were indications they increased acceptance of incorrect predictions.
Bansal et al., 2021, "Does the Whole Exceed its Parts?" (CHI). Explanations increased the rate at which people accepted the model's recommendation whether or not it was correct, and did not reliably improve joint human-AI accuracy.
Taken with Zhang and colleagues, this is the central negative result in the current literature, and it is aimed directly at the most common trust intervention in shipping products.
Buçinca, Malaya and Gajos, 2021, "To Trust or to Think" (CSCW). Cognitive forcing interventions, meaning design changes that require the person to engage before seeing the AI's answer, such as asking for their own judgement first or introducing a deliberate delay, reduced overreliance relative to a standard explanation interface.
The paper also reports the cost honestly: participants preferred the interfaces that produced worse outcomes, and the effect interacted with individual differences in need for cognition. This is the clearest demonstrated tradeoff in the field between decision quality and satisfaction.
Kizilcec, 2016, "How Much Information?" (CHI). Transparency about an algorithmic process helped maintain trust when the outcome violated the user's expectation, but higher levels of transparency did not monotonically increase trust and could reduce it.
Amershi et al., 2019, "Guidelines for Human-AI Interaction" (CHI). Eighteen guidelines synthesised from industry and academic sources and refined through evaluation with practitioners. The relevant cluster is the one about being clear on what the system can do and how well, and about supporting efficient correction and dismissal.
Google's People + AI Guidebook. Practitioner guidance, not research. Useful for its concrete pattern vocabulary, and it should not be cited as evidence.
Analysis
The objective is calibration, and most products do not measure it
Almost no shipping product measures the calibration of user reliance. The available metrics, acceptance rate and satisfaction, are both maximised by overreliance, which means a team optimising them is selecting for the failure mode.
The measurable form of calibration requires knowing when the system was wrong and whether the user caught it. That means logging user overrides against subsequent ground truth, which most products could do and almost none do.
The three bases require three different interface mechanisms
Performance is addressed by showing observed reliability in the user's own context rather than a stated global accuracy figure, since Yin and colleagues indicate the stated figure is overwritten by experience anyway.
Process is addressed by making the shape of the system's competence legible, meaning the classes of input it handles well and badly, in the user's own domain terms.
Purpose is addressed by stating what the feature is optimised for and who benefits. This is the basis that product teams handle worst, and it is where a mismatch is least recoverable, because a user who concludes the system's objective is not theirs has formed a stable belief that no performance improvement addresses.
The popular mechanisms have been tested and mostly fail
Local explanations, feature attributions and reasoning traces increase acceptance without improving accuracy. Confidence scores are, at best, mixed. Stated accuracy is overwritten by observed behaviour.
The reason these persist is not that they are believed to work. It is that they are cheap to build, they look responsible, and their effect on acceptance is easy to observe while their effect on calibration is not measured. That is a metric problem more than a knowledge problem.
The supported mechanisms are unpopular by construction
The interventions with the better evidence share a property: they slow the user down or reduce their confidence.
Cognitive forcing functions reduce overreliance and reduce satisfaction. Pre-announcing failure modes preserves trust through a fault and makes the product sound less capable at the moment of first use. Making uncertainty explicit is honest and makes output feel weaker.
Any team that ships these will see their preference metrics fall. This is a real organisational obstacle and there is no version of the finding that avoids it.
Uncertainty needs to be a first-class state, not a default
Lee and See's resolution property requires that trust discriminate between situations. A user can only do that if the interface distinguishes between situations, which requires the system to be able to say that it does not know.
Most interfaces cannot. A binary state model collapses unknown into one of the two values, usually the safer-looking one, and every unknown then presents itself with the same confidence as a determination. This forces trust to be global, which is precisely the low-resolution condition the framework identifies as unsafe.
The design requirement is that any state a system infers must have three values rather than two, and that the third must be rendered distinctly rather than styled as a weaker version of the others.
Provenance and reversibility are trust architecture, not trust messaging
Given Muir and Moray's asymmetry, the highest-value trust investment is at the point of failure rather than before it. Two mechanisms act there.
Provenance, meaning a durable, inspectable record of what the system did and on what basis, which lets a user diagnose a fault instead of generalising from it. Without it, a single visible error is attributed to the whole system, which is the specificity failure in its worst form.
Reversibility, meaning a real undo that restores prior state, which converts a fault from a loss into an inconvenience.
Both are architectural. Neither can be added by copywriting, and both are frequently deferred because they are invisible when the system is working.
The fluency problem, which the older literature did not face
This is the note's own argument rather than a cited finding, and it is the most important qualification on everything above.
The automation studied in the human factors tradition failed in ways that were detectable in principle. A gauge disagreed with another gauge, a recommendation contradicted visible evidence, an autopilot produced an attitude the pilot could feel. The operator had cues, and the research question was whether they attended to them.
Generative systems remove the cue. A fabricated citation has the same typography, register and confidence as a real one. A wrong summary reads exactly like a right one. The error is not merely hard to notice, it is stylistically identical to correctness, which means the detectability the older studies assumed is absent.
Two consequences follow. The process basis of trust becomes much harder to establish, because the system's competence boundary cannot be inferred from the appearance of its output. And verification has to be structural, meaning citations that resolve, values that reconcile against a source, actions that produce a diff, rather than perceptual.
Findings
-
The objective in the literature is appropriate reliance, not high trust. Both overreliance and underreliance are failures, and product metrics currently select for the first.
-
Trust is formed on three separable bases, performance, process and purpose, and interventions on one do not substitute for the others.
-
Local explanations and reasoning traces have been tested and increase acceptance of output regardless of its correctness. They should not be counted as calibration mechanisms.
-
Stated accuracy figures are weak and are overwritten by the user's observed experience within a small number of interactions.
-
Cognitive forcing functions have the strongest demonstrated effect on overreliance, and they measurably reduce user satisfaction. The tradeoff is real and should be made explicitly.
-
Pre-announcing the system's failure modes protects trust through a fault, which is the reverse of the usual instinct to present maximal capability at first use.
-
Trust falls faster than it recovers, so provenance and reversibility, which act at the moment of failure, are higher-value investments than any pre-failure messaging.
-
A system that cannot represent an unknown state forces global rather than situational trust, which the framework identifies as the unsafe condition.
-
The generative case is harder than the literature it borrows from, because fluent errors remove the detectability cue those studies assumed.
Limitations
Almost all of the recent evidence comes from short laboratory tasks. Participants are frequently crowdworkers making a series of unfamiliar single decisions with no consequences and no accumulated domain knowledge. Real reliance forms over months of use in a domain the person knows, and effects measured in a forty-minute session may not survive that.
Trust is measured by self-report scales of contested validity. Where behavioural measures are used, they are usually agreement rate, which conflates trust with agreement and with effort avoidance.
The transfer from process control to product design is not established. The foundational studies concern trained operators supervising physical systems with certification regimes. A consumer clicking through a suggestion is in a different situation, and this note assumes the mechanisms transfer while conceding that the magnitudes do not.
Sampling is narrow. The corpus is predominantly North American and European, with student and crowdworker populations. Dispositional trust varies across cultures, and Hoff and Bashir treat that as a substantial factor, so the generalisability of specific effect sizes is limited.
The corpus was selected without a protocol. Papers reporting negative results on explanations were sought deliberately, because they contradict common practice, which means this review may overstate how settled that negative result is. There is active disagreement in the field about when explanations help, and a systematic review would represent that disagreement better than this note does.
The fluency argument is unevidenced. It is a reading of a disanalogy, offered as reasoning. No study cited here tested whether error detectability moderates the effect of explanations, and the claim would need that study to stand.
No mechanism recommended here has been validated in the author's work. Two projects in the author's record are relevant and neither supplies evidence. The Hublix work involved an observation of older users, described in the project documentation, which is a usability observation and not a trust study. Veena Studio's documented interface makes provenance and reversal claims that were never tested with users. The practitioner position here is that of a designer who has read the literature and has never measured reliance in anything shipped.
Implications
For measurement. Log overrides against subsequent ground truth. Calibration is measurable and currently unmeasured, and no design recommendation here can be evaluated without it.
For interface design. Show observed performance in the user's context rather than a global accuracy claim. Describe the competence boundary in the user's domain terms. State what the feature optimises for. Give any inferred state a third value for unknown and render it distinctly.
For failure design. Announce the failure modes before the first use rather than after the first complaint. Build the audit trail and the undo before the explanation panel, since they act where trust actually breaks.
For explanation features. Keep them for debugging and for learning, and remove them from the safety argument. If an explanation is the only control on a consequential decision, the decision is uncontrolled.
For organisations. Expect the supported interventions to reduce satisfaction scores. Decide in advance which metric governs, because the tradeoff will otherwise be resolved silently in favour of the metric that is reported weekly.
For further work. The study this note wants and cannot find is a longitudinal one: reliance calibration in a real product over months, with logged overrides and known ground truth, comparing an explanation-led interface against a forcing-function interface. Everything in the current corpus is a proxy for that.
Conclusion
The literature has been clear for two decades that the goal is calibrated reliance rather than high trust, and product practice has not absorbed it. The mechanisms that dominate shipping products have been tested and increase acceptance without improving accuracy. The mechanisms with better support impose friction and reduce the metrics teams report.
The strongest available recommendations are architectural rather than presentational. Represent uncertainty as a real state. Record what the system did in a form a user can inspect. Make the action reversible. Announce the failure modes early, so that the first fault confirms an expectation instead of destroying one.
And the generative case is harder than its source literature, because a fluent error looks exactly like a correct answer. That removes the perceptual cue those studies relied on, and it means verification has to be built into the product rather than left to the user's attention.
References
- J. D. Lee and K. A. See, "Trust in Automation: Designing for Appropriate Reliance", Human Factors 46(1), 2004, 50–80.
- B. M. Muir and N. Moray, "Trust in automation. Part II: Experimental studies of trust and human intervention in a process control simulation", Ergonomics 39(3), 1996, 429–460.
- R. Parasuraman and V. Riley, "Humans and Automation: Use, Misuse, Disuse, Abuse", Human Factors 39(2), 1997, 230–253.
- K. A. Hoff and M. Bashir, "Trust in Automation: Integrating Empirical Evidence on Factors That Influence Trust", Human Factors 57(3), 2015, 407–434.
- M. T. Dzindolet, S. A. Peterson, R. A. Pomranky, L. G. Pierce and H. P. Beck, "The role of trust in automation reliance", International Journal of Human-Computer Studies 58(6), 2003, 697–718.
- M. Yin, J. Wortman Vaughan and H. Wallach, "Understanding the Effect of Accuracy on Trust in Machine Learning Models", CHI 2019.
- Y. Zhang, Q. V. Liao and R. K. E. Bellamy, "Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making", FAT 2020*.
- G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro and D. S. Weld, "Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance", CHI 2021.
- Z. Buçinca, M. B. Malaya and K. Z. Gajos, "To Trust or to Think: Cognitive Forcing Interventions Reduce Overreliance on AI in AI-assisted Decision-making", Proceedings of the ACM on Human-Computer Interaction 5(CSCW1), 2021.
- R. F. Kizilcec, "How Much Information? Effects of Transparency on Trust in an Algorithmic Interface", CHI 2016.
- S. Amershi, D. Weld, M. Vorvoreanu, A. Fourney, B. Nushi, P. Collisson, J. Suh, S. Iqbal, P. N. Bennett, K. Inkpen, J. Teevan, R. Kikin-Gil and E. Horvitz, "Guidelines for Human-AI Interaction", CHI 2019.
- Google, People + AI Guidebook. Practitioner guidance, not research. https://pair.withgoogle.com/guidebook/
- Hublix usability observation with users aged over fifty, and the three-state permission model. Recorded as B12 and B13 in
SOURCE_INVENTORY.md.
Related reading in this archive
- Designing with AI in the interface: the same problem stated as four guarantees
- Agents, autonomy and the human in the loop
- Latency, trust and the two variables everyone conflates
- Hublix: the three-state problem in a home automation interface
- Veena Studio: provenance and reversibility in a creative tool