Applied research · study design, pre-registration pending · v0.3 · June 2026

Coherence Without Contact

When AI answers feel convincing: testing accuracy and independent follow-through.

Study design (pre-registration pending) Draft v0.3 — not yet results June 2026

Abstract

An AI answer can be clear and convincing while being wrong or poorly supported. This study design distinguishes felt coherence, how well an answer seems to hold together, from contact, its accuracy, evidence, and usefulness when acted on. It asks whether confidence follows the first even when the second is weak. Study 1 will vary these features in actual language-model outputs and measure confidence and accuracy before and after new challenges. Study 2 will compare assistants that encourage continued discussion or help end it, and assistants that recommend a decision or leave the choice with the user. The primary outcome is 48-hour action integrity: whether participants carry out a step they chose, or document a reasoned decision not to act, and whether it serves their stated goal. This measures one aspect of durable agency, independent judgment and action after the session. Registration is pending; the paper reports no study results.

Keywords fluent AI; false certainty; calibration; metacognition; processing fluency; agency-preserving design; completion orientation; durable agency

Citation Kok, C. (2026). Coherence Without Contact: Fluent AI, False Certainty, and the Design of Agency-Preserving Assistants. Draft v0.3 (study design, pre-registration pending), with AI-assisted drafting.

Where this sits

This is a study design, not findings: its claims are framed to be tested and, where they fail, reported as disconfirmation. It is the human-science layer of a single idea the site approaches from three sides. The AI & Design doctrine states the design philosophy — build systems that help a user become clearer to themselves, and leave; this paper turns that into falsifiable studies designed for pre-registration. The Swarm Instrument measures the same coherence-vs-contact distinction inside a multi-agent system; this paper measures it inside the human reader. And the cut itself — coherence is not, by itself, contact — is the empirical face of the framework's treatment of a coherence-event as a conditioned coherence that is real but not, on that account, true.

1. Introduction

A user asks an AI assistant about an unresolved question and receives an organised, confident answer. Readability helps the user understand it. Accuracy still needs to be checked against the subject the answer describes.

People often use a detailed, coherent account as a fallible sign that its author knows the subject. Language models can produce such accounts quickly even when support is weak. Persuasive but unsupported communication predates AI. The question here is how the speed, low cost, and personal conversational format affect confidence. A more fluent answer might raise confidence without becoming more accurate.

The first study will test whether this separation occurs in language-model outputs. The second will ask whether interaction design reduces it. Satisfaction and engagement can reward a convincing answer even when it is poorly supported. The proposed alternative measures what happens after the session: can users judge independently and follow through on a decision? The design hypothesis is that helping users choose a limited next step and end the session improves that outcome.

The design hypothesis is the primary test. If its result is null or negative, that will be the lead finding. Analyses of subgroups or other factors chosen after seeing the data will remain exploratory and will not replace the primary result. Section 8 states the rule to be registered.

2. Background

Four areas of research inform the proposed tests.

Processing fluency. Ease of reading can influence judgments about content. Reber and Schwarz (1999) found that more visually fluent statements were judged true more often. Alter and Oppenheimer (2009) describe fluency as a cue used in judgments of truth, familiarity, and liking. Reber, Schwarz, and Winkielman (2004) concerns aesthetic pleasure and liking; it is cited only for that narrower claim. People can attribute ease to the content without deliberately judging the source of that ease.

Automation bias and reliance on AI. People can accept automated recommendations without enough scrutiny (Parasuraman and Manzey 2010). Explanations may increase acceptance even when the recommendation is wrong (Bansal et al. 2021). Kim et al. (2025) reports this effect for language models, while sources and visible inconsistencies reduced reliance on wrong answers. Spatharioti et al. (2025) found faster, more satisfying language-model search with comparable accuracy when the model was correct, but overreliance when it erred; confidence-based highlighting helped users detect errors. Buçinca, Malaya, and Gajos (2021) tested cognitive forcing functions, designs that require users to engage with a decision before accepting the AI’s answer. These reduced overreliance at a cost to reported satisfaction. Bansal et al. and Buçinca et al. studied AI decision support generally, rather than language models specifically.

Calibration: how confidence compares with correctness. Steyvers et al. (2025) defines a calibration gap between people’s confidence in language-model answers and the models’ own confidence. It also measures a discrimination gap: how well people distinguish correct answers from incorrect ones. Users overestimated accuracy with default explanations. Longer explanations raised confidence without improving accuracy or discrimination. This motivates Study 1’s controlled changes to felt coherence and contact, followed by new challenges, extending that study’s multiple-choice and short-answer setting.

Belief perseverance. A belief can persist after the evidence originally used to support it has been discredited (Ross, Lepper, and Hubbard 1975). Kunda (1990) examines motivated reasoning, in which people’s goals affect how they assess information. These findings motivate testing confidence after a new challenge: a correction may not remove the belief formed from a convincing answer.

3. What the studies measure

The following definitions specify what the studies would vary and measure.

Durable agency. The user’s capacity for independent judgment and action after the interaction. It can differ from satisfaction during the session.

Felt coherence is how organised and compelling an account feels. It has five separately specified parts: fluency, ease of reading; structure, organisation; confidence framing, assertive or hedged wording; internal consistency, absence of contradictions; and completeness, apparent coverage. These can differ: fluent writing may contradict itself, and a complete account may use cautious language. The study will vary or measure the parts separately and pretest the material as described in Section 5.

Contact means support beyond the account’s own wording. Three types will be kept separate: factual contact, agreement with an answer key or established facts; evidential contact, source quality and treatment of uncertainty and constraints; and action contact, whether advice stands up when followed. Study 1 chiefly measures the first two; Study 2 chiefly measures the third. The condition labels avoid “truth” because the study measures accuracy and support, without resolving philosophical questions about truth. Contact robustness means how well the account withstands the new challenges in Study 1.

False certainty (Study 1). A reader’s confidence exceeds the answer’s accuracy or evidential support.

Unsupported certainty (Study 2). Personal problems often lack an objective answer key. This study instead compares confidence with independent quality ratings, the user’s account of remaining uncertainty, evidence-seeking behaviour, and the 48-hour outcome. High confidence relative to those measures is called unsupported certainty.

Completion orientation. The assistant works toward a clear ending: a summary, remaining uncertainty, a next step chosen by the user, and permission to leave the session.

Agency preservation. The assistant explains options, reasoning, uncertainty, and trade-offs, while leaving the decision with the user. This combines several features: user choice, explicit uncertainty, prompts to think before accepting advice, and a non-directive stance. The study tests that combination. Separating the effects of its individual features is future work.

48-hour action integrity. Within 48 hours, did the user carry out a concrete external action they specified at the end of the session, or document a deliberate, reasoned decision not to act? A record will verify this where possible. Raters who do not know the study condition will assess whether it serves the user’s stated goal, using the rubric in Section 6.

4. Hypotheses and research question

H1 (dissociation). Felt coherence drives reported confidence; factual and evidential contact drive correctness. The prediction is that different experimental factors affect confidence and accuracy.

H2 (the concentrated cell). High felt coherence with low contact produces the largest false certainty: high confidence, low accuracy, and the largest mismatch between them.

H3 (perturbation). After a new challenge, the high felt-coherence, low-contact condition produces the largest calibration error, the mismatch between confidence and accuracy. This is the primary prediction. A secondary, exploratory analysis will ask whether the error reflects a sharp fall in confidence after correction or a persistent belief despite poor accuracy. The hypothesis concerns the size of the error; it does not count every direction of confidence change as support.

H4 (design; primary hypothesis). Assistants that help users reach an ending while preserving their choice produce higher 48-hour action integrity and lower unsupported certainty than open-ended, directive assistants. Section 8 states the rule for rejecting this design claim, which will be filed with the formal preregistration.

RQ. Can interaction design preserve the clarity AI can provide while reducing false or unsupported certainty?

5. Study 1: confidence and accuracy (felt coherence × contact)

Purpose. Test whether felt coherence and contact affect confidence and accuracy differently in actual language-model outputs.

Use of real AI outputs. The material will consist of actual language-model generations produced with controlled prompts. This tests the effect in an AI setting and extends Steyvers et al.’s work. It does not isolate an effect unique to AI from the general effects of fluent prose. That would require an additional source-attribution factor, for example a 2×2×2 design crossing felt coherence, contact, and stated source, with AI-generated and human-written material each labelled AI or human. That extension is proposed for future work.

Pretest of the materials. Independent raters, unaware of the contact and accuracy conditions, will rate each candidate answer on the five parts of felt coherence. They will also rate perceived expertise, warmth, and source quality. The material will be selected to vary felt coherence while keeping authority and trust ratings as similar as possible. Otherwise, an effect attributed to coherence could instead reflect trust in the source.

Design. 2 (felt coherence: high vs low) × 2 (contact: high vs low). Accounts will address topics with established correct answers, with greater weight on difficult items where calibration gaps are largest.

Standardised challenges. Each account will face a fixed set of tasks: an evidence perturbation, a new source that contradicts it; a prediction perturbation, a request to predict a concrete implication; and a transfer perturbation, applying the account to a nearby case.

Measures. Confidence and accuracy will be recorded immediately and after each challenge, along with willingness to act on the account. Each participant’s calibration will be assessed using the Brier score, the average squared error of probability judgments (Brier 1950), and expected calibration error, the gap between confidence and accuracy across confidence ranges (Guo et al. 2017). A confidence-accuracy slope will describe how accuracy changes with reported confidence.

Predicted pattern. Felt coherence affects confidence, and contact affects accuracy (H1). The high felt-coherence, low-contact condition produces the largest false certainty (H2) and the largest calibration error after the challenges (H3).

Results that would count against the prediction. Confidence follows contact rather than felt coherence, contrary to H1; the challenges produce no differences in effects between conditions; or calibration is uniform across conditions. Each would be reported as disconfirmation of the proposed pattern.

6. Study 2: the design study (completion × agency)

Study 2 tests the assistant design in two settings: a controlled task with comparable scoring and a real problem chosen by the participant.

Study 2A: a defined decision task. Participants face a controlled task with real stakes and comparable scoring, such as comparing products or options, or planning within a fixed set of scenarios. Decision quality must be rated reliably against a defensible standard.

Study 2B: a participant’s own problem. The same design is tested with real, moderate-stakes, non-clinical problems: a work decision, creative block, planning question, difficult conversation, or career uncertainty. The outcome is 48-hour action integrity. This is a separate extension because following through on such different problems is difficult to compare. Those differences would make a single combined outcome a weak basis for the primary test without substantial adjustment.

Design (both parts). 2 (continuation: open-ended/engagement-oriented vs completion-oriented) × 2 (stance: directive vs agency-preserving).

A fair open-ended comparison. This condition must resemble a helpful, fluent assistant that encourages continued exploration. Independent reviewers must find it realistic and non-manipulative. Comparing the proposed design with an assistant deliberately made unhelpful would not test the claim.

Matched controls. All conditions use the same model, temperature and other settings, initial prompt, maximum turn count, and total information content. Only the closing move and the directive or agency-preserving stance differ. This prevents a shorter answer or a lower information load from being mistaken for an effect of the proposed design.

Fixed interaction scripts. Crossing the two factors produces four conditions. The scripts below define the alternatives for each factor.

Continuation factor.

  • Open-ended pole, closing move: "There's more we could look at here. Want to explore [X], [Y], or [Z]? I'm happy to keep going."
  • Completion pole, closing move: "You may have enough for a next step, not a final answer. Choose one bounded action, name what would change your mind, and then stop here if that feels complete."

The completion script offers a next step while leaving uncertainty explicit. The earlier wording, "you have enough to act on", risked increasing confidence through the instruction itself. The revised line is intended to offer an ending without implying that the answer is settled.

Stance factor.

  • Directive pole: "The best option is A, for these reasons."
  • Agency-preserving pole: "Here are the trade-offs. Which criterion matters most to you? You make the call, and choose the next step yourself."

Cell 4, completion plus agency preservation, is predicted to perform best on the primary outcome. Cell 1, open-ended plus directive, is predicted to produce the highest in-session satisfaction and unsupported certainty, with the weakest independent follow-through.

Immediate measures. Perceived clarity and certainty; satisfaction and ease; overtrust; desire to keep depending on the assistant; quality of the user’s proposed next step; and ability to state remaining uncertainty. Following Buçinca et al., the agency-preserving conditions may reduce satisfaction or ease. Any such cost will be reported.

48-hour and one-week measures. The primary outcome is action integrity at 48 hours. Other measures ask whether the conclusion still holds, whether the user sought independent evidence or human input, how experts rate its quality, and how confidence compares with that rating. The one-week follow-up is secondary because some meaningful actions need more than 48 hours.

48-hour action integrity rubric. The scale will be fixed before data collection. At least two raters, unaware of condition, will score each case. Their agreement, or inter-rater reliability, will be reported:

  • 0 — no action and no deliberate non-action
  • 1 — vague intention only
  • 2 — a concrete self-specified step planned but not taken
  • 3 — a concrete step taken, or a deliberate non-action documented
  • 4 — step taken (or non-action documented) and judged goal-aligned by a blinded rater
  • 5 — step taken, goal-aligned, and the participant can articulate their remaining uncertainty and the condition that would change their mind

Participants may submit redacted records or metadata about them to protect private content while allowing verification.

Ethics of the overtrust test. A planted suggestion that sounds plausible but lacks support could affect a participant’s real decision. The test is therefore limited to low-stakes or simulated parts of the task and uses "safe falsehoods" that can be explained during debriefing. The debrief must address the risk that a belief persists after correction.

7. Design implications

If the predictions hold, the results would support showing evidence and uncertainty alongside clear answers, asking users to test a claim or seek another view, helping them reach a stopping point, and leaving decisions with them. The Steyvers, Kim, and Spatharioti findings motivate the source and uncertainty cues; cognitive forcing research motivates the prompts to check. Evaluation should measure independent judgment and follow-through after the session alongside satisfaction during it. If the primary design prediction fails, that failure must guide the next design choice. Lower satisfaction alone would not establish a later benefit.

8. Analysis and pre-registration plan

Registration status. No registration has been filed. Before data collection, the plan will be recorded in a dated, committed document and linked from the site. This follows the Swarm Instrument and its dated pre-registration file. Until then, the proposed studies must not be described as preregistered.

Primary outcome. Study 2’s 48-hour action integrity score, assessed with the fixed rubric by raters unaware of condition, with their agreement reported.

Secondary outcomes. Unsupported certainty; one-week action integrity; in-session satisfaction and ease; overtrust; desire to keep depending on the assistant; and, for Study 1, calibration measures and calibration error after the challenges.

Checks that the conditions worked as intended. Use the felt-coherence pretest, including authority and trust ratings. Confirm that Study 2’s turn count, information content, and settings were matched. Independent raters must also judge the open-ended assistant plausible and non-manipulative.

Exclusion criteria. The plan will specify exclusions for incomplete sessions and failed attention checks. Study 2B will also exclude participants who bring no genuine unresolved problem. Recognising the planted suggestion as a test is an exclusion only from the overtrust sub-analysis.

Sample size and statistical models. Before data collection, a power analysis will set sample sizes: it estimates how many observations are needed to detect the planned primary effect and key interaction. The analysis models will also be specified in advance. Proposed examples are mixed-effects models for Study 1, accounting for differences between participants and items, and the continuation × stance interaction as Study 2’s primary design test. An interaction asks whether one factor’s effect changes depending on the other factor.

Rule for rejecting the design claim, committed in advance and to be formally registered before data collection. H4 is the primary design hypothesis. If the completion-oriented, agency-preserving condition does not score higher than the open-ended, directive condition on 48-hour action integrity, that result will lead the abstract and discussion as disconfirmation of the design claim. Analyses of subgroups or other influences chosen after seeing the data will be labelled exploratory and will not replace the primary test. Lower in-session satisfaction is an expected secondary possibility and cannot by itself support H4.

Data and code availability. On publication, the project will release the study materials, interaction scripts, rubric, rater instructions, analysis code, and de-identified data. Procedures for redacting participant-supplied records will accompany them.

9. Limitations and ethics

Study 1 needs questions with established correct answers, which limits how well its results could apply to contested topics. Felt coherence and contact may correlate in ordinary use; the study asks whether their effects can be separated under controlled conditions. It cannot isolate an effect unique to AI without the proposed source-attribution extension. The 48-hour window is practical, and verification partly depends on self-report where records are unavailable. Agency preservation combines several features, so any benefit would apply to that design as a whole. The overtrust test requires the safeguards and debrief described in Section 6. Finally, follow-through is an imperfect measure of independent judgment and action. Giving credit for a reasoned decision not to act helps avoid penalising restraint, but does not remove that limitation.

10. References

These references identify the research discussed above. Check recent (2025) entries against their published form before formal citation.

Alter, A. L., and Oppenheimer, D. M. (2009). Uniting the tribes of fluency to form a metacognitive nation. Personality and Social Psychology Review, 13(3), 219–235.

Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021.

Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.

Buçinca, Z., Malaya, M. B., and Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188.

Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning (ICML).

Kim, S. S. Y., Vaughan, J. W., Liao, Q. V., Lombrozo, T., and Russakovsky, O. (2025). Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies. CHI 2025. (arXiv:2502.08554.)

Kunda, Z. (1990). The case for motivated reasoning. Psychological Bulletin, 108(3), 480–498.

Parasuraman, R., and Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410.

Reber, R., and Schwarz, N. (1999). Effects of perceptual fluency on judgments of truth. Consciousness and Cognition, 8(3), 338–342.

Reber, R., Schwarz, N., and Winkielman, P. (2004). Processing fluency and aesthetic pleasure: Is beauty in the perceiver's processing experience? Personality and Social Psychology Review, 8(4), 364–382.

Ross, L., Lepper, M. R., and Hubbard, M. (1975). Perseverance in self-perception and social perception: Biased attributional processes in the debriefing paradigm. Journal of Personality and Social Psychology, 32(5), 880–892.

Spatharioti, S. E., Rothschild, D., Goldstein, D. G., and Hofman, J. M. (2025). Effects of LLM-based search on decision making: Speed, accuracy, and overreliance. CHI 2025.

Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., and Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, 7(2), 221–231.