Does Your Dashboard Make Skipping Checks Look Normal?
In collaborative AI workflows, effective verification may depend not only on how much checking capacity exists but on whether people choose to invoke it; interfaces that expose peer behavior can supply social signals, although the verification-specific effect has not yet been demonstrated in a human workplace.
In a recent computer simulation, 400 agents repeatedly had three choices. They could solve a task themselves, accept an AI answer, or pay a cost to check the answer first.
At one point, the researcher changed a single kind of incentive. Seeing neighbors accept AI output without checking it could now make unchecked acceptance more attractive. The AI did not become more accurate. The agents did not become less capable. The model simply began rewarding a visible peer action.
The effect was large. As the strength of that peer-behavior term increased, the simulated verification rate fell from about 29 percent to 0.2 percent. A severity-weighted measure of unchecked reliance rose from about 30 percent to 52 percent. Those figures come from Ahana Biswas’s agent-based model of AI reliance, submitted on August 20 and listed by arXiv as accepted at CSS 2026.
The tempting conclusion is that checking is contagious in reverse: watch enough colleagues skip it, and eventually you will skip it too.
The paper does not establish that.
Its social effect is an assumption built into the model. The model asks what would happen if visible unchecked use carries a meaningful social payoff. Its parameters are not calibrated to a workplace, and its main heterogeneous simulation changes smoothly rather than displaying the abrupt tipping or hysteresis that appears in some of the paper’s mathematical cases. Even the 52 percent figure is easy to misread: it is not the percentage of agents making mistakes, but a weighted proxy for unverified use when the AI result is imperfect.
What the simulation does establish is narrower and, for workflow design, more interesting. A review system can have plenty of checking ability sitting inside it while the amount of checking actually performed changes for another reason: people may decide not to use the capacity they have.
That turns a familiar engineering question into a behavioral one.
Capacity is not the same as use
When a team worries that AI-generated work is moving faster than humans can inspect it, the obvious quantities are physical. How many reviewers are available? How long does each check take? Which outputs deserve the most scrutiny?
Those are real constraints. They are not the only ones.
Research on human-AI reliance commonly records an individual decision: accept the system’s answer, reject it, or seek more evidence. A 2026 systematic review of 67 HCI papers found a fragmented mix of behavioral, self-report, performance, and process measures, often built around that person-level choice. The broader trust literature has long included organizational and social context, so the point is not that researchers forgot that people have coworkers. It is that the act of checking is usually something one person eventually does or does not do.
A person can decline to check because checking is expensive, because the AI appears reliable, because a deadline is close, because a policy permits it, or because an interface makes the extra work awkward. A randomized study published this month offers a useful rival to any social explanation. Among 200 older Chinese adults shown static AI-chat screenshots, a bundle of source labels, uncertainty language, and a verification prompt increased reported verification intention even though no peer behavior was shown. The behavioral proxy—opening optional source information—also moved upward, but its confidence interval included no effect.
In other words, an interface can change checking without a social mechanism at all.
Biswas’s model asks what happens when collaboration adds one more input: what other people appear to be doing.
A colleague’s action can mean two different things
Suppose you see that a coworker used an AI assistant and the resulting code passed every test. You have learned something about the tool. The colleague’s action came with evidence about the outcome.
Now suppose you see only that twenty coworkers accepted AI suggestions today. You have learned something different. Perhaps the tool is good. Or perhaps you have learned what people around here apparently do.
Those two interpretations are easy to blend together. They are not the same mechanism.
The first is ordinary learning from other people’s experience. In Biswas’s model, when peers exchange comparable information about AI quality, their beliefs can become more similar without necessarily shifting the group’s average reliance. The paper also shows an important boundary: if influential people systematically transmit different beliefs, information alone can move the group. A network does not need social pressure to create a collective effect.
The second interpretation is the paper’s more speculative one. Visible unchecked use changes the attractiveness of unchecked use directly. The action functions as a norm rather than merely as evidence.
Do real people treat visible AI use that way? There is evidence for the broader social channel, though not yet for checking itself.
A study published online in the European Management Journal in July used a survey of 299 employees, a controlled online experiment with 150 participants, and a longitudinal field study involving 74 teams and 241 employees. In the experiment, seeing high coworker adoption of generative AI produced both positive and negative expectations even when direct performance information was withheld. In the field study, positive expectations were associated with sustained later adoption, and a more competitive team climate amplified the formation of both positive and negative expectations.
That is useful evidence that visible coworker AI use can become part of another employee’s decision environment. It is not evidence that seeing coworkers skip verification causes someone else to skip it. Adoption and checking are different behaviors. One makes use of a tool more common; the other spends time questioning its output.
The missing distinction matters because a shared work interface can show both behaviors very unevenly.
What the interface leaves behind
Completion is easy to display. A card closes. A pull request is approved. A report is published. An AI suggestion is accepted. These events create durable status changes.
Verification often disappears into the process that produced them. Someone opened the source. Reran the calculation. Tested an edge case. Compared another model. Asked a colleague. Rejected a suspicious sentence and tried again. Unless the system deliberately records those actions, the final interface may preserve the green check and erase almost everything that made the green check trustworthy.
That creates a simple asymmetry: the behavior easiest for colleagues to observe may be acceptance, not checking.
If people sometimes use visible peer behavior as a social cue, then a dashboard is doing two jobs at once. It reports the past to a manager. It may also supply information about what looks ordinary to everyone who sees it next.
This is an inference from the evidence, not an established workplace effect. The current studies do not show that an acceptance-heavy dashboard causes verification to fall. They do show enough to make neutrality a hypothesis rather than a default: coworker AI behavior can influence later AI adoption, non-social interface changes can alter verification intentions, and a formal model demonstrates how a peer-behavior term could make checking frequency feed back on itself.
That also means the obvious fix—make checking visible—has not been earned.
A visible norm can point the wrong way
Social signals are not automatically helpful simply because the desired behavior is admirable.
A classic field experiment on household electricity use makes the problem concrete. Households were shown how their energy consumption compared with the neighborhood average. High-use households cut consumption. Low-use households, however, increased it. The descriptive norm had inadvertently told efficient households that most people used more power. Adding a small sign of social approval for low consumption eliminated that boomerang effect.
An AI-review dashboard is not an electric bill. The incentives, stakes, and reference groups are different. But the experiment exposes a boundary that carries over cleanly: displaying what most people do can normalize the wrong behavior if the underlying norm is bad.
A system that announces that only a small minority of colleagues verified an AI answer might encourage more checking. It might also advertise that almost everyone gets away without it. A “checked” badge might signal care. It might also turn checking into the cheapest action that produces a badge.
The outcome that matters is not the number of check marks. It is whether errors are found, reliance becomes better calibrated, and accepted work survives downstream use.
That is why the most important experiment has not yet been run.
The experiment the dashboard needs
Keep the tasks and AI outputs the same. Change only what participants can see about peers.
One group sees that peers accepted outputs. Another sees that peers verified them. A third sees no peer actions. Then separately vary whether the results of those peer actions are visible.
That last split is crucial. If someone checks more after learning that a colleague caught an error, ordinary learning about AI quality could explain the change. If someone checks more after seeing the act of verification even when its outcome is hidden, the case for a normative effect becomes stronger.
Run the sequence repeatedly. Measure actual checking, not only stated intention. Count errors detected, time spent, corrections made, and the quality of the final work. Vary expertise, task difficulty, accountability, and team culture. A well-designed study could then separate peer norms from several rivals that already have good reasons to matter: model quality, check cost, direct interface friction, policy, and informational learning.
Until that experiment exists, “checking is contagious” is a hypothesis.
The more defensible conclusion is quieter. A team can possess review capacity that it does not invoke. In a collaborative system, the traces left by earlier work may become one of the inputs to the next person’s decision about whether to use that capacity.
Imagine the row in a work queue after an AI-produced task is approved. The green check remains. The source tabs, rerun tests, discarded drafts, and second opinions usually do not. A manager sees status. A colleague may also see a norm.
We do not yet know how much that second signal changes behavior. But the row is no longer safe to treat as only a record of the past.
Sources
- Ahana Biswas, “Modeling AI Overreliance as a Complex Adaptive System”, arXiv, submitted August 20, 2026; listed as accepted at CSS 2026.
- Boxiang Yu, Lu Wu, Tingting Li, Yifang Lin, and Yimo Shen, “Visible Hands in Technology Diffusion: Colleague GAI Adoption as Dual Cognitive Signal in Competitive Workplace Ecosystems”, European Management Journal, available online July 23, 2026.
- Mohammad Naiseh, Deniz Cemiloglu, Huseyin Dogan, and Nan Jiang, “How is Reliance on AI Measured? Mapping and Evaluating Measurement Approaches in Human-AI Interaction”, SSRN working paper, posted May 29, 2026; revised July 10, 2026.
- Natalie C. Benda et al., “Trust in AI: why we should be designing for APPROPRIATE reliance”, Journal of the American Medical Informatics Association, 2022 issue; published online November 2, 2021.
- Jun’an Yu et al., “Effects of a Safety User Interface Bundle on Verification Intentions in Generative AI Chat Use Among Older Chinese Adults: Randomized Vignette Survey”, Journal of Medical Internet Research, August 14, 2026.
- P. Wesley Schultz, Jessica M. Nolan, Robert B. Cialdini, Noah J. Goldstein, and Vladas Griskevicius, “The Constructive, Destructive, and Reconstructive Power of Social Norms”, Psychological Science, May 2007.