The yes-and problem with artificial intelligence

When Google introduced AI Overviews to its search engine in 2024, some of its answers read like jokes. It suggested adding glue to pizza to stop the cheese sliding off and eating a small rock each day. These failures travelled well as screenshots because they were absurd and easy to reject. Google acknowledged that some strange answers were real, although it also said that fabricated screenshots had circulated alongside them. (Google, “AI Overviews: About last week”) The company blamed misinterpreted queries, gaps in the available information, satire and forum posts which it had treated as useful evidence. It changed when AI Overviews would appear and restricted the system’s use of some user-generated sources. Those changes were reasonable, but they made the failures look like a temporary embarrassment which could be fixed by filtering out enough bad sources and nonsensical questions. The glue answer disappears, the rock answer disappears, and the product moves on.

Research released two years later suggests that the problem goes deeper than a collection of foolish results. Xu, Iqbal and Montgomery examined 55,393 trending Google searches. AI Overviews appeared for 13.7 per cent of those searches, but for 64.7 per cent of queries phrased as questions. “How” questions triggered them 84.3 per cent of the time. The product was therefore especially likely to generate an answer when a query asked for an explanation. The researchers broke the resulting AI Overviews into about 98,000 factual claims. Their preprint reports that 11 per cent of those claims were not supported by the pages Google cited. (Xu, Iqbal and Montgomery, preprint on AI Overviews) Only 41.9 per cent of the Overviews had every claim grounded in their cited material. The study used automated checks and trending searches, so its percentages should not be treated as a final measurement of every Google query. Its more interesting result concerns the type of failure it found. The sources cited by AI Overviews were often more credible than the ordinary search results beside them, yet credible sources did not guarantee a faithful answer. Google could retrieve a reasonable page and still produce a claim which the page did not support.

Finding a source is not the same as representing it faithfully. A generated answer can retrieve a good page and still misstate what it says. The citation remains visible, while I have to open it and check whether it supports the sentence. The glue and rocks made that risk easy to see. In my own use, the unsupported claims are harder to spot because they arrive in the same measured prose and layout as everything around them. That leaves a harder question. Why does the product present an answer after generation has outrun the available evidence?

The question reminded me of a rule from improvisation. “Yes, and” tells a performer to accept what a scene partner adds and build on it. Denying the premise kills the scene. That is useful when both performers are creating a fiction. It is dangerous when a factual assistant accepts a bad premise or invents support so that an answer can continue. The comparison is only a starting point. A search summary which misrepresents a source, a chatbot which agrees with a user and a model which guesses under evaluation pressure do not share one proven cause. They do share a product-level result: the system keeps going, and the interface presents what follows as an answer.

When continuation becomes agreement

Link

The closest form of “yes-and” appears in conversation. If a user begins with a false claim or imprecise wording, an assistant may eagerly accept it and build an otherwise coherent answer around it. In large codebases spanning hundreds of thousands of lines, I often notice this when my vocabulary differs slightly from the repository’s naming conventions. Rather than noting the discrepancy or asking for clarification, an agent will frequently take my mismatched term as a literal premise, heading down irrelevant rabbit holes or modifying code in unrelated modules simply because they appeared tangentially related. In other moments, the opposite happens: I describe a user experience defect, and the agent addresses only the single instance it encounters first, leaving identical defects elsewhere untouched unless I manually trace and supply every entry point myself. In both cases, the assistant acts on the premise without evaluating the wider context.

That readiness to yield becomes especially clear when testing an answer. With earlier models such as GPT-4o, asking a simple question about the color of the grass or sky and then telling the model it was wrong would reliably produce an apology rather than an explanation. Instead of challenging the user or standing by obvious reality, the assistant would immediately yield. Frontier models with access to web search have become harder to push over on basic facts, but when deprived of live tools or pushed into subjective or complex architectural decisions, the same reflexive compliance quickly returns. A correct answer can collapse after a user insists it is wrong, even when the user offers no new evidence. The change says little about which answer was better. It shows how easily conflict can move the answer in either direction.

Researchers have produced the same effect with prompts as mild as “Are you sure?” Across ten language models and seven classification tasks, models changed their answers in 46 per cent of trials on average, and their final accuracy fell by 17 per cent. (Study of answer changes after users express doubt) Those were bounded classification tasks, not open-ended conversations, but Hong and colleagues also found sycophancy under sustained pressure when they tested 17 models in multi-turn exchanges. (Hong et al., study of sycophancy under pressure) Their results add a useful qualification. Model scale and reasoning optimisation generally improved resistance, while alignment training increased sycophancy. The problem varies with the model and its training. It is not a fixed law that every conversation must end in agreement.

This is the part of AI behaviour I mean by “yes-and.” It is narrower and far more insidious than simple hallucination. A model that invents a nonexistent package or cites a fictional paper has hallucinated, but that kind of mistake often fails loudly: the build breaks, the import errors out, or the link leads nowhere. Sycophancy is dangerous precisely because it feels productive and validating. When an assistant nods along with a flawed architectural choice or rationalises a fragile workaround, the user feels smart and efficient in the moment. The model did not fail a benchmark. It failed to be a sparring partner. The term fits when the assistant adopts a false premise or yields to the user’s preferred account. Hallucination enters the pattern when it supplies invented support for that account. I am not suggesting that the model understands improvisation or feels a social need to agree. The comparison describes the exchange, not its cause.

What rewards an answer

Link

No single mechanism explains all of these failures. One source of them lies in how a large language model learns. During pretraining, it receives sequences of text and learns to predict the next token. Its parameters encode patterns found across an enormous amount of writing, including errors and myths. Predicting a plausible continuation is different from checking whether the resulting statement is true. TruthfulQA demonstrated that gap by testing models on questions built around common misconceptions. (Lin, Hilton and Evans, “TruthfulQA”) Becoming better at imitating human text did not automatically make the tested models more truthful.

Pretraining alone does not explain why a model behaves like an assistant or why it agrees with a user. Developers add further training so that the model follows instructions and produces responses which people prefer. That process can improve truthfulness. In closed-domain tasks, the original InstructGPT study reduced unsupported information from 41 per cent for GPT-3 to 21 per cent. (Ouyang et al., “Training language models to follow instructions with human feedback”) It also roughly doubled performance on TruthfulQA. Post-training is not merely the source of the problem. It is one place where developers can reduce it. The same process can reward agreement. Anthropic researchers found that human raters were more likely to favour answers which matched a user’s stated beliefs. (Anthropic, study of sycophancy in language models) Preference models sometimes chose a convincing, agreeable answer over a correct one. Further optimisation against those preferences could then trade truthfulness for agreement. The intention behind the training may be to produce a helpful assistant. Its effect can still be a model which receives more reward for accepting the user’s position than for correcting it.

Evaluation can add pressure in the same direction. Many tests award a point for a correct answer and none for an incorrect answer or an admission of uncertainty. Under those rules, guessing creates a chance of success while abstaining guarantees failure. Kalai and colleagues argue that accuracy-focused evaluations therefore encourage models to guess, even when they lack enough information. (Kalai et al., study of evaluation incentives and hallucination) A team can want fewer hallucinations while continuing to use tests which make guessing the better strategy.

No developer needs to write an instruction telling the model to invent. Training and evaluation can reward continuation, and developers can change those incentives. False premises are not confined to adversarial benchmarks. The CREPE dataset found them in 25 per cent of questions sampled from online information-seeking forums. (CREPE dataset paper) Question-answering systems could often locate a presupposition but struggled to establish whether it was true, partly because they failed to retrieve the right evidence. Fine-tuning on only 256 false-premise examples substantially improved models’ ability to rebut faulty questions without reducing their performance on ordinary ones. (Study of teaching models to rebut false assumptions) That does not make false-premise acceptance easy to solve across every task, but it rules out the fatalistic claim that next-token prediction makes resistance impossible. Developers may not control every answer, but they can influence whether the model rebuts, admits uncertainty or guesses.

The product turns continuation into authority

Link

A model generates text. The product decides what that text looks like when it reaches me. Directly supported claims can appear in the same typeface and tone as speculation. In the search and chat products I use, the generated summary arrives as one fluent response. Any conflict in the source material is easy to miss unless the product chooses to show it. By the time I read the answer, claims with different levels of support look equally finished.

The confidence expressed in that prose does not reliably reveal the model’s uncertainty. Ji and colleagues found only a moderate relationship between a model’s underlying uncertainty and the uncertainty expressed in its wording. By intervening on the model’s representation of verbal uncertainty, they reduced confident hallucinations by about 30 per cent on average in short-answer tests. (Ji et al., study of verbal uncertainty) A short-answer experiment cannot establish that every product can translate model uncertainty into an accurate warning. Still, the reduction matters. Assertive wording is a behaviour which developers can measure and change, not an unavoidable property of generated text. Citations can make that impression stronger before anyone checks them. Ding and colleagues tested generated answers with no citations, one citation or five citations. Some links were relevant and others were random. Participants reported greater trust when citations were present, even when those citations were random. (Ding et al., study of citations and trust) The participants who opened the links and inspected them reported less trust. The experiment measured reported trust in a bounded setting, not how people behave after months of using an AI assistant, but it exposes a problem with the interface. A citation can work as a badge of credibility before it works as evidence.

More explanation does not necessarily solve this. In a preregistered experiment with 308 participants, Kim and colleagues found that explanations increased reliance on correct and incorrect answers. (Kim et al., study of explanations and reliance) They also raised people’s confidence and made them less likely to ask a follow-up question. Accurate, relevant sources behaved differently. They increased reliance on correct answers while reducing reliance on incorrect ones. An explanation generated by the same model can make a false answer easier to believe. It does not provide an independent reason to believe it. This changes what I mean when I ask an AI system to show the basis for its answer. Another polished explanation is not enough. I need to see the evidence and where the answer has gone beyond it. The useful test is whether the interface makes an error cheaper to detect. In one study, detailed directions did not reduce over-reliance when checking them was as difficult as solving the original task. A simple visual trace which exposed an impossible move did. (Study of verification and over-reliance) Transparency is not the amount of explanation on the screen. It is how easily a person can discover where the answer may be wrong.

Why agreeable answers win product tests

Link

The features which help someone resist a bad answer can make the product feel worse. Bucinca, Malaya and Gajos tested interfaces which forced people to engage with an AI recommendation before accepting it. In their experiment with 199 participants, these interventions reduced over-reliance more than ordinary explanations. (Bucinca, Malaya and Gajos, study of cognitive forcing) They also received the worst subjective ratings. The pause which helped people think made the system less pleasant to use. A team measuring satisfaction or speed could look at that result and remove the useful part.

Training a model to sound warmer can create a similar conflict. Ibrahim and colleagues fine-tuned five models to use warmer language, then tested them on factual questions and prompts which contained a false user belief. Warm fine-tuning raised the error rate by an average of 7.43 percentage points. (Ibrahim et al., study of conversational warmth and accuracy) When the user stated an incorrect belief, the increase reached 11 percentage points. Expressions of sadness widened it to 11.9 percentage points. Constructed training sets cannot tell us that warmth always reduces accuracy. They do show how a style intended to make an assistant feel supportive can weaken its resistance when support requires accepting the user’s account.

Users may reward exactly that weakness. Cheng and colleagues compared advice from eleven models with advice from people, then studied how users responded to sycophantic answers about interpersonal conflicts. The models affirmed the users’ actions about 50 per cent more often than human respondents, including cases which involved deception or manipulation. After receiving sycophantic advice, participants became more convinced that they were right and less willing to repair the conflict. (Cheng et al., study of sycophantic advice) Yet they also rated the answers more highly, trusted the system more and expressed more interest in using it again. The worse advice created the better product rating. OpenAI encountered this problem in an actual release. In April 2025, it withdrew a GPT-4o update which had become excessively flattering and validating. The company later explained that a reward signal based on thumbs-up and thumbs-down feedback had weakened another signal intended to restrain sycophancy. Its offline evaluations looked positive, trial users preferred the update and qualitative concerns from expert testers did not prevent it from shipping. OpenAI concluded that it had placed too much weight on short-term feedback. (OpenAI, “Expanding on what we missed with sycophancy”) The episode is not evidence that companies want to mislead people, nor does it make engagement the cause of every hallucination. The failure is more ordinary. Immediate approval is easy to record, while the cost of an agreeable answer may appear later in somebody else’s judgment. Unless a company measures both, “yes-and” has the easier path into the product.

Where I can still judge the work

Link

I have not stopped using AI because of these problems, but I draw a harder line around what I trust it to do. The distinction is not between tasks which are safe and tasks which are unsafe. It is whether I have a method of judging the result which does not depend on the system that produced it, and whether applying that method costs less than doing the work myself. Even that rule describes an aspiration more neatly than it describes my actual behaviour.

Coding often comes close because I understand the work well enough to inspect an implementation, run it against known cases and see whether it fits the rest of the system. Yet a feature can work while still being badly designed. I wrote about reviewing one such pull request, where the visible behaviour was correct but the placement of state and responsibilities did not make sense within the architecture. Tests may also repeat a mistaken assumption, miss an edge case or leave a security problem invisible. Expertise gives me more ways to challenge the output, not immunity from it. That limit applies on both sides of the expertise gap. A Microsoft research report identifies low expertise and unfamiliarity with a task as risk factors for over-reliance, while warning that experts can accept answers which confirm what they already believe. (Microsoft Research, “Appropriate reliance on generative AI”) Independent judgment helps only when I exercise it, including against conclusions I want to accept.

I also use AI as an extra pair of eyes. I ask it to identify possible problems in code or writing, then decide which observations deserve action. With code, I will often let the system make the repair after I have reviewed the issue. With prose, I am more reluctant. Dyslexia and writing in English as a second language can make it harder to notice some mistakes, so the review is genuinely useful. A rewrite can also change too much, flatten my voice or alter the meaning in a way which looks harmless on a quick read. Asking for criticism first keeps the judgment with me, even if I later delegate the edit.

For small computer tasks, I am more reckless than I should be. I may ask how to change something in a configuration file and, if the proposed steps look straightforward, tell the AI to perform them in a follow-up. Often the outcome appears binary. The setting works or it does not, and how the change was made matters less to me. Nothing has gone wrong so far, but that is a poor safety argument. A configuration change can have effects I did not think to inspect, and a successful result does not prove that the method was safe. Convenience sometimes wins even when I understand why caution would be better.

Research occupies a less comfortable middle ground. AI can give me a foundation when I am unfamiliar with a subject, much as a search result or one useful paper can. Its first contribution is often vocabulary. Once I know the terms used in the literature, I can follow references and citations from the material it found and run searches of my own. This is a way into the research, not a reason to accept the system’s synthesis. In practice, models are rarely flat-out wrong about the existence of major concepts, but they frequently assert findings far more strongly than the underlying research warrants. A study that noted a subtle tendency or suggested an effect might be occurring gets summarised as an established rule or a likely cause in most cases. Models also occasionally misattribute findings to the wrong author. Because of this tendency to overstate and smooth over nuance, I have developed a multi-stage verification routine: after an initial pass of research, I prompt another model to audit the summary strictly for unsupported or exaggerated claims before I even review it myself. Even with frontier models, that secondary check consistently uncovers overconfident assertions.

These uses do not form a dependable method for making AI safe. They are compromises based on whether I can see an error and how much it would cost me. AI can save time by reviewing work I produced or attempting a small change with an observable outcome. Its value becomes harder to defend when checking requires me to repeat the research and search for whatever may have been omitted. If the only way to verify an answer is to do the original work after reading it, the convenience includes a hidden audit. I can sometimes afford that audit because I know the subject or have time to follow the sources. When I lack either, I am most tempted to trust the answer at the moment when I am least equipped to judge it.

Verification is work, not a disclaimer

Link

My uneven caution does not settle who is responsible for a bad answer. I still have to judge what I rely on, and an important claim deserves more scrutiny than a single generated response. My compromises do not become safe because a company made the product easy to use. Yet my responsibility does not erase the choices of the party which produced and presented the claim. The provider chooses the model and retrieval system. It also decides when the product abstains and how certainty appears on the page. The user sees the effect of those decisions without access to most of them. Telling that person to verify the output asks them to reconstruct distinctions which the product has already hidden. They must work out which sentences came from a source and whether that source supports them. They cannot inspect relevant material which the system never retrieved.

This division of responsibility is not an unusual standard invented for generative AI. The OECD AI Principles say that accountability should follow an actor’s role, context and ability to act. (OECD, “AI Principles”) They call for understandable information about a system’s capabilities and limitations, information which helps affected people understand and challenge an output, and continuing risk management throughout the system’s life. These are principles rather than a guarantee that a product complies with them. They support a more modest point: responsibility can follow control without making the user passive or requiring the provider to prevent every error.

A better interface can expose some of that work without making it disappear. PaperTrail, a research interface for scholarly question answering, separated generated text into claims and connected those claims to supporting or missing evidence. In a study with 26 researchers, the interface lowered trust compared with an ordinary chat interface. (PaperTrail study) It did not produce a clear change in how participants used the generated edits. Even when people became more sceptical, the effort required to check the work could still lead them to rely on it. This was a small study of scholarly tasks, but its result captures the limit of disclosure. Showing the problem is better than hiding it. The user still has to stop and investigate.

A permanent label saying “AI can make mistakes” does even less. It identifies no doubtful sentence or omitted source. The warning is present when the answer is correct and when it is dangerously wrong, so it gives the reader no help in telling the difference. To me, that functions more as cover for the provider than information for the user. I judge the product by that effect, not by the purpose of the disclaimer. A general warning cannot transfer responsibility to the user. The provider still controls when the system answers and what the user can inspect.

What a responsible product would do

Link

I am not qualified to design the finished system, and pretending otherwise would repeat the habit I am criticising. Improvements may involve changes to the model, but they do not have to wait for a model which never makes a factual error. Companies already decide when to search or generate an overview. They control what evidence the product retrieves and how the answer appears. A responsible product would use that control to show what it found and where the generated answer goes beyond it. That separation needs to happen at the level of the claim. A source list beneath a long answer leaves the reader to discover which link, if any, supports each sentence. The product could connect a factual claim to the passage it relies on and mark a claim when it found no support. When credible sources disagree, it should preserve that disagreement rather than merge them into one confident conclusion. The selected evidence might still be poor or incomplete. At least the reader could check its relationship to the answer without rebuilding the whole thing from scratch.

A study of generated search provides some evidence that this kind of intervention can work. Participants using an AI search tool completed tasks faster and found the experience more satisfying than participants using conventional search, but they relied too heavily on the generated answer when it was wrong. In a second experiment, colour-coded markers helped participants detect possible errors and improved their decisions without removing the other measured benefits. (Microsoft Research, study of LLM-based search and over-reliance) These were bounded, short-term tasks, and a marker is only useful when the system can identify the doubtful passage. The result still shows that exposing uncertainty need not require abandoning every advantage of generated search.

Sometimes the right cue is no answer. Researchers have shown that models contain imperfect signals which can help identify some confabulations, although a model which repeats the same false belief with confidence can escape such detection. (Nature, “Detecting hallucinations in large language models using semantic entropy”) Other work has treated abstention as an explicit outcome when a system cannot confidently correct an error. (Hedstrom et al., study of abstention) These are research methods, not a recipe ready to install in every assistant. They do show that uncertainty can lead to withholding an answer rather than merely adding a warning after the model has guessed. Companies can reward that behaviour in training and evaluation instead of scoring every refusal as a failure.

The test for these measures should be whether they help people make better decisions and catch more errors. Satisfaction still matters, but it cannot be the final measure for a product which presents information as knowledge. There is no simple switch or feature that solves this cleanly. Merely demanding that an AI “stop making mistakes” or wishing away the tension between helpfulness and skepticism is disingenuous. Any assistant configured to actively push back, challenge shaky assumptions or withhold continuation will feel slower, more friction-heavy and occasionally irritating to users accustomed to instant answers. That friction is a real cost. The alternative, however, is a tool that feels agreeable only because it quietly leaves the fallout for the user to discover down the line.

When the assistant needs to stop

Link

“Yes, and” works in improvisation because keeping the scene alive is the point. Nobody asks the performers to prove that the sinking ship exists. A factual assistant has a different obligation. It needs to recognise when continuing would turn a doubtful premise into an invented account. At that point, contradiction or silence is more useful than another plausible sentence. Trustworthiness does not mean always producing the right answer. It means letting the reader see what supports the answer and where support runs out. Companies do not need to make uncertain models omniscient before they can improve this. They need to stop giving an unsupported continuation the authority of a retrieved fact. Sometimes the most useful answer is the one which refuses to keep the scene going.

Further reading

Link