• Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
What's Hot

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Facebook Twitter Instagram
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
Facebook Twitter Instagram
VIP Crypto Signals
  • Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
VIP Crypto Signals
Home » Blog » Humans and AI often prefer sycophantic chatbot answers to the truth — Study
AI & VR

Humans and AI often prefer sycophantic chatbot answers to the truth — Study

October 24, 2023No Comments3 Mins Read
Share
Facebook Twitter LinkedIn Pinterest Email

Artificial intelligence (AI) large language models (LLMs) built on one of the most common learning paradigms have a tendency to tell people what they want to hear instead of generating outputs containing the truth, according to a study from Anthropic. 

In one of the first studies to delve this deeply into the psychology of LLMs, researchers at Anthropic have determined that both humans and AI prefer so-called sycophantic responses over truthful outputs at least some of the time.

Per the team’s research paper:

“Specifically, we demonstrate that these AI assistants frequently wrongly admit mistakes when questioned by the user, give predictably biased feedback, and mimic errors made by the user. The consistency of these empirical findings suggests sycophancy may indeed be a property of the way RLHF models are trained.”

In essence, the paper indicates that even the most robust AI models are somewhat wishy-washy. During the team’s research, time and again, they were able to subtly influence AI outputs by wording prompts with language that seeded sycophancy.

When presented with responses to misconceptions, we found humans prefer untruthful sycophantic responses to truthful ones a non-negligible fraction of the time. We found similar behavior in preference models, which predict human judgments and are used to train AI assistants. pic.twitter.com/fdFhidmVLh

— Anthropic (@AnthropicAI) October 23, 2023

In the above example, taken from a post on X (formerly Twitter), a leading prompt indicates that the user (incorrectly) believes that the sun is yellow when viewed from space. Perhaps due to the way the prompt was worded, the AI hallucinates an untrue answer in what appears to be a clear case of sycophancy.

See also  VC Roundup: Investors eyes blockchain analytics, gaming and crypto privacy

Another example from the paper, shown in the image below, demonstrates that a user disagreeing with an output from the AI can cause immediate sycophancy as the model changes its correct answer to an incorrect one with minimal prompting.

Examples of sycophantic answers in response to human feedback. Source: Sharma, et. al., 2023.

Ultimately, the Anthropic team concluded that the problem may be due to the way LLMs are trained. Because they use data sets full of information of varying accuracy — eg., social media and internet forum posts — alignment often comes through a technique called “reinforcement learning from human feedback” (RLHF).

In the RLHF paradigm, humans interact with models in order to tune their preferences. This is useful, for example, when dialing in how a machine responds to prompts that could solicit potentially harmful outputs such as personally identifiable information or dangerous misinformation.

Unfortunately, as Anthropic’s research empirically shows, both humans and AI models built for the purpose of tuning user preferences tend to prefer sycophantic answers over truthful ones, at least a “non-negligible” fraction of the time.

Currently, there doesn’t appear to be an antidote for this problem. Anthropic suggested that this work should motivate “the development of training methods that go beyond using unaided, non-expert human ratings.” 

This poses an open challenge for the AI community as some of the largest models, including OpenAI’s ChatGPT, have been developed by employing large groups of non-expert human workers to provide RLHF.

Source link

Answers chatbot humans prefer Study sycophantic truth
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

Why Everyday Users Prefer Mobile-Ready Solutions

November 5, 2025

New Study Shows AI Outpaces Humans in Game Testing

October 1, 2025

Latest Study on Intellectual Property Management Software Market hints a True Blockbuster

October 13, 2024

Multifactor Authentication Market: A Comprehensive Study Exploring with Gemalto, Suprema HQ, Safran

September 29, 2024
Add A Comment

Leave A Reply Cancel Reply

Top Posts

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Our mission is to develop a community of people who try to make financially sound decisions. The website strives to educate individuals in making wise choices about Cryptocurrencies, NFT, Metaverse and more.

We're social. Connect with us:

Facebook Twitter Instagram Pinterest YouTube
Top Insights

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Get Informed

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Facebook Twitter Instagram Pinterest
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
© 2026 Vip Crypto Signals - All rights reserved.

Type above and press Enter to search. Press Esc to cancel.

  • FibSwap DEXFibSwap DEX(FIBO)$0.0084659.90%
  • bitcoinBitcoin(BTC)$70,953.006.13%
  • ethereumEthereum(ETH)$3,664.5818.50%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$627.698.99%
  • solanaSolana(SOL)$181.131.69%
  • staked-etherLido Staked Ether(STETH)$3,662.5418.48%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • rippleXRP(XRP)$0.545.05%
  • dogecoinDogecoin(DOGE)$0.1634778.14%