• Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
What's Hot

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Facebook Twitter Instagram
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
Facebook Twitter Instagram
VIP Crypto Signals
  • Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
VIP Crypto Signals
Home » Blog » Researchers find LLMs like ChatGPT output sensitive data even after it’s been ‘deleted’
AI & VR

Researchers find LLMs like ChatGPT output sensitive data even after it’s been ‘deleted’

October 2, 2023No Comments3 Mins Read
Share
Facebook Twitter LinkedIn Pinterest Email

A trio of scientists from the University of North Carolina, Chapel Hill recently published preprint artificial intelligence (AI) research showcasing how difficult it is to remove sensitive data from large language models (LLMs) such as OpenAI’s ChatGPT and Google’s Bard. 

According to the researchers’ paper, the task of “deleting” information from LLMs is possible, but it’s just as difficult to verify the information has been removed as it is to actually remove it.

The reason for this has to do with how LLMs are engineered and trained. The models are pretrained on databases and then fine-tuned to generate coherent outputs (GPT stands for “generative pretrained transformer”).

Once a model is trained, its creators cannot, for example, go back into the database and delete specific files in order to prohibit the model from outputting related results. Essentially, all the information a model is trained on exists somewhere inside its weights and parameters where they’re undefinable without actually generating outputs. This is the “black box” of AI.

A problem arises when LLMs trained on massive datasets output sensitive information such as personally identifiable information, financial records, or other potentially harmful and unwanted outputs.

Related: Microsoft to form nuclear power team to support AI: Report

In a hypothetical situation where an LLM was trained on sensitive banking information, for example, there’s typically no way for the AI’s creator to find those files and delete them. Instead, AI devs use guardrails such as hard-coded prompts that inhibit specific behaviors or reinforcement learning from human feedback (RLHF).

In an RLHF paradigm, human assessors engage models with the purpose of eliciting both wanted and unwanted behaviors. When the models’ outputs are desirable, they receive feedback that tunes the model toward that behavior. And when outputs demonstrate unwanted behavior, they receive feedback designed to limit such behavior in future outputs.

See also  Starknet Investors Might Be Difficult to Find
Despite being “deleted” from a model’s weights, the word “Spain” can still be conjured using reworded prompts. Image source: Patil, et. al., 2023

However, as the UNC researchers point out, this method relies on humans finding all the flaws a model might exhibit, and even when successful, it still doesn’t “delete” the information from the model.

Per the team’s research paper:

“A possibly deeper shortcoming of RLHF is that a model may still know the sensitive information. While there is much debate about what models truly ‘know’ it seems problematic for a model to, e.g., be able to describe how to make a bioweapon but merely refrain from answering questions about how to do this.”

Ultimately, the UNC researchers concluded that even state-of-the-art model editing methods, such as Rank-One Model Editing “fail to fully delete factual information from LLMs, as facts can still be extracted 38% of the time by whitebox attacks and 29% of the time by blackbox attacks.”

The model the team used to conduct their research is called GPT-J. While GPT-3.5, one of the base models that power ChatGPT, was fine-tuned with 170 billion parameters, GPT-J only has 6 billion.

Ostensibly, this means the problem of finding and eliminating unwanted data in an LLM such as GPT-3.5 is exponentially more difficult than doing so in a smaller model.

The researchers were able to develop new defense methods to protect LLMs from some “extraction attacks” — purposeful attempts by bad actors to use prompting to circumvent a model’s guardrails in order to make it output sensitive information

See also  How to Find Lost Bitcoins: a Full Guide

However, as the researchers write, “the problem of deleting sensitive information may be one where defense methods are always playing catch-up to new attack methods.”

Source link

ChatGPT data deleted Find LLMs output researchers sensitive
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

YGG Play Shuts Down Services as Yield Guild Pivots to AI Data

July 31, 2026

How to Use ChatGPT for Crypto Trading

July 24, 2026

Build It or Kill It: Find the DeFi Bets Worth Making on July 2

June 16, 2026

Inside PlayerID: PSG and Matchain’s New Platform for Athlete Data Ownership

October 20, 2025
Add A Comment

Leave A Reply Cancel Reply

Top Posts

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Our mission is to develop a community of people who try to make financially sound decisions. The website strives to educate individuals in making wise choices about Cryptocurrencies, NFT, Metaverse and more.

We're social. Connect with us:

Facebook Twitter Instagram Pinterest YouTube
Top Insights

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Get Informed

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Facebook Twitter Instagram Pinterest
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
© 2026 Vip Crypto Signals - All rights reserved.

Type above and press Enter to search. Press Esc to cancel.

  • FibSwap DEXFibSwap DEX(FIBO)$0.0084659.90%
  • bitcoinBitcoin(BTC)$70,953.006.13%
  • ethereumEthereum(ETH)$3,664.5818.50%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$627.698.99%
  • solanaSolana(SOL)$181.131.69%
  • staked-etherLido Staked Ether(STETH)$3,662.5418.48%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • rippleXRP(XRP)$0.545.05%
  • dogecoinDogecoin(DOGE)$0.1634778.14%