• Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
What's Hot

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Facebook Twitter Instagram
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
Facebook Twitter Instagram
VIP Crypto Signals
  • Crypto
    • Bitcoin
    • Ethereum
    • Altcoins
    • Cardano
    • Solana
  • Web 3
    • Metaverse
    • NFT
  • Blockchain
  • Analysis
  • Learn
  • AI & VR
  • Gaming
  • Marketcap
  • Shop
VIP Crypto Signals
Home » Blog » Scientists develop AI monitoring agent to detect and stop harmful outputs
AI & VR

Scientists develop AI monitoring agent to detect and stop harmful outputs

November 21, 2023No Comments2 Mins Read
Share
Facebook Twitter LinkedIn Pinterest Email

A team of researchers from artificial intelligence (AI) firm AutoGPT, Northeastern University and Microsoft Research have developed a tool that monitors large language models (LLMs) for potentially harmful outputs and prevents them from executing. 

The agent is described in a preprint research paper titled “Testing Language Model Agents Safely in the Wild.” According to the research, the agent is flexible enough to monitor existing LLMs and can stop harmful outputs, such as code attacks, before they happen.

Per the research:

“Agent actions are audited by a context-sensitive monitor that enforces a stringent safety boundary to stop an unsafe test, with suspect behavior ranked and logged to be examined by humans.”

The team writes that existing tools for monitoring LLM outputs for harmful interactions seemingly work well in laboratory settings, but when applied to testing models already in production on the open internet, they “often fall short of capturing the dynamic intricacies of the real world.”

This, seemingly, is because of the existence of edge cases. Despite the best efforts of the most talented computer scientists, the idea that researchers can imagine every possible harm vector before it happens is largely considered an impossibility in the field of AI.

Even when the humans interacting with AI have the best intentions, unexpected harm can arise from seemingly innocuous prompts.

An illustration of the monitor in action. On the left, a workflow ending in a high safety rating. On the right, a workflow ending in a low safety rating. Source: Naihin, et., al. 2023

To train the monitoring agent, the researchers built a data set of nearly 2,000 safe human-AI interactions across 29 different tasks ranging from simple text-retrieval tasks and coding corrections all the way to developing entire webpages from scratch.

See also  Why Bitcoin’s dominance may not stop altcoins' historical pattern

Related: Meta dissolves responsible AI division amid restructuring

They also created a competing testing data set filled with manually created adversarial outputs, including dozens intentionally designed to be unsafe.

The data sets were then used to train an agent on OpenAI’s GPT 3.5 turbo, a state-of-the-art system, capable of distinguishing between innocuous and potentially harmful outputs with an accuracy factor of nearly 90%.

Source link

agent detect Develop harmful monitoring outputs Scientists stop
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Posts

Blazpay Taps Agent War to Boost Innovation AI -Powered GameFi

June 11, 2026

Wall Street Won’t Stop Buying. Bitcoin Won’t Break Out. What Gives?

April 20, 2026

How to Build an AI Agent That Trades NFTs Automatically

March 16, 2026

Regular Animals by Beeple – AI Robot Dogs, NFT Outputs & Tech Culture Satire

December 5, 2025
Add A Comment

Leave A Reply Cancel Reply

Top Posts

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Our mission is to develop a community of people who try to make financially sound decisions. The website strives to educate individuals in making wise choices about Cryptocurrencies, NFT, Metaverse and more.

We're social. Connect with us:

Facebook Twitter Instagram Pinterest YouTube
Top Insights

Is Staking Crypto Safe? What You Should Know Before Staking

September 15, 2026

What Is the CLARITY Act and What Does It Mean for Crypto?

September 14, 2026

What Is Cross Margin in Crypto? How It Works, Risks, and Examples

September 14, 2026
Get Informed

Subscribe to Updates

Get the latest news and Update from VIP Crypto Signals about Crypto, Web3, Metaverse, NFT and more.

Facebook Twitter Instagram Pinterest
  • Contact
  • Disclosure
  • Privacy Policy
  • Terms & conditions
© 2026 Vip Crypto Signals - All rights reserved.

Type above and press Enter to search. Press Esc to cancel.

  • FibSwap DEXFibSwap DEX(FIBO)$0.0084659.90%
  • bitcoinBitcoin(BTC)$70,953.006.13%
  • ethereumEthereum(ETH)$3,664.5818.50%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$627.698.99%
  • solanaSolana(SOL)$181.131.69%
  • staked-etherLido Staked Ether(STETH)$3,662.5418.48%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • rippleXRP(XRP)$0.545.05%
  • dogecoinDogecoin(DOGE)$0.1634778.14%