Daily Guardian UAEDaily Guardian UAE
  • Home
  • UAE
  • What’s On
  • Business
  • World
  • Entertainment
  • Lifestyle
  • Sports
  • Technology
  • Travel
  • Web Stories
  • More
    • Editor’s Picks
    • Press Release
What's On

Apple’s latest refurb drop brings cheaper MacBooks, iPhone 16 Plus, and Apple Watches

August 15, 2026

CM Hypermarkets Opens Fifth UAE Outlet in Ajman’s Jurf

August 15, 2026

Android 17 is getting a feature iPhone users have had for years

August 15, 2026

Apple has a plan to fix its memory chip problem, but the US government says no

August 15, 2026

WhatsApp will finally let you react with the emoji you actually want, but there’s a catch

August 15, 2026
Facebook X (Twitter) Instagram
Finance Pro
Facebook X (Twitter) Instagram
Daily Guardian UAE
Subscribe
  • Home
  • UAE
  • What’s On
  • Business
  • World
  • Entertainment
  • Lifestyle
  • Sports
  • Technology
  • Travel
  • Web Stories
  • More
    • Editor’s Picks
    • Press Release
Daily Guardian UAEDaily Guardian UAE
Home » Anthropic aims to fix one of the biggest problems in AI right now
Technology

Anthropic aims to fix one of the biggest problems in AI right now

By dailyguardian.aeJuly 2, 20242 Mins Read
Share
Facebook Twitter LinkedIn Pinterest Email

Hot on the heels of the announcement that its Claude 3.5 Sonnet large language model beat out other leading models, including GPT-4o and Llama-400B, AI startup Anthropic announced Monday that it plans to launch a new program to fund the development of independent, third-party benchmark tests against which to evaluate its upcoming models.

Per a blog post, the company is willing to pay third-party developers to create benchmarks that can “effectively measure advanced capabilities in AI models.”

“Our investment in these evaluations is intended to elevate the entire field of AI safety, providing valuable tools that benefit the whole ecosystem,” Anthropic wrote in a Monday blog post. “Developing high-quality, safety-relevant evaluations remains challenging, and the demand is outpacing the supply.”

The company wants submitted benchmarks to help measure the relative “safety level” of an AI based on a number of factors, including how well it resists attempts to coerce responses that might include cybersecurity; chemical, biological, radiological, and nuclear (CBRN); and misalignment, social manipulation, and other national security risks. Anthropic is also looking for benchmarks to help evaluate models’ advanced capabilities and is willing to fund the “development of tens of thousands of new evaluation questions and end-to-end tasks that would challenge even graduate students,” essentially testing a model’s ability to synthesize knowledge from a variety of sources, its ability to refuse cleverly worded malicious user requests, and its ability to respond in multiple languages.

Anthropic is looking for “sufficiently difficult,” high-volume tasks that can involve as many as “thousands” of testers across a diverse set of test formats that help the company inform its “realistic and safety-relevant” threat modeling efforts. Any interested developers are welcome to submit their proposals to the company, which plans to evaluate them on a rolling basis.











Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Keep Reading

Apple’s latest refurb drop brings cheaper MacBooks, iPhone 16 Plus, and Apple Watches

Android 17 is getting a feature iPhone users have had for years

Apple has a plan to fix its memory chip problem, but the US government says no

WhatsApp will finally let you react with the emoji you actually want, but there’s a catch

U.S. courts will now make government use of spyware tools public

Google’s hyped-up RGB HiLight on the Pixel 11 Pro had me excited for nothing

LG wants to build humanoid robots, and NVIDIA is giving it the brains

Sonos Ace Ultra may launch cheaper than the original Ace, new leak reveals

3 underrated movies you can watch for free this weekend (August 14-16)

Editors Picks

CM Hypermarkets Opens Fifth UAE Outlet in Ajman’s Jurf

August 15, 2026

Android 17 is getting a feature iPhone users have had for years

August 15, 2026

Apple has a plan to fix its memory chip problem, but the US government says no

August 15, 2026

WhatsApp will finally let you react with the emoji you actually want, but there’s a catch

August 15, 2026

Subscribe to News

Get the latest UAE news and updates directly to your inbox.

Latest Posts

U.S. courts will now make government use of spyware tools public

August 15, 2026

Google’s hyped-up RGB HiLight on the Pixel 11 Pro had me excited for nothing

August 15, 2026

LG wants to build humanoid robots, and NVIDIA is giving it the brains

August 15, 2026
Facebook X (Twitter) Pinterest TikTok Instagram
© 2026 Daily Guardian UAE. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.