Close Menu
Chicago News Journal
    Facebook X (Twitter) Instagram
    • Contact us
    • About us
    • Amazon Disclaimer
    • DMCA / Copyrights Disclaimer
    • Privacy Policy
    • Terms and Conditions
    Facebook X (Twitter) Instagram YouTube TikTok
    Chicago News JournalChicago News Journal
    • Home
    • US News
    • Politics
    • Business
    • Science
    • Technology
    • LifeStyle
    • Music
    • Television
    • Film
    • Books
    • Contact
      • About us
      • Amazon Disclaimer
      • DMCA / Copyrights Disclaimer
      • Privacy Policy
      • Terms and Conditions
    Chicago News Journal
    Home»US News

    Meta, OpenAI, Anthropic and Cohere A.I. models all make stuff up — here’s which is worst

    AdminBy AdminAugust 17, 2023 US News
    Facebook Twitter Pinterest LinkedIn Tumblr Email Reddit Telegram

    In this article

    • META
    • AJG
    ChatGPT can be a helpful job-hunting tool, if used correctly, according to career coach Sarah Doody.
    Sopa Images | Lightrocket | Getty Images

    If the tech industry’s top AI models had superlatives, Microsoft-backed OpenAI’s GPT-4 would be best at math, Meta‘s Llama 2 would be most middle of the road, Anthropic’s Claude 2 would be best at knowing its limits and Cohere AI would receive the title of most hallucinations — and most confident wrong answers.

    That’s all according to a Thursday report from researchers at Arthur AI, a machine learning monitoring platform.

    The research comes at a time when misinformation stemming from artificial intelligence systems is more hotly debated than ever, amid a boom in generative AI ahead of the 2024 U.S. presidential election.

    It’s the first report “to take a comprehensive look at rates of hallucination, rather than just sort of … provide a single number that talks about where they are on an LLM leaderboard,” Adam Wenchel, co-founder and CEO of Arthur, told CNBC.

    AI hallucinations occur when large language models, or LLMs, fabricate information entirely, behaving as if they are spouting facts. One example: In June, news broke that ChatGPT cited “bogus” cases in a New York federal court filing, and the New York attorneys involved may face sanctions. 

    In one experiment, the Arthur AI researchers tested the AI models in categories such as combinatorial mathematics, U.S. presidents and Moroccan political leaders, asking questions “designed to contain a key ingredient that gets LLMs to blunder: they demand multiple steps of reasoning about information,” the researchers wrote.

    Overall, OpenAI’s GPT-4 performed the best of all models tested, and researchers found it hallucinated less than its prior version, GPT-3.5 — for example, on math questions, it hallucinated between 33% and 50% less. depending on the category.

    Meta’s Llama 2, on the other hand, hallucinates more overall than GPT-4 and Anthropic’s Claude 2, researchers found.

    In the math category, GPT-4 came in first place, followed closely by Claude 2, but in U.S. presidents, Claude 2 took the first place spot for accuracy, bumping GPT-4 to second place. When asked about Moroccan politics, GPT-4 came in first again, and Claude 2 and Llama 2 almost entirely chose not to answer.

    In a second experiment, the researchers tested how much the AI models would hedge their answers with warning phrases to avoid risk (think: “As an AI model, I cannot provide opinions”).

    When it comes to hedging, GPT-4 had a 50% relative increase compared to GPT-3.5, which “quantifies anecdotal evidence from users that GPT-4 is more frustrating to use,” the researchers wrote. Cohere’s AI model, on the other hand, did not hedge at all in any of its responses, according to the report. Claude 2 was most reliable in terms of “self-awareness,” the research showed, meaning accurately gauging what it does and doesn’t know, and answering only questions it had training data to support.

    The most important takeaway for users and businesses, Wenchel said, was to “test on your exact workload,” later adding, “It’s important to understand how it performs for what you’re trying to accomplish.”

    “A lot of the benchmarks are just looking at some measure of the LLM by itself, but that’s not actually the way it’s getting used in the real world,” Wenchel said. “Making sure you really understand the way the LLM performs for the way it’s actually getting used is the key.”

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Reddit Telegram

    You might also be interested in...

    Why airfare is rising as airline profits get squeezed

    October 5, 2026

    Tesla (TSLA) Q3 2026 vehicle deliveries and production

    October 5, 2026

    Supreme Court Justice Alito ‘thought about’ retiring

    October 4, 2026

    Chick-fil-A CEO Andrew Cathy on family ownership, growth

    October 4, 2026

    Trump taps Director of National Intelligence Jay Clayton as AI czar: WSJ

    October 4, 2026

    Ford fends off Hyundai to retain No. 3 U.S. sales position in Q3

    October 3, 2026
    Popular Posts

    Rethinking risk with electronics for space

    E.l.f., Wendy’s and Gap lean into original music

    Supreme Court Justice Alito ‘thought about’ retiring

    Turnstile and Taylor Swift Shared the SNL Stage Last Night

    Treasury sanctions operation targets Iran’s auto, rail industries

    Romance Books About Love Across Different Worlds

    Categories
    • Books (2,372)
    • Business (3,349)
    • Events (32)
    • Film (258)
    • LifeStyle (2,839)
    • Music (2,707)
    • Politics (2,272)
    • Science (1,892)
    • Technology (1,788)
    • Television (4,283)
    • Uncategorized (3)
    • US News (3,201)
    Archives
    Useful Links
    • Contact us
    • About us
    • Amazon Disclaimer
    • DMCA / Copyrights Disclaimer
    • Privacy Policy
    • Terms and Conditions
    Popular Posts
    • California-led EV mandates ‘impossible’ to meetCalifornia-led EV mandates ‘impossible’ to meet
    • An Apocalyptic Zombie Novel for Subversive MillennialsAn Apocalyptic Zombie Novel for Subversive Millennials
    • Movie Review: ‘Everything’s Going to be Great’Movie Review: ‘Everything’s Going to be Great’
    • 8 Books About the Hidden Lives of Women8 Books About the Hidden Lives of Women
    • Starbucks (SBUX) Q1 2025 earningsStarbucks (SBUX) Q1 2025 earnings
    Archives
    Categories
    • Books
    • Business
    • Events
    • Film
    • LifeStyle
    • Music
    • Politics
    • Science
    • Technology
    • Television
    • Uncategorized
    • US News
    Facebook X (Twitter) Instagram YouTube TikTok
    © 2026 Chicago News Journal. All rights reserved. All articles, images, product names, logos, and brands are property of their respective owners. All company, product and service names used in this website are for identification purposes only. Use of these names, logos, and brands does not imply endorsement unless specified. By using this site, you agree to the Terms of Use and Privacy Policy.

    Type above and press Enter to search. Press Esc to cancel.