Anthropic introduces Claude 3.5 Sonnet, matching GPT-4o on benchmarks

Enlarge (credit: Anthropic / Benj Edwards)

On Thursday, Anthropic announced Claude 3.5 Sonnet, its latest AI language model and the first in a new series of “3.5” models that build upon Claude 3, launched in March. Claude 3.5 can compose text, analyze data, and write code. It features a 200,000 token context window and is available now on the Claude website and through an API. Anthropic also introduced Artifacts, a new feature in the Claude interface that shows related work documents in a dedicated window.

So far, people outside of Anthropic seem impressed. “This model is really, really good,” wrote independent AI researcher Simon Willison on X. “I think this is the new best overall model (and both faster and half the price of Opus, similar to the GPT-4 Turbo to GPT-4o jump).”

As we’ve written before, benchmarks for large language models (LLMs) are troublesome because they can be cherry-picked and often do not capture the feel and nuance of using a machine to generate outputs on almost any conceivable topic. But according to Anthropic, Claude 3.5 Sonnet matches or outperforms competitor models like GPT-4o and Gemini 1.5 Pro on certain benchmarks like MMLU (undergraduate level knowledge), GSM8K (grade school math), and HumanEval (coding).

Read 17 remaining paragraphs | Comments

What's your reaction?

Excited

Happy

In Love

Not Sure

Silly

Anthropic introduces Claude 3.5 Sonnet, matching GPT-4o on benchmarks

What's your reaction?

Louisiana’s Ten Commandments Commandment Is Classic Public Schooling. LA GATOR Is, but Almost Wasn’t, the Solution.

SCOTUS Upholds a Tax on Stock Ownership in Narrow Opinion

Russia takes unusual route to hack Starlink-connected devices in Ukraine

Google goes “agentic” with Gemini 2.0’s ambitious AI agent features

AI company trolls San Francisco with billboards saying “stop hiring humans”

Leave a reply Cancel reply

More in:Editor's Pick

AMD’s trusted execution environment blown wide open by new BadRAM attack

Reddit debuts AI-powered discussion search—but will users like it?

Ten months after first tease, OpenAI launches Sora video generation publicly

Your AI clone could target your family, but there’s a simple defense

Posts List

Milei Has Deregulated Something Every Day

Reddit debuts AI-powered discussion search—but will users like it?

Itinerant Baseball Team May “Need” More Taxpayer Funds

Share

What's your reaction?

You may also like

Leave a reply Cancel reply

More in:Editor's Pick

Posts List

Latest Posts