AI Briefing
KO

Claude 2

·2024.08.06 04:55

Key point

Anthropic unveiled its new model Claude 2, featuring substantial improvements in performance, safety, and context window.

Details

Anthropic announced Claude 2, a new model with significant improvements in performance and safety. This model is available through the API and a new beta website, claude.ai.

Key benchmark results are as follows:

  • Scored 76.5% on the multiple-choice section of the Bar exam (up from 73.0% for Claude 1.3)
  • Achieved a score in the 90th percentile on the GRE reading and writing exams
  • Scored 71.2% on Codex HumanEval (a Python coding test), a significant jump from 56.0%
  • Scored 88.0% on GSM8k (grade-school math problems)

Users can leverage a context window of up to 100K tokens to process hundreds of pages of technical documentation or books, and can also write long documents such as memos or letters in a single pass.

On the safety front, red-teaming evaluations showed a 2x improvement in the ability to provide non-harmful responses compared to Claude 1.3. The beta service is currently available in the US and UK, with plans to expand globally in the future.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.