AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#benchmark
The latest AI and developer news about #benchmark, with the original source and a short summary.
Feed
Trending
Tags
Settings
#benchmark - page 15 | AI Briefing
H Company Unveils UI-Specialized Holo2 Model
HuggingFace Blog
·
2026.02.04 02:00
Emirati Dialect Evaluation Benchmark Alyah Released
HuggingFace Blog
·
2026.01.27 19:00
IBM Unveils Benchmark for Industrial AI Agents
HuggingFace Blog
·
2026.01.21 15:00
Industrial Fieldwork Through an Inclusive Lens: Inclusive AI Agent Benchmark (1)
NAVER CLOVA
·
2026.01.12 00:00
Dissecting AI Agent Evals
Anthropic Engineering
·
2026.01.09 00:00
Evaluating AI's Ability to Conduct Scientific Research
OpenAI Blog
·
2025.12.16 18:00
Qwen3-Omni-Flash-2025-12-01: An Upgrade That Listens, Sees, and Follows More Intelligently
Qwen
·
2025.12.09 06:00
Stable Cinemetrics: A Structural Taxonomy and Evaluation for Professional Video Generation
Stability AI Research
·
2025.10.02 00:00
Hugging Face Unveils RTEB, a Retrieval Benchmark
HuggingFace Blog
·
2025.10.01 09:00
Gaia2 and ARE Released for Agent Evaluation
HuggingFace Blog
·
2025.09.22 09:00
AI Agent Benchmark TextQuests Released
HuggingFace Blog
·
2025.08.12 09:00
STEM and Code Benchmark '3LM' Released for Arabic LLMs
HuggingFace Blog
·
2025.08.01 23:00
Video LMM Benchmark 'TimeScope' Released
HuggingFace Blog
·
2025.07.23 09:00
FutureBench, a Benchmark for Predicting Future Events, Released
HuggingFace Blog
·
2025.07.17 09:00
Solar Pro 2
Upstage
·
2025.07.10 10:00
Introducing HealthBench
OpenAI Blog
·
2025.05.12 19:00
LCLM Evaluation Benchmark HELMET Released
HuggingFace Blog
·
2025.04.16 09:00
BrowseComp: A Benchmark for Browsing Agents
OpenAI Blog
·
2025.04.10 19:00
Arabic LLM Evaluation Leaderboard and AraGen Update
HuggingFace Blog
·
2025.04.08 09:00
DeepSeek-V3-0324 Model Released
HuggingFace Blog
·
2025.03.27 03:00
Introducing the SWE-Lancer Benchmark
OpenAI Blog
·
2025.02.18 19:00
HF Improves Math Evaluation on LLM Leaderboard
HuggingFace Blog
·
2025.02.14 09:00
DABstep Benchmark for Data Agents Released
HuggingFace Blog
·
2025.02.04 09:00
Big Bench Audio Benchmark Released
HuggingFace Blog
·
2024.12.20 09:00
Previous
12
13
14
15
16
Next
Previous
7
8
9
10
11
12
13
14
15
16
Next