AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#evaluation
The latest AI and developer news about #evaluation, with the original source and a short summary.
Feed
Trending
Tags
Settings
#evaluation - page 3 | AI Briefing
Agent Evaluation Readiness Checklist
LangChain Blog
·
2026.03.27 23:00
How to Build Evals for Deep Agents
LangChain Blog
·
2026.03.27 00:00
Enterprise AI Deployment
AI21 Labs
·
2026.03.10 20:00
RAG, You've Heard of It… But How Do You Apply It to Your Service?
Woowa Brothers
·
2026.03.10 11:00
What We Learned Using 2 Trillion Tokens for Category Classification
Karrot
·
1
·
2026.02.27 20:00
Orchestrating Modular AI Agents
AI21 Labs
·
2026.02.26 21:00
monday Service + LangSmith: A Code-Centric Evaluation Strategy Built from Day One
LangChain Blog
·
2026.02.18 17:00
Verifying the "bash is all you need" hypothesis
Vercel Blog
·
2026.01.22 22:00
Stable Cinemetrics: A Structural Taxonomy and Evaluation for Professional Video Generation
Stability AI Research
·
2025.10.02 00:00
Gaia2 and ARE Released for Agent Evaluation
HuggingFace Blog
·
2025.09.22 09:00
How to Write Effective Tools for Agents: Working with Agents
Anthropic Engineering
·
2025.09.11 00:00
Enterprise Knowledge Agents
AI21 Labs
·
2025.08.25 21:00
RAG Evaluation Using LLM as a Judge
Mistral AI
·
2025.04.09 09:00
Arabic LLM Evaluation Leaderboard and AraGen Update
HuggingFace Blog
·
2025.04.08 09:00
Implementing Agent Observability with Arize Phoenix
HuggingFace Blog
·
2025.02.28 09:00
How to Quickly Get Started with LLM Evaluation Using OpenEvals
LangChain Blog
·
2025.02.27 03:00
A Case Study of LLM-as-a-Judge-based RAG Evaluation
HuggingFace Blog
·
2024.10.28 09:00
BenCzechMark: Evaluating Czech LLMs
HuggingFace Blog
·
2024.10.01 09:00
LiveCodeBench Leaderboard Released
HuggingFace Blog
·
2024.04.16 09:00
Previous
1
2
3
Next
Previous
1
2
3
Next