AI Briefing
KO

NVIDIA AI-Q Ranks #1 on DeepResearch Bench

·2025.08.05 04:51

Key point

NVIDIA's AI-Q Blueprint, leveraging Llama Nemotron models, ranked #1 among open source stacks on DeepResearch Bench.

Details

NVIDIA's AI-Q Blueprint took the #1 spot on Hugging Face's DeepResearch Bench 'LLM with Search' leaderboard. This blueprint proved that highly sophisticated agentic workflows rivaling closed models can be achieved using open source models alone.

Core Tech Stack

  • Llama 3.3-70B Instruct: The foundation model for structured report generation.
  • Llama-3.3-Nemotron-Super-49B-v1.5: A model optimized via NAS (Neural Architecture Search) and knowledge distillation, excelling at multi-step reasoning, tool use, and reflection.
  • NVIDIA NeMo Retriever & Agent toolkit: Handles multimodal retrieval and orchestration of complex agentic workflows.

Strengths of Llama Nemotron

  • Reasoning mode support: Can switch between normal chat mode and deep Chain-of-Thought mode via system prompts.
  • Efficient deployment: Supports 49B parameters and a 128K context window, and can run efficiently on a single H100 GPU.

Benchmark Results AI-Q was validated through hallucination detection, multi-source synthesis, citation reliability, and RAGAS metrics. On DeepResearch Bench, which covers over 100 real-world research tasks, it achieved an overall score of 40.52, the best performance among fully open-licensed stacks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.