AI Briefing
KO

Financial-RLVR-10K: 10,000 Financial Reasoning Dataset Released for GRPO Training

·2026.08.16 23:48

Key point

Financial-RLVR-10K, a 10,000-sample financial reasoning dataset containing executable Python code, has been released.

Details

The Financial-RLVR-10K dataset has been released, designed for RLVR (Reinforcement Learning from Verifiable Rewards), GRPO, and PPO fine-tuning of open reasoning models such as Qwen, Llama, and DeepSeek.

Existing LLM-as-a-judge approaches are costly and can provide ambiguous reward signals. This dataset provides Python solution code for every problem, designed to enable deterministic rewards (1.0) through execution in a sandbox.

Key Features:

  • 1,950 Adversarial Logic Traps: Accounting for 19.5% of the dataset, these train models to ignore invalid mathematical conditions (e.g., r <= g in the Gordon Growth model) and judge correct logic.
  • Core Financial Domains: Includes DCF valuation, Black-Scholes option pricing, and corporate WACC.
  • Open Source: Distributed under the MIT license, allowing free use by anyone.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.