AI Briefing
KO

arc-task-gen: 11x Cheaper Inference Costs Than GPT-5.6, New Task Generation for ARC-AGI-1

pathwaycom/arc-task-gen

·2026.08.24 23:06

Automatically generates new reasoning tasks aligned with the distribution of the public ARC-AGI-1 dataset. It creates a private evaluation set to avoid problems exposed to existing benchmarks and measure how well models perform on unseen problems. The generated tasks.json files follow the standard ARC format and are compatible with existing evaluation harnesses.

This repository is used to validate the performance of the BDH-CQ model. BDH-CQ is an architecture combining non-Transformer-based recursive latent space reasoning. At the 150M parameter configuration, it achieved a pass@2 score of 29.5% on the public ARC-AGI-1 evaluation set. Following the price reduction on July 30, the inference cost per task is 11 times lower than GPT-5.6 Luna(Low).

ARC-AGI-1 Efficiency Frontier and Model Performance Comparison Chart
ARC-AGI-1 Efficiency Frontier and Model Performance Comparison Chart

Lukas Kaiser, co-author of the Transformer architecture, and NYU researcher Richard Jong, among others, independently reproduced the results. Pre-training experiments ranging from 1B to 600B parameters also demonstrated scaling similar to Transformers while maintaining recursive latent reasoning capabilities. Released under the MIT license, it is suitable for researchers evaluating the generalization capabilities of frontier models.

GitHub
GitHub repository

pathwaycom/arc-task-gen

Generates original ARC-AGI-1-style tasks distribution-matched to the public eval set.

Python

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.