arc-task-gen: 11x Cheaper Inference Costs Than GPT-5.6, New Task Generation for ARC-AGI-1
pathwaycom/arc-task-gen
About the project
Automatically generates new reasoning tasks aligned with the distribution of the public ARC-AGI-1 dataset. It creates a private evaluation set to avoid problems exposed to existing benchmarks and measure how well models perform on unseen problems. The generated tasks.json files follow the standard ARC format and are compatible with existing evaluation harnesses.
This repository is used to validate the performance of the BDH-CQ model. BDH-CQ is an architecture combining non-Transformer-based recursive latent space reasoning. At the 150M parameter configuration, it achieved a pass@2 score of 29.5% on the public ARC-AGI-1 evaluation set. Following the price reduction on July 30, the inference cost per task is 11 times lower than GPT-5.6 Luna(Low).

Lukas Kaiser, co-author of the Transformer architecture, and NYU researcher Richard Jong, among others, independently reproduced the results. Pre-training experiments ranging from 1B to 600B parameters also demonstrated scaling similar to Transformers while maintaining recursive latent reasoning capabilities. Released under the MIT license, it is suitable for researchers evaluating the generalization capabilities of frontier models.
pathwaycom/arc-task-gen
Generates original ARC-AGI-1-style tasks distribution-matched to the public eval set.
Python
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.