AI Briefing
KO

Autonomous Research System Praxist Released

·2026.09.01 09:00

Key point

Sapient Intelligence released Praxist, a research system that autonomously improves existing projects, claiming it achieved higher performance than Claude Code on MLE-bench at approximately 1/12 the cost.

1 / 3

Details

Sapient Intelligence released Praxist, a system that autonomously improves research projects with measurable metrics. Praxist goes beyond hyperparameter tuning by changing methodologies, architectures, and strategies themselves. It repeats a loop where research peers running in parallel generate results, evaluators convert these into structured evidence, and a planning panel synthesizes them to set the research agenda for the next generation.

Differentiation from AutoML and Operating Principle

Unlike AutoML, which only adjusts parameters within a predefined search space, Praxist escapes local optima and passes verified mechanisms and evidence to the next generation through the Deep Innovation Gate and Quality-Diversity allocation. While it uses Codex or Claude Code as manipulation interfaces, it does not replace their interactive agent capabilities but layers a persistent research loop and lifecycle control on top.

MLE-bench Performance and Cost Efficiency

According to the paper results released by the researchers, Praxist won 60 medals (80.0%) on MLE-bench, a standardized set of 75 tasks, with 49 of them being gold medals. This is higher than the comparison baseline, Claude Code based on Claude Opus 4.8 (55 medals, 34 gold medals). Notably, the model expenditure cost for Praxist was measured at $3,054, approximately 1/12 of Claude Code's $38,370. However, since these are the developer's own measurements, re-verification is necessary for actual adoption.

Application Conditions and Limitations

Praxist is suitable for research teams that already have a working baseline and automated evaluation paths, but where the best improvement path is unclear. It is unsuitable for initial exploration stages or tasks where superiority cannot be determined by metrics. Scientific assumptions and domain constraints are owned by the project side, while Praxist handles only orchestration and evidence protocols. Currently, CPython 3.11/3.12 Linux environments are officially supported, while macOS and others are classified as compatibility targets.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.