Apodex has launched TRACES, a new benchmark designed to evaluate how well artificial intelligence (AI) systems can tackle real-world problems where the answer is not yet known, rather than testing recall against a fixed dataset.
Apodex, the company behind what it calls Discoverative AI, said TRACES transforms real-world problems into executable environments where AI systems can observe, act, use tools, learn from feedback and work toward verifiable outcomes. The benchmark draws on 423 high-value problems assembled from a survey of 561 industries across 16 sectors, and evaluates both the final outcome and the process an AI system used to reach it.
Six capabilities anchor the evaluation
TRACES scores AI systems across six capabilities: Tools (selecting and correctly interpreting external tools), Repair (correcting its own errors once feedback arrives), Alternatives (weighing competing hypotheses as evidence accumulates), Coherence (holding logic intact across a long chain of work), Evidence (grounding conclusions in observation, data or citation) and Scope (stating the conditions under which a conclusion holds).
“TRACES is a benchmark designed specifically to evaluate progress in discoverative AI. It brings together sophisticated efforts in scouting high-value real-world problems, assembling the tools and data needed to build executable environments, and developing a novel scoring system that evaluates not only outcomes but also the discovery process,” said Dr. Sheng Wang, Lead Scientist, Apodex.
- 423 high-value problems assembled from a survey of 561 industries across 16 sectors
- Six scored capabilities: Tools, Repair, Alternatives, Coherence, Evidence, Scope
- Submissions evaluated by both an outcome verifier and a process verifier, with independent review of disputed scores
- Open for participation from both solver-system builders and organisations proposing new problems
Scoring the process, not just the final answer
Apodex said process verification is central to the benchmark because scientific discovery rarely offers a single graded answer.
“The TRACES process verification is what makes the benchmark unique. In scientific discovery, the answer is one line at the end of hundreds of judgments – what to try next, when the evidence is enough, when to abandon a hypothesis. The capability lives there, and scoring only the last line throws away almost all of it,” said Brian Wang, AI Research Scientist, Apodex.
The launch extends Apodex’s focus on AI systems built for open-world research problems, an area the company, led by founder and chief executive Tianqiao Chen, has positioned around its Self-Evolving Solver architecture. TRACES is open for participation, with teams able to submit solver systems for evaluation or propose new problems from their own field for Apodex to build into executable environments.



Share your thoughts