Apodex Launches TRACES, a New AI Discovery Benchmark

Apodex has launched TRACES, a new benchmark designed to evaluate how well artificial intelligence (AI) systems can tackle real-world problems where the answer is not yet known, rather than testing recall against a fixed dataset.

Apodex, the company behind what it calls Discoverative AI, said TRACES transforms real-world problems into executable environments where AI systems can observe, act, use tools, learn from feedback and work toward verifiable outcomes. The benchmark draws on 423 high-value problems assembled from a survey of 561 industries across 16 sectors, and evaluates both the final outcome and the process an AI system used to reach it.

Six capabilities anchor the evaluation

TRACES scores AI systems across six capabilities: Tools (selecting and correctly interpreting external tools), Repair (correcting its own errors once feedback arrives), Alternatives (weighing competing hypotheses as evidence accumulates), Coherence (holding logic intact across a long chain of work), Evidence (grounding conclusions in observation, data or citation) and Scope (stating the conditions under which a conclusion holds).

“TRACES is a benchmark designed specifically to evaluate progress in discoverative AI. It brings together sophisticated efforts in scouting high-value real-world problems, assembling the tools and data needed to build executable environments, and developing a novel scoring system that evaluates not only outcomes but also the discovery process,” said Dr. Sheng Wang, Lead Scientist, Apodex.

  • 423 high-value problems assembled from a survey of 561 industries across 16 sectors
  • Six scored capabilities: Tools, Repair, Alternatives, Coherence, Evidence, Scope
  • Submissions evaluated by both an outcome verifier and a process verifier, with independent review of disputed scores
  • Open for participation from both solver-system builders and organisations proposing new problems

Scoring the process, not just the final answer

Apodex said process verification is central to the benchmark because scientific discovery rarely offers a single graded answer.

“The TRACES process verification is what makes the benchmark unique. In scientific discovery, the answer is one line at the end of hundreds of judgments – what to try next, when the evidence is enough, when to abandon a hypothesis. The capability lives there, and scoring only the last line throws away almost all of it,” said Brian Wang, AI Research Scientist, Apodex.

The launch extends Apodex’s focus on AI systems built for open-world research problems, an area the company, led by founder and chief executive Tianqiao Chen, has positioned around its Self-Evolving Solver architecture. TRACES is open for participation, with teams able to submit solver systems for evaluation or propose new problems from their own field for Apodex to build into executable environments.

Author


Discover more from techcoffeehouse.com

Subscribe to get the latest posts sent to your email.

Use promo code “TCH15” to get 15% off on checkout.

Share your thoughts

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading