METR introduces Expenditure Horizon metric (metr.org)

🤖 AI Summary
METR has introduced a novel metric called "expenditure horizon," which measures an AI agent's optimization capabilities against human labor costs. This metric is significant because it quantifies the point at which an AI becomes less cost-effective than a human for a given task, providing a clearer picture of how much AI can accelerate its own development. By analyzing the NanoGPT speedrun, METR estimates that each 1% improvement in optimization costs about $2,500 in human labor, while preliminary agentic runs reveal expenditure horizons ranging from $0 to $3,000 after spending over $10,000. The expenditure horizon combines traditional R&D benchmarks with a continuous scoring system, enhancing the precision of capability assessment for various optimization tasks. Importantly, it allows for a more interpretable analysis of AI contributions to complex problems already optimized by humans. However, limitations exist, including the method's reliance on smooth returns to optimization and the assumption that agent returns diminish more rapidly than human returns. This innovative approach positions the expenditure horizon as a vital tool for understanding the cost-effectiveness and long-term evolution of AI in research and development contexts, steering the community towards hybrid optimization strategies that leverage both human and AI strengths.
Loading comments...
loading comments...