10-task GLM 5.3 harness bench: Claude, OpenCode, pi, zcode, Hermes and 3code (capocasa.dev)

🤖 AI Summary
Recent benchmarks evaluating the performance of various AI harnesses—Claude, OpenCode, pi, zcode, Hermes, and 3code—on a set of ten verified software engineering tasks have unveiled intriguing insights into their efficiency and capability. Notably, 3code emerged as a frontrunner, solving 9 out of 10 tasks while utilizing only 5 million tokens, making it a highly cost-effective option. In contrast, Claude Code, while robust in performance, exhibited a tendency to consume significantly more tokens without completing the challenging "astropi" task. These findings suggest that developers should weigh their choices based on both performance and token efficiency, especially in budget-constrained environments. The benchmarks serve as a valuable tool for identifying potential inefficiencies and guiding future model improvements. For example, OpenCode demonstrated a strong cache rate but incurred a higher token cost, highlighting a trade-off between token usage and performance efficiency. Furthermore, pi has impressed with its adaptability despite lacking model-specific tuning, underscoring the merit of systematic approaches like 3code's targeted fine-tuning. As the independent researcher behind these tests, Carlo plans to expand the scope of such benchmarks, paving the way for more informed choices in the rapidly evolving AI landscape.
Loading comments...
loading comments...