Are AI Labs Pelicanmaxxing? (dylancastillo.co)

🤖 AI Summary
Simon Willison, an AI researcher, has turned a playful prompt of "Generate an SVG of a pelican riding a bicycle" into a notable benchmark for testing large language models (LLMs). This informal metric, dubbed "pelicanmaxxing," has raised questions about whether AI labs might be optimizing their models specifically to excel at this quirky scenario. To investigate, Willison generated 1,008 SVGs across seven leading models, analyzing their outputs with a scoring judge and feature extraction techniques. The results reveal that AI labs are not systematically enhancing their models for this prompt. Ratings of pelicans and bicycles from various models did not show significant improvements over other animal-vehicle combinations, indicating that there’s little evidence of intentional "pelicanmaxxing." Notably, while all generated pelican-on-bicycle images faced right—a unique trend—the overall quality of these images did not exceed the models' capabilities with other prompts. Willison's findings suggest that, despite the humorous tone of the benchmark, it serves as an useful tool for evaluating LLM performance without bias from overly specialized training. The full exploration is detailed in his analysis, available on GitHub.
Loading comments...
loading comments...