
dylancastillo.co
July 22, 2026
11 min read
64/100
Summary
Simon Willison has tested major LLM releases using the prompt “Generate an SVG of a pelican riding a bicycle,” creating an informal benchmark in AI. The results often receive significant engagement on platforms like Hacker News, sparking discussions about the benchmark's usefulness.
Key Takeaways
Community Sentiment
Positives
Concerns

Kimi K3, and what we can still learn from the pelican benchmark
Jul 17, 2026

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7
Apr 16, 2026

Gemini 3 Deep Think drew me a good SVG of a pelican riding a bicycle
Feb 14, 2026

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge
May 3, 2026

Agentic coding notes from Galapagos Island
Jul 4, 2026