Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Privacy

|

Cookies

|

Contact
llmsai-benchmarksdeveloper-toolscommunity-engagement

Are AI Labs Pelicanmaxxing?

Are AI labs pelicanmaxxing? – Dylan Castillo

dylancastillo.co

July 22, 2026

11 min read

🔥🔥🔥🔥🔥

64/100

Summary

Simon Willison has tested major LLM releases using the prompt “Generate an SVG of a pelican riding a bicycle,” creating an informal benchmark in AI. The results often receive significant engagement on platforms like Hacker News, sparking discussions about the benchmark's usefulness.

Key Takeaways

  • Simon Willison's prompt "Generate an SVG of a pelican riding a bicycle" has become a popular informal benchmark for evaluating AI models.
  • An experiment generating 1,008 SVGs across seven AI models found no significant improvement in the quality of pelican-on-bicycle images compared to other animal-vehicle combinations.
  • The analysis utilized GPT-5.6 Luna for scoring and Gemini 3.1 Flash-Lite for feature extraction to assess the generated images.
  • The results indicate that AI labs do not demonstrate superior performance in generating pelicans on bicycles compared to other prompts.
Read original article

Community Sentiment

Mixed

Positives

  • Dylan's methodology here is significantly more robust than anything I was considering, showcasing a solid approach to evaluating model performance.
  • I'm glad someone ran the numbers on this; it's a refreshing take on quantitative analysis of AI outputs that often get dismissed.
  • The outputs of Qwen3.7 Plus have to be seen to be believed; there's clearly impressive capability on display here.

Concerns

  • Using a LLM to generate complex SVGs is pretty niche; this makes me question the general applicability of the benchmark.
  • There's a lack of evidence that these models have trained on specific combinations like pelicans on bicycles, which raises doubts about their generalization capabilities.
  • The argument that AI labs could be sucking up all the training material without real improvement is disheartening and suggests stagnation.

Related Articles

Kimi K3, and what we can still learn from the pelican benchmark

Kimi K3, and what we can still learn from the pelican benchmark

Jul 17, 2026

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Apr 16, 2026

Gemini 3 Deep Think

Gemini 3 Deep Think drew me a good SVG of a pelican riding a bicycle

Feb 14, 2026

An open-weights Chinese model just beat Claude, GPT-5.5, and Gemini in a programming challenge - ThinkPol

Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge

May 3, 2026

Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island

Agentic coding notes from Galapagos Island

Jul 4, 2026