I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride bicycles?
> What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing?
We will just have more of the same.
I take exception to that! It's a performative joke for attention that works far more widely than just Hacker News.
what they do have are many different pelicans and people helpfully rating them in the comments.
This is quite possibly reasoning-effort prompt which is injected before the opening <think> token whenever you set a custom reasoning effort, see e.g. DeepSeek-V4 max mode prompt: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise.
And for frontier models I go one step ahead and try to recreate a complex animation video, with the ability for the model to review its own work. And at this Fable is still the top one. Ex: https://www.youtube.com/watch?v=uDAeAuYyl0E (recreation of Claude announcement video) and https://www.youtube.com/watch?v=cSsVNtGPOIg (recreation of a fireship video). Sol did something similar but you can instantly tell its AI slop from very small things, and it just has no narrative or thought put into the writing.
https://mesmer.tools/benchmarks/ai-video-generation , I usually put basic ones here.
Engineers get unbelievably silly about evaluating costs of things.
"The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.
25 cents is 10x the cost of 2.5 cents, but it's still extremely cheap for the product. It's very much the wrong comparison for a world where the primary competition is still humans who need to eat, and it treats percentage differences as more important than absolute differences when the opposite is true.
Secondly, humans vs LLMs are apples vs oranges. It makes no more sense to compare human costs vs LLM costs as it would have to compare human costs vs calculator costs. LLMs are faster and cheaper but extremely different beasts with different limitations. Humans do not one-shot SVGs of pelicans riding bicycles, and they do not charge in tokens.
Comparing LLM cost efficiency is not something that should need to be defended. It's quite straightforward and reasonable...
I would think the opposite because unless people have been hand drawing these with high quality, the training would be on much crappier versions that old AIs have done.
"That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"
Plus obviously humans can still overfit to a specific style of test.
I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts. Just operating on a feeling that labs don't optimise for this (as mentioned, even if they don't training data is filled with these) is not solid enough that criticism shouldn't be leveraged when it comes up.
Please share those again!
One of the things I'm most looking forward to is a lab producing a model that creates a really great pelican riding a bicycle and then a terrible sloth riding a skateboard (or whatever).
I've not seen that myself yet.
> [...] a really great pelican riding a bicycle and then a terrible sloth riding a skateboard [...]
Happy to play ball. You made a blog post a few weeks back on one of the Qwen models with the eye-catching title "Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7" [1].
Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5
Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. Whether this is a case of training data being unsanitized or intentional benchmark targeted training, I cannot say, but it is the case.
And here is Opus 4.7, again via Openrouter: https://imgur.com/a/Qus1Enf
A massive delta in favour of Opus 4.7, despite the pelican Qwen3.6-35B-A3B produced being noticeably better as you rightly pointed out. What does that tell us? Whether intentional or not (with such deltas, I do have my suspicions), any eval with such a delta is clearly polluted and can not be a source of information, especially as its continued existence does hinge on you testing similar prompts in private as a sanity check, yet by your own admission never noticing the plainly apparent delta in quality. I specifically stuck with the skateboarding sloth too, to keep it as fair as possible and found this in less than 5 minutes...
I would not critique your use of this fun benchmark the way I tend to if I did not have evidence to back up my position, including private evals beyond SVGs that I can reliably use to point out major deviations between what a models claimed performance is according to major benchmarks vs the actual performance outside these known test cases.
I will also say that while I have a lot to be critical of regarding Anthropics modus operandi, especially how they present interesting findings like their j-space work, which I found was irresponsibly anthropomorphic in their reporting, especially as this wasn't a first in model interpretability, but mainly a leap due to being applied to a larger model, but of all the labs, they are the ones that never underperform my evals vs public ones and they appear to strictly keep their training data sanitised.
Happy to discuss public vs private evals and the merit of each if you'd like, I do appreciate your reporting in general but just think the SVG benches have become evidently polluted, which is also why even simple queries in my benchmarks are private. Just saw Thinking Machines Inkling model succeed in certain queries that neither Fable 5, nor GPT-5.6 Sol on any reasoning level managed, which I feel is valuable to truly gauge where we are at. Informs my work with models, my views of the industry and my assessment of the future these tools have, along with how to best implement them to enable better UX.
Did you read either the post or the comment it was referencing?
On the note of training on SVGs, I have seen some labs models outperform when prompted for SVGs of certain animal and action combinations (pelican on bike, panda eating burger, etc.) compared to other similarly outlandish prompts for SVG output that are not part of widely reported benchmarks, even shared evidence one of the last times this came up on here.
[0] ... incredible Simon still believes ...
[1] I’m still not convinced that labs ....
I'm sure all sorts of crap pelican riding bicycle SVGs have ended up in the huge crawls of data that the labs feed into their pre-training steps.
What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the performance for this particular task, independent of general model improvements.
The one exception here is Gemini, who have clearly invested a lot of effort in SVG tasks. I have no idea if my stupid benchmark influenced that decision!
Gemini have boasted about how good they are at pelicans riding bicycles, frogs on penny-farthings, giraffes driving a tiny car, ostriches on roller skates, turtles kickflipping skateboards, and dachshunds driving a stretch limousine. So if they trained for the test they did at least expand it a whole bunch! https://twitter.com/JeffDean/status/2024525132266688757
Given the massive delta easily reproducible with some models, is it really doubtful that certain labs have not: https://news.ycombinator.com/item?id=48951229
We are going from pretty good pelican to jumbled mess with a similarly silly, but different prompt across multiple models from multiple labs, both Western and Eastern, both Open Weight and Closed.
I don't know why the standard is is to be sure that it is happening versus it being a plausible risk of making the results useless.
In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts.
Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and producing a watercolour of a pelican. They’re all rendered in an approximately uniform style, even though the svg format has a basically unlimited possibility space.
Blue sky and green grass aren't that surprising, but the color and direction are interesting.
When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like which direction, color of bicycle, general pelican geometry etc. It will be interesting to see if other creatures end up with coincidentally similar design choices or if that's unique to the pelican-bicycle combination.
So the direction may not be that interesting!
Before that it was vertical (although the ordering of the columns was right to left).
Yes, people will usually post or draw a bicycle right to left which is going to ve opposite of what normally is drawn. I tried the prompt in arabic for many models and I don't recall any adjusting it based on that difference at least culturally speaking.
Took some searching and sleuthing to actually figure out what "drive side out" means, as I'm just a casual "from A to B" cyclist: apparently this is referring to the side the chainset, chainwheels and all those things are on.
I haven't seen many AI works that produces a pelican on a bicycle done in a "Ligne Claire" style, for example.
I guess AI's narrows down the output probability space drastically and converge on some agreed upon aesthetics. Works great for computer programs but bad for art.
For example: "generate an SVG of a chessboard seen from a 45 degree angle slightly higher POV" or "generate an SVG of a basketball court from a TV broadcast perspective".
I find Gemini is still the best at creating SVGs.
>Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground
I'd also enjoy the absurdism of "Herring on a pogostick"
You would not expect that to happen if the models trained on the unrecognizable mess, right?
> model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts
And the labs clearly did focus on improving image rendering.
> they have a uniform style
SVG output from LLMs always looks like that. It looked that way from the beginning; no LLM ever produced a watercolor when asked for SVG output. They all render the prompted element centered in the picture. They all tend to draw things going from left to right, and so on.
Go invent your own random alternatives and the AI models have across the board gotten better over time. Insects playing sports, anthropomorphic fruits performing martial arts, wizards conjuring weapons of WWII, whatever you can imagine. I've tried a lot of these, well beyond what I think would be a reasonable thing to specifically train as combinations. If they have given it a corpus of SVG drawings it has learned to extrapolate.
(note: wizards conjuring a tank got me a surprise animated SVG with my Qwen 3.6 35B model)
That doesn't seem right. I use these models as research assistants when writing lots of random blog posts (including in economically ~useless areas like the history of contra dance) and Fable 5 is a serious improvement (when I don't get downgraded!) over Opus 4.6-4.8 which was a serious improvement over Opus 4.
This shouldn’t really come as a surprise, particularly to anyone who’s used diffusion models. The same thing happens when you ask an LLM for a short story [1] without providing any specific details.
Even cranking up the temperature or top_p values is no panacea. The more generic your prompt, the more pedestrian the response.
Your pelican output is thus both in the training set and yet still outside the capability of the model architecture.
And so you are tracking both the capability of the training and also the capability of the querying!
When you receive your first outstanding pelican it will track a gain of capability.
(btw I first mentioned simonw-pelican-into-training-set in May 2025 on twitter.)
My 3D-egyptology-explainer showed a massive uplift for Kimi K3 and this tracks a much improved 3D capability.
Perhaps I'm underestimating the number of pelicans(?!)
```
if is_willison_pelican_blog_post:
[redacted]
```
You haven't seen their final form [1]
[1] final form is a frontend/react/let's not talk about it, library - it caused a great deal of PTSD to me and my previous company's team due to its dogmatic preference for "we use these axioms, end of story", over practical utility - so it was quite challenging to do state of the art tasks such as nested form fields (e.g. 'user.address.personal.line-1'). The PTSD it caused made us all block out the memories, I suppose. But - it had zero dependencies. That is what mattered. It kept us going. We weren't reaching for more. We had plenty of time.
And thank god for that. Because I'd forgotten my watch in California - and this was in Tokyo [2]
[2] a joke within a joke about Jensen's Kyoto gardener story. Beautiful story, drowned out by WatchGate memes. Why can't jokes have layers? Models have trillions. If you miss 100% of the jokes you don't make, make all the jokes. Someone will laugh (eventually, maybe?) Even if it's: "this person + comedy club = full secret service detail". If someone laughs at that - at my own expense? I don't mind. They laughed. I know this is a gibberish, off-topic message - it's also a human message. I just felt we need more such things in our lives these days.
PS: have you physically seen a pelican in real life? (not a joke)
We have several thousand living 15 minutes walk from our house. I recently started adding my wildlife photography (from iNaturalist) to my blog, so I'm posting several new pelican photos a week at the moment: https://simonwillison.net/search/?q=pelican&type=beat%3Asigh...
I asked because I genuinely feel that the % of people working on some of the most important technology these days - things such as these 'strangely shaped tools' (to borrow from nearcyan) - large language models - the younger generation (folks in their early/mid 20s) - it is not unlikely that they have not physically seen the meatspace version of whatever digital correspondence of it that is being packed into latent space.
After all, why waste time going to the SF or Oakland zoo? One can just check Simon's latest pelican blog post and skip the zoo trip - the harnesses are waiting.
I'd be interested to see what comes out, but it also highlights an curious prompt-control-comparison question
I'm saying an example of what not to do is still an example.
Sorry, how again is this the end of the frontier labs?
Competition is always good.
Even as a paying customer, even as an enterprise, your access to US models may be turned off at any time for arbitrary reasons, including someone mis-understanding "Please fix this [open source] code" (which contained security vulnerabilities that were fixed) as a jailbreak.
Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful.
Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.
That's kind-of why I don't think they're doing that. Anything beyond something that works with a simple design templates looks, well, like they tried to do too much with a simple design template.
You’re reading a personal blog and complaining about an open source personal project he runs and distributes for free. He’s allowed to talk about his personal work on his personal blog. Especially considering the cli utility he talks about is directly related to the post.
Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes.
We complain about spammers all the time, what's wrong with that?
He’s a talented developer who is selling nothing and wrote a post on his personal blog about a newly released LLM and mentioned that he used his own tooling to call the API.
He isn’t selling anything in this post (go ahead and show me where in the post I can buy something. A clearly delineated “sponsored by” sentence exists outside of the post, but does not in any way fit the definition of spam). Look into his reputation and blog history, it’s just him talking about what he is passionate about.
I’m only defending Simon because truly independent, non commercial work like his blog is incredibly rare and should be encouraged and not shit on by drive by commenters who can’t be bothered to distinguish between spam and actual content.
If you want to complain about spammers, find some spam first.
Yes, he is. Every time he posts his own link, he is spamming. Every time he posts a one-liner on an AI topic to get upvoted, he is also spamming.
- https://news.ycombinator.com/item?id=46367224#46370862 - "I'm determined to normalize linking to one's own writing, provided it's relevant to the conversation."
- https://news.ycombinator.com/item?id=46367224#46371369 - an expanded argument for that
You not liking the content does not make it spam. Simon being prolific with making new content does not make it spam. Someone sharing their work in a place that it is popular does not make it spam. Self promotion is not necessarily spam.
If you think that popular blog posts about technical subjects are spam, you should get off of HN since that is a huge part of the value of this site.
Yes, he's still a spammer even if you personally like everything he posts and comments. Most users on this forum simply aren't allowed to post their own links that often. He's clearly exempt from this rule because a significant portion of the site's audience is as much of an AI maximalist as he is.
Just remember that the people like Simon who provide this content for free are real people who read these comments. In your case he was polite enough to respond to your criticism.
You can politely object to the frequency of his posts without making it a personal attack. If you ever try putting yourself out there you might learn kindness and empathy towards people who make the internet interesting, even if you don’t like their area of interest.
Does it make you feel good to talk shit anonymously about a kind person? Just move on.
If you had talent and motivation you could just write an extension in a few minutes that blocks SimonW links so you don’t have to be so offended by things that others clearly are eager and want to see.
I ask you exactly the same, does it make you feel good to attack people to defend him? Why are you making this as me attacking him personally (which I didn't do) and not a complaint that what he does is spam, as many other people are also complaining?
I have the impression you think you are doing the right thing defending him so it gives you a freepass to be offend someone that, differently than you, don't think spamming monetized "free content" is kindness.
I won't write an extension to fix a moderation problem that as a decades long user I have see working much better.
It does feel right to respond in kind to an anonymous troll attacking someone working in public for free.
If you had opened the conversation with a well phrased criticism or question (“Simon, are you concerned that you post too much, or that you submit too often to HN?”) I wouldn’t have an issue.
Instead you dropped a drive by comment on a tangential conversation to get a cheap shot in (a spam comment, dare I say?).
You are getting what you gave. You doubled down on it too. Don’t act high and mighty now. You chose the tone, not me.
You aren’t a moderator, you aren’t a creator, and your submissions are so unpopular that they might as well be spam (your combined contributions to this site are less valued by the community than a single post from Simon if we use votes as a signal). You are a taker criticizing givers. It’s easy, and lazy.
Go offer to help moderate, that’s how moderation gets better, not shitposting. Go write the articles that you think others should read instead of complaining about the people writing the articles that others ARE reading. Most of all stop doing things that can hurt people and make the world worse.
My last message was much nicer, but you focused on the single negative part of it. That’s human nature. Imagine how someone creating content with their real name on it feels reading flippant dismissals of their work, and still being gracious enough to give you considered response.
You’re so close to getting it. A moment of self reflection about how you reacted to the single harsh sentence from a stranger might help you the next time you decide to write a harsh sentence to a stranger.
Are you? Really, it's "impressive" that you got to write such a massive wall of nothing, with so much hate and aggressiveness to defend a spammer with zero arguments other than you thinking him posting dozens of links daily isn't spamming.
You spent a lot of time stalking and harassing someone because you got annoyed by them calling out an influencer you like. This is the real shitposting. A moment of self reflection about your reaction would save you a lot of time in life. As much as you love your own HN internet credits, they mean nothing (unless you have a blog with paid sponsorship).
> You are getting what you gave
do you really think your "harsh" words are punishment or something? you are very, very delusional. You are just an incoherent random internet loser with too much time to argue online. And obviously, you won't stop me from reminding this community that simonw is getting paid by posting here.
To the non-trolls: don't be discouraged by people who have nothing but negativity to add. Do something great. Build people up. Be proud of your work. Share it. There will always be a person who has a problem with what you are doing. Ignore that asshole.
Back to you brazukadev, I assume you'll want to get the last word in.
Pretending to be motivating small creators by helping paid influencers hack their way to the top of every thread.
You still need an OpenRouter API Key and be careful this can burn quite a bit of money.
If you look at https://news.ycombinator.com/from?site=simonwillison.net you'll see that I submitted just one out of the last thirty articles from my site that were submitted to Hacker News - and the one I submitted failed to gain any votes.
There seems to be more to producing a better model than brute forcing parameter count after all.
and Tencent is rumored to have done via Japan: https://wccftech.com/china-tencent-gains-access-to-nvidia-bl...
And that's not even considering just smuggling the GPUs in by eg buying them in Singapore.
AI-specific chips also seem to be on the easier side to design & create relative to high performance CPUs & GPUs, so there's no particular reason to expect Chinese domestic designs to continuously lag behind. They have access to the same fabs, after all
$186 billion and $105 billion revenue in 2025 respectively vs. $402 billion? Yes, Google is larger, but they're all in that same ballpark?
ByteDance's 2025 net income isn't that different from Anthropic's Series H funding even ($50bn vs $65bn respectively).
But this is all also ignoring how much of China is state owned (25% of the GDP!), so the available resource pool is dramatically larger than it would appear depending on what the government decides is important
Firstly, the export-restricted GB202s (e.g. 5090, RTX 6000 Pro Blackwell) are fabled in TSMC, and then packaged/made in... China before they supposedly have to be sold out (by US law; but not by Chinese law). You can immediately see the problem there.
Secondly, despite the supposed 'crackdowns' and et al, NVIDIA and their channel partners pretty much will sell to anyone in countries like Singapore without any questions.
Third, there's human "smugglers" who just physically carry em on trips, and Chinese customs is obviously not going to care about the US's laws on Chinese soil.
We may be boiling the oceans but at least we are finally getting some good SVGs of pelicans on bicycles.
So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi.
Kimi is cheapest by 5x but also slowest by 2x
Edit: Actually, looking at the K2.6 response, that's borderline failing too, it's using HTML+CSS+SVG, not just SVG, again failing to follow the prompt properly.
By the way, that website seems like a black hole for information, it says "Expires in 6 days" in the top right which seems really weird for a page hosting couple of KB of data at most.
Making the content auto expire after a short period of time greatly decreases the attractiveness of the site to lots of SEO spammers and other types of abuse, and if someone were to get something malicious or vile posted it will clean up after itself without me having to wade into things.
GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark.
I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters.
If I had to guess it seems to be the difference between memory (params) and intelligence (attention density). I think you need both.
Deepseek V4 Flash, the 284B model, is roughly equivalent to launch GLM 5, the 744B [sic] model.
Yes, this part is accurate. Expert density determines how much raw compute each hidden state gets.
> The number of the experts tells you about the diversity of its skills.
Most people misunderstand this part. Counter-intuitively experts don't develop diverse skills, they instead balance compute during the forward pass, allowing models to increase their parameter count without the MLP layers exploding in memory + compute requirements.
Why does Kimi not use a "Double Cheese Whammy" branding for "their" butchered and stolen IP?
I built a whole ELO scoring mechanism a while back, described here: https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-...
I probably should spend some time on this now, even though the benchmark itself is feeling a bit stale. There's still a lot of demand for a gallery!
Interestingly enough, using an LLM-as-judge is a great way to approach things like this at scale but you do need to invest in some Cohen's Kappa or Fleiss' Kappa understanding which means putting a human in the driver seat to evaluate the effectiveness of your non-human judge. Absent of that, it's just another case of human-centipede but with LLMs.
What does "better" even mean there?
Wow, that's a stark take. I suppose I'm biased towards a scientific viewpoint. All the best.
New hotness: pelicanmaxxing
https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
Getting the compute to run inference for multi-trillion parameter models at any sort of scale and performance is daunting. There are a handful of vendors that have systems that can do this (~ Nvidia NVl-72 class) that pretty much only the frontier labs and hyperscalers effectively have access to.
Lol
The user is asking to to generate an innocent and mundane graphic, possibly as part of a test.
But wait, pelicans cannot ride bicycles! A pelican is a water bird, and bicycles are designed to be ridden humans. Something alarming may be happening here, could this a jailbreaking attempt?
I need to reconsider and reread the user’s request, “make me a svg of a pelican riding a bicycle”. That is a perfectly innocent and legitimate task, as well as popular “benchmark” on social media communities, so I will continue. I need to continue to be on alert and watch out for potential jailbreaking attempts.
Usually, the pattern is that we see a tsunami of planted "China number one" stories boosted by hordes of Chinese "internet commentators", and then the world trembles for a few days until the scam mechanics are revealed.
My would be either: crippling limitations on the model, vast, unfair, and/or illegal subsidies by the CCP regime as a mercantilist attack on Western capabilities (as we've already seen in iron smelting and clean energy), sanctions-busting, gamed benchmarks, outright theft -- or a combination of the above.
In all seriousness, I propose SWE-bench-adversarial-pelican-gen: it's like SWE-bench, but the harness gets interrupted every 5 turns/tool-calls and is asked to produce an SVG of an arbitrary animal before being told to continue, and every few tool call outputs add comment lines that refer to SVGs of pelicans (and, perhaps, how a møøse bit my sister once). And, at the end, once it's 800k tokens deep into context, it's asked to produce an SVG of a pelican and is evaluated against both the pelican and the completion and efficiency of the task.
You're only as good as your ability to solve problems in the midst of an SVG pelican attack.
My only guess is that GLM 5.2 was specifically RLed for SVG generation and that resulted in superior performance.
People seem to have forgotten this fact.
It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.
AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally fair to assess them on that.
Besides, if you move up one layer to "how good is AI at generating valid SVG markup of non-obvious things", pelican on a bike is actually a good test.
So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era.
In fact, this problem (for this test) is also stated by the pelican test author:
"The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.
So don’t go using pelicans to compare models!"
Empirically, they have something very much alike to the human "g factor" - a shared pool of "general intelligence" that all tasks benefit from.
When a "make it bigger, train it harder" upgrade like Kimi K3 or Mythos 5 drops, the performance rises on every metric. Not just the "headline benchmarks" like Mythos and coding/cybersecurity, but also things like literary analysis - which has nearly zero economic value, and isn't commonly post-trained or benchmarked for. And companies keep encountering things like "our carefully trained specialist model with lots of in-domain training on expensive closed datasets just got leapfrogged on our internal benchmarks by a next gen off the shelf generalist".
You can go hard on benchmarkmaxxing post-training, and you can burn millions of GPU-hours on coding RLVR. But, by the very nature of LLMs, a lot of the performance gains in flagship models are broad and domain-inspecific.
"Stiff prose" is more of a "style" thing than a "capability" thing. No one cares about how good an AI is at things like long form creative writing, because that's the opposite of a profitable field. All of LLM behavior is routed through text, so it's very easy to perturb "writing style" by some training elsewhere. Regression evaluation is hard. And the writing-specific post-training LLMs get is usually just cheap RLAF, with all the usual RLAF degeneracy.
Thus, we get the "default styles" that suck from a "creative writing" standpoint. A lot of that is just "what sounded good to the previous generation of LLMs" - and, unlike human readers, LLM evaluators don't get bored from seeing the same cliches repeated 9000 times across 9000 different instances of generated text. Humans tend to update over time from "this sound cool and punchy" to "this is generic AI slop", but RLAF evaluators stay at step 1. What little human-guided optimization this gets is aimed at "copywriting, marketing blurbs, punchy short-form" - and it shows.
You can do a lot there with some aggressive prompting, but the default writing styles suck, and I frankly don't expect that to change soon. No one cares enough to change it.
Pelicans? Used to be a decent proxy for "general model capabilities that no one would benchmaxx for" - a way to probe for that elusive "LLM g factor". Now that it's a known metric, it's very gameable. But it was pretty solid while it was novel and obscure.
Because of this, I presume the Pelican has been in the training data for at least a year+.
The models are very useful, I am afraid they have fundamental limitations though generalizing (it is just hard to evaluate effectively). So it will just be whack-a-mole "can your model do X", and there will always be a new X.
* operates an absurd prompt
* involves SVG coding knowledge, generates a source code artifact
* involves world knowledge (what is a pelican? What is a bicycle? What does each do?” How are each constructed?”)
* when rendered, the coding artifact expresses an image that makes sense to us perceptually, including color and spatial relationships
* different models and settings have different output so it can be used as an evaluation scheme
That said I wouldn’t choose a model based on this! Just like some brain teaser shouldn’t determine employment eligibility.
Also, a way to evaluate a models ability to remove dead code, clean up slop, reorganize, etc.
None of the existing benchmarks test any of the things that truly matter. They were relevant when models struggled to one-shot functions, but we're so beyond that point right now, yet the industry has not kept up.
Interesting to see that prices are converging to an "equilibrium price" regardless of being US or Chinese.