Muse Spark 1.3 (developer.meta.com)
simonw 9 hours ago
llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.
And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
jmkni 9 hours ago
Definitely an upgrade over 1.2
drusepth 9 hours ago
simonw 9 hours ago
The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.
piker 9 hours ago
m12k 8 hours ago
johntb86 8 hours ago
vunderba 9 hours ago
It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.
In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.
polyterative 9 hours ago
srcreigh 8 hours ago
collabs 8 hours ago
There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.
A medium blog post says
> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?
ACCount37 6 hours ago
It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.
Now we have AIs capable of that and more, and no one bats an eye.
twoodfin 2 hours ago
Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.
Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.
How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?
sly010 21 minutes ago
irthomasthomas 4 hours ago
Gigachad 6 hours ago
werdnapk 5 hours ago
postalcoder 9 hours ago
daemonologist 9 hours ago
(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)
georgemcbay 8 hours ago
While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.
kibae 8 hours ago
labcomputer 6 hours ago
Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.
Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.
Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.
So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.
threetonesun 8 hours ago
ModernMech 9 hours ago
optimalsolver 9 hours ago
reaperducer 8 hours ago
Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.
Much like a mother pelican, they regurgitate what they've been fed.
BeetleB 7 hours ago
https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...
I plan to update it with more pelicans from all the models released since.
(Spoiler alert: They haven't improved much since then).
xhrpost 7 hours ago
murkt 2 hours ago
fc417fc802 2 hours ago
> GPT-5.1 Codex
> monstrosity
What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.
porphyra 7 hours ago
__MatrixMan__ 6 hours ago
I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.
SV_BubbleTime 4 hours ago
They’re all so close in proportions.
elfly 3 hours ago
bodeadly an hour ago
tomrod 9 hours ago
Also 3X token use vs. 1.2
_puk 8 hours ago
jonahx 9 hours ago
Fergusonb 9 hours ago
jonplackett 9 hours ago
jttnr 8 hours ago
EugeneOZ 8 hours ago
Thank you for doing this, I love your benchmark the most!
drob518 8 hours ago
gpt5 4 hours ago
hollowturtle 8 hours ago
tintor 8 hours ago
nojs 6 hours ago
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
fc417fc802 2 hours ago
BeetleB 7 hours ago
cheesecakegood 5 hours ago
tintor 8 hours ago
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
TiredOfLife 8 hours ago
Bird knees bend same way human ones do
0xbadcafebee 7 hours ago
"Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
Or at least they’re not doing it in a plainly obvious manner."ipsum2 7 hours ago
leumon 6 hours ago
wewewedxfgdf 6 hours ago
Aced it, got the job as a senior software engineer.
The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".
fuddle 6 hours ago
sroussey 4 hours ago
labrador 4 hours ago
hunterpayne 3 hours ago
rattray 3 hours ago
bensyverson 3 hours ago
nycdatasci 3 hours ago
smashah 3 hours ago
latentsea 3 hours ago
keeda 33 minutes ago
UltraSane 4 hours ago
treebeard901 an hour ago
salutis 3 hours ago
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
pavs 5 hours ago
I am guessing its not super common, but it happens just so you know.
andytratt 5 hours ago
m00dy 3 hours ago
simonw 13 minutes ago
dwaite 30 minutes ago
dhon_ 29 minutes ago
superfrank 9 hours ago
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
MangoCoffee 8 hours ago
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
KptMarchewa 8 hours ago
TiredOfLife 8 hours ago
dakolli 7 hours ago
dcl 5 hours ago
dakolli 2 hours ago
dcl 2 hours ago
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
avarun an hour ago
superfrank 6 hours ago
CGamesPlay 5 hours ago
sejje 8 hours ago
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
monkpit 40 minutes ago
water-drummer 8 minutes ago
Iolaum 7 minutes ago
bertili 9 hours ago
cbg0 9 hours ago
bermudi 8 hours ago
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
gpt5 4 hours ago
dominotw 9 hours ago
WASDx 9 hours ago
dakolli 7 hours ago
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
skybrian 7 hours ago
switchbak 3 hours ago
comicjk 6 hours ago
dakolli 6 hours ago
atemerev 5 hours ago
Well, that's about the same validity as "In Western astrology..." or "in flat earth theory..."
monkpit 35 minutes ago
lukewarm707 3 hours ago
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
monkpit 36 minutes ago
nl 6 hours ago
That's the opposite of parasitic.
switchbak 3 hours ago
cycrutchfield 2 hours ago
yipinwong an hour ago
People already started using contributor API, and your input is irrelevant.
notatoad 3 hours ago
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
caconym_ 3 hours ago
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
jdm2212 2 hours ago
zackify 40 minutes ago
israrkhan 3 hours ago
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
jmward01 7 hours ago
popularonion an hour ago
We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.
I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.
Lucasoato 9 hours ago
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
dbbk 8 hours ago
ctolsen 7 hours ago
pqdbr 7 hours ago
cdelsolar 12 minutes ago
neuronic 6 hours ago
switchbak 3 hours ago
wrsh07 5 hours ago
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
apodolny 7 hours ago
Gecko4072 9 hours ago
WASDx 9 hours ago
7734128 9 hours ago
0xbadcafebee 7 hours ago
HDBaseT 6 hours ago
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
owaiswiz 6 hours ago
HDBaseT 2 hours ago
userbinator an hour ago
wrsh07 5 hours ago
wxw 9 hours ago
Definitely shows how important a user data flywheel is for RL and model improvement.
majerep 9 hours ago
jumploops 8 hours ago
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
2001zhaozhao 8 hours ago
(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)
dbbk 8 hours ago
winstonp 8 hours ago
hadlock 7 hours ago
HDBaseT 6 hours ago
Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
a012 4 hours ago
cnxhk 8 hours ago
IIIIIllIIII an hour ago
ydna404 2 hours ago
MitziMoto an hour ago
coolcoder613 5 hours ago
keyle 4 hours ago
dcl 5 hours ago
alexboehm 5 hours ago
dcl 5 hours ago
This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?
fibonacci112358 8 hours ago
ryanschaefer 6 hours ago
Aurornis 6 hours ago
maciejgryka 7 hours ago
gehsty 7 hours ago
Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?
I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?
phyrex 7 hours ago
israrkhan 3 hours ago
souvlakee 9 hours ago
LZ_Khan 7 hours ago