Muse Spark 1.3 (developer.meta.com)

464 pointsby bvaldivielso9 hours ago311 comments
https://research.meta.ai/blog/introducing-muse-spark-1-3

simonw 9 hours ago

  llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

4.2266 cents, 38 seconds.

For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.

UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.

And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

jmkni 9 hours ago

lol

Definitely an upgrade over 1.2

drusepth 9 hours ago

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

simonw 9 hours ago

It's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction.

The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.

piker 9 hours ago

I was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon.

m12k 8 hours ago

It's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language?

johntb86 8 hours ago

Someone studied this (among other thigns): https://dylancastillo.co/posts/pelicanmaxxing.html . Pelicans on bikes always face right in this test, but other animals on other transportation methods sometimes face left.

vunderba 9 hours ago

The more generic your prompt, the more generic the response. It's a regression to the "mean" of the training data aka GIGO for AI.

It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.

In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.

polyterative 9 hours ago

as a kid I did them like this. nobody told me to do that. are we all so similar?

collabs 8 hours ago

I sincerely believe I've never had a single original thought™ in my whole life.

There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me.

A medium blog post says

> Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work?

ACCount37 6 hours ago

Westworld is such a time capsule.

It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing there? It was totally a sci-fi premise.

Now we have AIs capable of that and more, and no one bats an eye.

twoodfin 2 hours ago

Indeed: “Our hosts began to pass the Turing test within the first year.”

Required sci-fi suspension-of-disbelief in 2017, and then at some point in the last few years we just blew by that one.

Later seasons of the show were much less dramatically satisfying, but also played out the consequences of the science of artificial intelligence demonstrating as a side-effect that human intelligence and free will might have as much of an uncertain foundation as that of machines.

How much data from the Panopticon, how many parameters would it take to train a model that could predict your responses?

sly010 21 minutes ago

I think the turing test is still very much load-bearing — if you know what I mean.

irthomasthomas 4 hours ago

Tesla had the same thought. He called himself an automata: "entirely controlled by the forces of the medium" It inspired him to create the first remote control vehicle.

Gigachad 6 hours ago

It's just the simplest most recognizable form of a house. Like how a smiley face is so generic and simplistic but everyone will know what it represents. Just two dots and a line yet it's easily and unambiguously understood to represent a human face and a happy emotion.

werdnapk 5 hours ago

Well, all the LLMs are being trained on previous pelicans, so they look the same.

postalcoder 9 hours ago

Search Google Images for "bicycle". Almost all bicycle product shots are staged the same way: side view, going left-to-right. It makes sense to me that given that skew in the training data, the model grounds itself in the bicycle.

daemonologist 9 hours ago

and furthermore, this is because the drivetrain is ~always on the right side of the bike - if you want to inspect or admire a bicycle you look at the right side, as you might look under the hood of a car.

(Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.)

georgemcbay 8 hours ago

> and furthermore, this is because the drivetrain is ~always on the right side of the bike

While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive.

kibae 8 hours ago

Since most languages read from left to right, rightward movement tends to read as forward progression. So when showing a bicycle in side profile, having it face right feels more naturally like it’s moving forward.

labcomputer 6 hours ago

I can’t tell you why it’s always on the right, but it’s always on the same side because of network effects.

Bicycle frames are not fully symmetric left-right because you need things like a mount point for the derailleur hanger, and optionally affordances to keep the chain off the stays when the wheel is removed.

Those things have to be on the same side as the chain. Bikes designed for disc brakes additionally need a mount point for the brake caliper on the opposite side from the chain.

Additionally, rear wheels are not symmetric: the spokes on the chain side connect to the hub closer to the plane of the rim. That is, they are more perpendicular to the wheel’s rotational axis than spokes on the opposite side (which is why you should always mount a single pannier on the chain side). This asymmetry is to provide space for the gears.

So once the industry decided to put the chain on the ride, you can’t very well make a group set designed for a left chain if you want it to work on the vast majority of frames.

threetonesun 8 hours ago

Product shots yes, people riding them its more like 50/50. Also if you search for a specific bicycle race you'll find more going right to left.

ModernMech 9 hours ago

Yes, I do a thing where I ask the machine to generate responses in the form of a lizard talking to a cat. The lizard is always a green gecko and the cat is always orange, which I never specify.

optimalsolver 9 hours ago

Sun is missing a few rays and not wearing sunglasses.

reaperducer 8 hours ago

Is there a reason these pelicans always have roughly the same composition

Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.

Much like a mother pelican, they regurgitate what they've been fed.

BeetleB 7 hours ago

Not when rendered via POV-Ray:

https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...

I plan to update it with more pelicans from all the models released since.

(Spoiler alert: They haven't improved much since then).

xhrpost 7 hours ago

Wow, I actually had this exact idea. I was specifically curious as to how well a given LLM could understand a DSL that hasn't changed much in a couple decades and doesn't have nearly as many examples to learn from online. Seems like it did alright, all things considered.

murkt 2 hours ago

Ohh, horizontal wheels. They’re about as good as I expected, models have pretty bad spatial awareness. I would expect Fable to be a bit better than old models, though.

fc417fc802 2 hours ago

I wonder how a multi-modal model would do with a harness and tool calling? Specifically a "render" command that produced an image output enabling it to iterate. (Well I see you did this manually with gemini 2.5 pro but I still think it would be interesting to explore various harness setups.)

> GPT-5.1 Codex

> monstrosity

What are you talking about? That's clearly a sci-fi pelican on a hoverboard (successor of the humble bicycle) wearing a visor. Truly visionary.

porphyra 7 hours ago

The canonical view of a bicycle is facing right. Usually, people want to draw/photograph/depict the side of the bicycle with the running gear, which is on the right side of the frame for historical reasons.

__MatrixMan__ 6 hours ago

The thing that distinguishes pelicans from other birds does so most strongly in profile. If you're looking straight at one, the throat pouch would be hidden by the beak.

I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.

SV_BubbleTime 4 hours ago

I’m a firm believer in pelicanmaxxing.

They’re all so close in proportions.

elfly 3 hours ago

well it is svg, it is doing it from circles and lines as primitives, it wants to do it simply and kind of builds the whole thing hierarchically. Making it 3d is way more complicated (as the POV example shows) and the prompt doesn't say 3d anyway

bodeadly an hour ago

Yes. It's because you are asking it to generate an image of a pelican riding a bicycle. If someone asked you to draw a pelican riding a bycycle, would you interpret that to mean using 3d photorealism? LLMs follow conventions. The convention for an animal riding a bike is to create a childish 2d line drawing.

tomrod 9 hours ago

What does the mean pelican look like at this point?

Also 3X token use vs. 1.2

_puk 8 hours ago

Red eyes and a tattoo?

jonahx 9 hours ago

If you have a grading rubric, huge points off for adding arms instead of using the wings as arms!

Fergusonb 9 hours ago

I think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is.

jonplackett 9 hours ago

Has any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?

jttnr 8 hours ago

I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?

EugeneOZ 8 hours ago

Absolutely BRUTAL! :)

Thank you for doing this, I love your benchmark the most!

drob518 8 hours ago

Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.

gpt5 4 hours ago

We should just consider the pelican bench as saturated and mostly meaningless.

hollowturtle 8 hours ago

Is there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too

tintor 8 hours ago

It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.

nojs 6 hours ago

The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.

It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.

fc417fc802 2 hours ago

You can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).

BeetleB 7 hours ago

You win this thread's prize:

https://news.ycombinator.com/item?id=49538333

cheesecakegood 5 hours ago

It also works as extremely effective engagement farming, for lack of a better phrase

tintor 8 hours ago

Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.

Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.

TiredOfLife 8 hours ago

https://en.wikipedia.org/wiki/Bird_feet_and_legs#/media/File...

Bird knees bend same way human ones do

0xbadcafebee 7 hours ago

For all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html

  "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
   Or at least they’re not doing it in a plainly obvious manner."

ipsum2 7 hours ago

All of the links show "Error: Gist API returned 403".

leumon 6 hours ago

next, try: "generate an svg of a human hand". this is a prompt where many models fail imo.

wewewedxfgdf 6 hours ago

I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle.

Aced it, got the job as a senior software engineer.

The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".

fuddle 6 hours ago

I also aced my interview by focussing on pelicancode problems, instead of leetcode problems.

sroussey 4 hours ago

You should post your source code you wrote here… ;)

labrador 4 hours ago

I interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.

hunterpayne 3 hours ago

"software developer"...you keep using that word. I do not think it means what you think it means.

rattray 3 hours ago

Is this for real

bensyverson 3 hours ago

It’s better than for real, it’s for LinkedIn

nycdatasci 3 hours ago

Sorry for your loss.

smashah 3 hours ago

499 connections :(

latentsea 3 hours ago

I interviewed as a software developer at Meta. They asked me to do a add legs to the player in a VR world. I couldn't do it. They hired me anyway.

keeda 33 minutes ago

Have you considered that's WHY they hired you?

UltraSane 4 hours ago

If you could actually hand write SVG code on the spot that looked like a realistic pelican riding a bike I would want to hire you for SOMETHING.

treebeard901 an hour ago

Or placed in an asylum next to the people who designed XML

salutis 3 hours ago

"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."

That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!

pavs 5 hours ago

FYI, your renderer breaks with error "git api access error 403", rate limiting error from git, when using cloudflare vpn.

I am guessing its not super common, but it happens just so you know.

andytratt 5 hours ago

excellent thread

m00dy 3 hours ago

I see no point having these pelicans used for anything related model qualification.

simonw 13 minutes ago

At this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!

dwaite 30 minutes ago

is there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.

dhon_ 29 minutes ago

I'm waiting for the models to start responding with "Oh hi Simon!"

superfrank 9 hours ago

I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.

I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.

I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.

Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.

MangoCoffee 8 hours ago

>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap

its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.

KptMarchewa 8 hours ago

I would imagine your interactions with it are more important than the output.

TiredOfLife 8 hours ago

Training on ai generated content is how the models got a big jump in capability

dakolli 7 hours ago

Every lab trains their models with AI generated code at this point.

dcl 5 hours ago

Hopefully, 'validated' AI code

dakolli 2 hours ago

What do you think you're doing when you accept an edit, press thumbs up, or don't ask for modifications after an edit.

dcl 2 hours ago

Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either - does this code do what the user actually asked - is this code actually 'good'

There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.

I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.

That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.

avarun an hour ago

This is exactly what RLVR is, and the reason that models have improved so much at verifiable domains like coding and math while not so much on unverifiable ones like writing and UI design.

superfrank 6 hours ago

Funny. I use it through Opencode Go which gives more use than I can use, but didn't realize it was actually free on Zen. Will switch to that I guess

CGamesPlay 5 hours ago

The useful training data is when you clarify your intent, when you tell the model a different approach would be better, when you consistently refactor towards Y and away from X, and so on. The training data isn’t the code, it’s the session transcript. (Anthropic would call this a “distillation attack” against their model, but in this case the model is you!)

sejje 8 hours ago

If it's a mistake, it should course-correct.

I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.

I have to think that's the future, somehow, and I'm really excited about it.

monkpit 40 minutes ago

Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.

water-drummer 8 minutes ago

Still waiting on them to release weights for Muse Spark 1.2, like they promised to. Wonder if they plan on doing the same for 1.3 which would be crazy

Iolaum 7 minutes ago

Zuck hinted at it n his twitter post but I doubt it.

bertili 9 hours ago

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

cbg0 9 hours ago

But is the score really reflective of the quality or are both models benchmaxxing?

bermudi 8 hours ago

Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.

gpt5 4 hours ago

Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.

dominotw 9 hours ago

how much of it is from reallocation of staff to ai training and labeling

WASDx 9 hours ago

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

dakolli 7 hours ago

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.

skybrian 7 hours ago

I’m retired so it won’t be replacing my labor :)

switchbak 3 hours ago

The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?

comicjk 6 hours ago

My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.

dakolli 6 hours ago

https://en.wikipedia.org/wiki/Commodity_fetishism

atemerev 5 hours ago

"In Marxist philosophy..."

Well, that's about the same validity as "In Western astrology..." or "in flat earth theory..."

monkpit 35 minutes ago

Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this

lukewarm707 3 hours ago

global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.

monkpit 36 minutes ago

You’d expect that, wouldn’t you? But, alas…

nl 6 hours ago

I'm using AI to build things I wouldn't (and/or couldn't) have built before.

That's the opposite of parasitic.

switchbak 3 hours ago

And the unabomber has entered the chat.

cycrutchfield 2 hours ago

Don’t you have some looms to break?

yipinwong an hour ago

Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.

notatoad 3 hours ago

when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

caconym_ 3 hours ago

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

jdm2212 2 hours ago

And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.

zackify 40 minutes ago

I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo

israrkhan 3 hours ago

Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

jmward01 7 hours ago

muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?

popularonion an hour ago

This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"

We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.

I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.

Lucasoato 9 hours ago

A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).

Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.

dbbk 8 hours ago

How is it not SOTA? It's beating 5.6 Sol.

ctolsen 7 hours ago

You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.

pqdbr 7 hours ago

Thats 3 months in AI years.

cdelsolar 12 minutes ago

what is it in dog years

neuronic 6 hours ago

Look at all the valuable software products that Fable 5.1 has produced since yesterday!

switchbak 3 hours ago

If it talks less like a robot, I’d call that a win!

wrsh07 5 hours ago

It's somewhat useful to note just for your own timelines that Fable was reportedly trained in February. I'm not sure when mythos 5.1 finished training, but muse spark 1.3 almost certainly finished more recently than that.

This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.

So you shouldn't think "wow, Facebook has caught up"

You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)

There are downsides too, but I'll discuss those separately somewhere

apodolny 7 hours ago

I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.

Gecko4072 9 hours ago

Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.

WASDx 9 hours ago

I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.

7734128 9 hours ago

Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.

0xbadcafebee 7 hours ago

I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models

HDBaseT 6 hours ago

A small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily.

It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.

[0] https://www.anthropic.com/research/small-samples-poison

owaiswiz 6 hours ago

doesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?

HDBaseT 2 hours ago

In theory, but in practice how difficult is that?

userbinator an hour ago

If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.

wrsh07 5 hours ago

Price segmentation at its finest

wxw 9 hours ago

“contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol.

Definitely shows how important a user data flywheel is for RL and model improvement.

majerep 9 hours ago

The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.

jumploops 8 hours ago

The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data.

The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).

Stats:

1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)

2001zhaozhao 8 hours ago

I have a feeling that Meta is not gonna like what people actually use the contributor model for lol.

(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)

dbbk 8 hours ago

Web Search doesn't have a discount on contributor pricing

winstonp 8 hours ago

It's the perfect model for open-source work because it's gonna end up in the training data anyway

hadlock 7 hours ago

There's a lot of value in agentic loop tool failure + recovery training data

HDBaseT 6 hours ago

Not to mention, this is hyper competitive against even Chinese providers given its multi-modal support.

Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.

a012 4 hours ago

Muse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR which is more attractive to big corps.

cnxhk 8 hours ago

artificial analysis results: https://x.com/ArtificialAnlys/status/2095247787277553929

IIIIIllIIII an hour ago

Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.

ydna404 2 hours ago

For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?

MitziMoto an hour ago

Some of us own and run companies? Cost per performance is a huge deal.

coolcoder613 5 hours ago

I have not tried Muse Spark for code, but I've been using it for a while to write Latin. I find it's one of the best at it, alongside Gemini. For example, I've recently been using it to translate the subtitles of the show I'm watching into Latin, to provide me with a bit more input. (I'm learning Latin, for context)

keyle 4 hours ago

I am very impressed by this model so far. It's faaast and it seems to be just intelligent enough to do really well. It's UI work (simple python UI) is very clean and functional. The UX was 'there'.

dcl 5 hours ago

Very keen to try this after using Claude Code over the last few months. Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?

alexboehm 5 hours ago

Just try opencode, it comes with 1.3 contributor free.

dcl 5 hours ago

Well thats very interesting. Thank you. Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness.

This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?

fibonacci112358 8 hours ago

Is everyone rushing to launch something before Astra tomorrow?

ryanschaefer 6 hours ago

For all of the comments about training: I thought that subscription plans for other models allow the same. Am I mistaken?

Aurornis 6 hours ago

It's a toggle. Some will automatically enable it and you have to turn it off. People who rapidly click through setup flows can miss it and leave it enabled.

maciejgryka 7 hours ago

Does anyone know what the license for this model is? Specifically any word on restrictions about what it can be used for?

gehsty 7 hours ago

As a product, would developers switch to a meta model/harness? I don’t think so.

Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?

I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?

phyrex 7 hours ago

Meta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else?

israrkhan 3 hours ago

it seems like gemini 3.8 flash is more capable and cheaper. The only reason i would use this is if i was willing to share my data with meta, and allow them to train on my data. In that case it becomes dirt cheap.

souvlakee 9 hours ago

Why they didn't use LLM to create html table instead of https://lookaside.fbsbx.com/elementpath/media/?media_id=1048...?

LZ_Khan 7 hours ago

Ha, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld.