Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...M5 Ultra Mac Studio Review (macstories.net)
simonw 8 hours ago
peri-cl 7 hours ago
Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).
/meta Here's a CSS filter that stops those nuisance chart animations,
macstories.net##*:style(animation: none !important; transition: none !important)redox99 7 hours ago
peri-cl 7 hours ago
tcdent 5 hours ago
Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.
skohan an hour ago
gpugreg 7 hours ago
beastman82 6 hours ago
I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
mathisfun123 6 hours ago
throwaway27448 5 hours ago
_hugerobots_ 5 hours ago
bigyabai 5 hours ago
_hugerobots_ 5 hours ago
bigyabai 2 hours ago
mathisfun123 5 hours ago
tom_ 5 hours ago
throwaway27448 4 hours ago
I don't get these weird parasocial emotional attachments/beefs people have with brands. Talk to a therapist.
mathisfun123 4 hours ago
selectodude 3 hours ago
mathisfun123 3 hours ago
brookst 13 minutes ago
nacs 6 hours ago
ProllyInfamous 4 hours ago
My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."
This seems apt. My next LLM machine will be closer to 96gb+ vRAM.
selectodude 4 hours ago
throwaway27448 6 hours ago
bigyabai 5 hours ago
No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.
throwaway27448 4 hours ago
Crossover works on macos, too. So does moltenvk, so does vanilla wine, etc etc. You can run most games without a hitch these days (allegedly, according to /r/macgaming). But I don't play video games so a GPU would probably be better off in some kid's computer.
bigyabai 2 hours ago
But of course, Apple doesn't allow that as part of their ecosystem. It's really a privilege to have MoltenVK perform worse than the fanmade HoneyKrisp driver. It's valuable when Apple refuses to sign AArch64 CUDA drivers for macOS. It's exciting to pay Crossover to support half of the library Proton offers for free.
Clearly, I'm some sort of ingrate that selfishly demands the best things, without considering how to accommodate the poor trillion-dollar megacorporation.
Eisenstein 5 hours ago
beastman82 5 hours ago
medvezhenok 3 hours ago
_hugerobots_ 5 hours ago
louthy 2 hours ago
Perhaps consider some non-offensive language for your comparison?
tomega2134 4 hours ago
throwaway219450 2 hours ago
32GB is still not that much. I would rather get a Spark and have the RAM to experiment with larger LLMs, even if it was slow.
cyanydeez 3 hours ago
fhub an hour ago
Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.
liuliu 6 hours ago
searealist 4 hours ago
GeekyBear 2 hours ago
RationPhantoms 7 hours ago
Maybe Apple is an acquisition away from changing that balance.
wlesieutre 7 hours ago
kridsdale1 6 hours ago
dagmx 6 hours ago
Apple just shifted to N2. They’re not going to be doing another major shift right away.
And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence https://www.tomshardware.com/tech-industry/semiconductors/ts...
smith7018 2 hours ago
[1] https://wccftech.com/apple-to-move-to-1-4nm-process-soon-to-...
GeekyBear 4 hours ago
> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators
https://www.tomshardware.com/tech-industry/semiconductors/ap...
aurareturn 4 hours ago
jmyeet 7 hours ago
This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.
I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.
Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.
The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.
pama 6 hours ago
kridsdale1 6 hours ago
wmf 6 hours ago
fragmede 2 hours ago
prmoustache 6 hours ago
jmyeet 5 hours ago
Each PC is probably going to cost ~$6k and you're talking about 8000W of electricity draw. That's going to consume multiple 20A circuits even at 240V. And the electricity ain't free either. A Mac Studio seems to draw ~500W max.
Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That's comparable bandwidth to the 100Gbps Infiniband option.
So you're talking about $50-60k of hardware and more power draw and more heat for something that will I'm sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I'm likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.
throw0101c 4 hours ago
Why Infiniband ("IB")? If it's for RDMA, that is possible with certain Ethernet cards/chipsets as well. Certainly Mellanox, but Broadcom:
* https://techdocs.broadcom.com/us/en/storage-and-ethernet-con...
and Intel as well:
* https://www.intel.com/content/www/us/en/support/articles/000...
Link level flow control or priority flow control needs to be supported on the switch ports as well.
glitchc 3 hours ago
happyopossum 16 minutes ago
Where are you buying 8 5090s for under $10k? With CPU, RAM, and (checks comment) infiniband hardware???
You're probably looking at a lot closer to $60k when all is said and done, and that's before you hire an electrician to run a sub panel for your homelab...
weee322 3 hours ago
every company make his own npu (without xai)
probaby in 2028 we will have more concurent firm on market place
traceroute66 6 hours ago
beastman82 6 hours ago
washadjeffmad 5 hours ago
nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it's a comfortable temperature blowing on my legs under the desk without warming my office in the summer.
I put together this system two years ago, so it's a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don't think I would spend $7000 to save 100W, though.
TacticalCoder 4 hours ago
Yeah people don't pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs.
For example I've got all but one of my virtual desktops that put the CPU in powersave mode: I don't need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand.
Now I don't do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.
ActorNightly 4 hours ago
nacs 6 hours ago
Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
peri-cl 6 hours ago
https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...
(Note it's a sparse MoE with only 6B active).
nacs 6 hours ago
That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
well_ackshually 4 hours ago
nacs 3 hours ago
If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).
bitexploder 2 hours ago
I paid $500 for the RAM in Nov 2023 :)
peri-cl 2 hours ago
No wonder Warren Buffet gave up and resigned.
api 5 hours ago
alex7o 5 hours ago
GeekyBear 4 hours ago
> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...
Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028.
https://www.tomshardware.com/tech-industry/semiconductors/ap...
karmakaze 4 hours ago
These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.
lhl 38 minutes ago
This is with llama.cpp. You can of course use vLLM/SGLang well on these cards and they're even faster. On vLLM w/ NVIDIA/Qwen3.8-27B-NVFP4 baseline has a prefill of about 13,000 tok/s. The baseline tok/s is 72 tok/s, but at mtp7, it's 157 tok/s, and w/ dflash7 that goes up to 215 tok/s. On mtp-bench, DFlash2 gets a hair under 300 tok/s w/ the code_python prompt.
srcreigh 7 hours ago
I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.
It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.
The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.
An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
slowin 7 hours ago
Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.
That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).
nowittyusername 5 hours ago
Octoth0rpe 2 hours ago
I think this is true, but also misses that a lot of us are just doing basic flask apps with a react front end. We don't need astra; Something sonnet 4.6 level locally is perfectly sufficient 95% of the time, and maybe 99% of the time.
brandon272 23 minutes ago
It's like watching a discussion about cars available to take on a 100km road trip. A new car gets released that is on par with a Toyota Corolla but it is dismissed as completely useless for a 100km trip because it doesn't have the seat massagers and air ride suspension that the new Escalades have.
The reality is that something like Sonnet 4.6 is still amazingly capable for so many programming tasks, especially if you already have some reasonable level of experience to steer it in the right direction.
And if you think Sonnet 4.6 is still worthwhile, then it seems undeniable that something like Qwen 3.8-27B is also worthwhile.
_hugerobots_ 5 hours ago
slowin 5 hours ago
_hugerobots_ 5 hours ago
slowin 4 hours ago
I'm also a huge fan of local models and think it's absolutely imperative that they continue to advance so we can move off of the Anthropic/OpenAI hosted models. It's important to accurately asses where we are in that journey though.
srcreigh 4 hours ago
slowin 4 hours ago
_hugerobots_ 4 hours ago
fhub an hour ago
Correctness matters much more than speed to me, but if I can get both, that’s obviously very interesting.
brandon272 21 minutes ago
zozbot234 6 hours ago
srcreigh 3 hours ago
isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.
> but this would decrease single-session performance even further
Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]
It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.
And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see
If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.
If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.
[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...
zozbot234 2 hours ago
sethd an hour ago
I ordered the same one for work so I could run more local agents at once (many iOS simulators and Xcode build processes).
sajithdilshan 8 hours ago
That’s like 12 years worth of OpenAI Pro subscriptions
geodel 7 hours ago
Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
vardump 7 hours ago
cyclopeanutopia 7 hours ago
prmoustache 5 hours ago
Kurtz79 7 hours ago
A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.
BatFastard 5 hours ago
Don't you mean an Apple to NVidea comparison?
qwytw 4 hours ago
Is there evidence that's true though? I mean gross margins on subscriptions being negative since the API is seemingly very profitable (if the price is compared with the cost of serving very large open models).
As long as there is pressure from other providers serving cheaper models that are somewhat competitive without having to incur any of the R&D costs raising prices will be tricky.
patrickmcnamara 6 hours ago
kridsdale1 6 hours ago
geodel 6 hours ago
geodel 6 hours ago
I think it goes without saying. And it is eminently evident over last couple of decades that from compute to storage to meals 3rd part providers have saved billions upon billions of dollars to enterprises and individuals alike by providing these essential services.
simonw 7 hours ago
Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.
criddell 7 hours ago
And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.
simonw 6 hours ago
bel8 41 minutes ago
It's a bold strategy cotton, lets see if it pays off for em.
hgoel 6 hours ago
ionwake 4 hours ago
SXX 2 hours ago
Might be if RAM prices get much more reasonable its gonna be 1/3 of the price, but it's very much possible its gonna be half or more.
And if you're buing Mac Studio and not some AI-only locked down board it's possible to reuse it for other purposes.
matt-p 3 hours ago
112233 7 hours ago
Razengan 7 hours ago
kridsdale1 6 hours ago
I appreciate the reference to RUSH: Red Barchetta in the final line.
woah 4 hours ago
glitchc 2 hours ago
Do they include footguns from pointer bugs?
ericmay 6 hours ago
[1] https://www.macworld.com/article/3238319/mac-studio-m5-max-r...
nowittyusername 5 hours ago
throw0101c 4 hours ago
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s
Octoth0rpe 2 hours ago
zamadatix 2 hours ago
Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".
mstaoru an hour ago
manyatoms 5 minutes ago
You run uncensored local models where you can ask questions that would get denied by public providers, or questions that you prefer them not to know the intricate details (like your financial planning)
tempoponet 7 hours ago
This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
_hugerobots_ 5 hours ago
ApolloFortyNine 7 hours ago
I didn't expect this to make the 5090 to look like a good deal.
nacs 6 hours ago
It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
orsorna 6 hours ago
peri-cl 5 hours ago
(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).
orsorna 3 hours ago
asimovDev 4 hours ago
Eisenstein 4 hours ago
nojs 4 minutes ago
hamiltont 4 hours ago
I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
Lwerewolf 4 hours ago
peri-cl 4 hours ago
For generation speed in isolation, yes.
GeekyBear 4 hours ago
akozak 6 hours ago
novaleaf 5 hours ago
dylan604 25 minutes ago
saagarjha an hour ago
happyopossum 11 minutes ago
liuliu 6 hours ago
mtsolitary 2 hours ago
SamuelAdams 6 hours ago
flounder3 5 hours ago
jjtheblunt 3 hours ago
Gracana 3 hours ago
If I switch to Mac OS, I have to sort out a package manager and install all the stuff that's missing, and when it comes to containers... they're just linux VMs. I'd happily cut out the weird proprietary middleman if I could.
crossroadsguy 6 hours ago
mjlee 3 hours ago
I'd be quite surprised if Mac OS alone needs more than 8GB, given that they sell the Neo with 8GB of RAM today.
odkdkekfkwjf 38 minutes ago
kokonokko1337 7 hours ago
Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
steve1977 5 hours ago
I get it on Windows systems, at least when someone wants to use Linux-type tooling. But macOS already supports pretty much all of that natively?
RunSet 5 hours ago
For starters, the source code.
steve1977 4 hours ago
Apart from that, for the UNIX part, the source is available for quite a few components:
https://github.com/apple-oss-distributions
notably also the kernel
throw0101c 4 hours ago
Strictly speaking, Apple can claim to ship a UNIX® operating system:
steve1977 4 hours ago
Gracana 4 hours ago
kokonokko1337 already said it was good enough to run LLMs, presumably RunSet isn't saying the source code is needed to run an inference server.
odkdkekfkwjf 37 minutes ago
dylan604 19 minutes ago
theplumber 6 hours ago
lowbloodsugar 2 hours ago
addaon 6 hours ago
snarfy 8 hours ago
andrekandre 6 hours ago
but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)
chasd00 6 hours ago
drdaeman 2 hours ago
I thought this only applies to LLM inference providers, but not raw GPU rentals.
Eisenstein 4 hours ago
drdaeman 2 hours ago
bel8 31 minutes ago
A $10/mo subscription to OpenCode Go would have done the job for you.
They have models like Kimi K3, Grok 4.6 , GLM-5.3, Mimo 2.6 Pro (launched today, already available) which are happy to follow your orders without accusing you of being a terrorist.
crorella 5 hours ago
BatchJob 7 hours ago
devy 7 hours ago
aenis 4 hours ago
Entry level serious hardware starts at 100k, and a bit better but still almost-useful grade is 200k (8x rtx pro, plus a nice epyc pairing). Thats the sort of thing a salaried expert lets their employer buy them for sort of serious work.
Anything really serious is well north of 1M - not including the housing and commercial grade mains connection. And at best that buys fast Kimi K3 or GLM.
12kaj2 6 hours ago
prmoustache 6 hours ago
saejox 5 hours ago
villgax 5 hours ago
slashtom 5 hours ago
sghiassy 7 hours ago
whalesalad 7 hours ago
jmull 6 hours ago
99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills.
Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.
BatFastard 5 hours ago
Even if you could get a frontier model, you would not be able to run it on any Mac. So speculating on what M7 or M9 will achieve in 5 years (if we even still exist) seems pointless.
sghiassy 5 hours ago
I don’t think Apple is going to lie down and cede AI to the cloud.
geodel 4 hours ago
How about writing mail to President and senators on AI doomsday scenario if frontier labs do not pace themselves?
That mini model on mac mini would scared to hell to do such thing. It need that rugged frontier model to speak truth to power.
BatFastard 2 hours ago
I would love to find an excuse to buy a 10,000 dollar machine! But I cant find one yet. My current cloud bill is in excess of 400 USD per month. Just can't achieve frontier model capabilities locally.
beastman82 6 hours ago
sghiassy 6 hours ago
whalesalad 5 hours ago
fragmede 5 hours ago
sghiassy 2 hours ago
CamperBob2 4 hours ago
Sam's address will probably be more riveting, imaginative, and terrifying than the last couple of Terminator screenplays. Legislators will lobby him to write the laws for them, and the ghost of Harlan Ellison will threaten to sue him.
ajross 4 hours ago
I really don't see who buys this, except people who want the Studio for some other reason. But nothing in the story says you want to fill racks with these instead of Blackwell or TPU parts; it's not even close.
sghiassy 2 hours ago
Think of a company like Apple moving onto your turf. They’re not going to cede AI to the cloud. They want their part of the pie.
So in 7 years, how much AI will be handled locally on your iPhone. And will you have repaid all the debt on your balance sheet before Apple eats your lunch
SXX 2 hours ago
It's way too easy for 1T+ frontier labs to ditch Nvidia. So Nvidia will also put effort to make sure there are open weights models and local hardware available.
And Apple will benefit from this too.
ajross 2 hours ago