It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s (github.com)
embedding-shape 12 hours ago
README could clearly make use of a cleanup, seems to be more like a session log dump now than a good introduction to the project for a new user. Maybe try something like "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project".
carloslfu 11 hours ago
Eufrat 11 hours ago
carloslfu 11 hours ago
Eufrat 11 hours ago
I have to image whatever style of writing this was trained on is a lot more pleasant to read and I feel bad for whoever writes like this now being associated as bad AI writing.
cyanydeez 9 hours ago
makira 9 hours ago
trollbridge 8 hours ago
ricardobeat 6 hours ago
This is the first line of the README. I can't believe people are becoming ok with this, and I'm 100% on the AI train.
mulemisterX 9 hours ago
ig0r0 8 hours ago
metadat 8 hours ago
pram 7 hours ago
hadlock 6 hours ago
Qwen: Looking at you for a new ~35B MoE! Please and thank you
kamranjon 5 hours ago
prometheus1992 12 hours ago
carloslfu 11 hours ago
Balooga 10 hours ago
prometheus1992 9 hours ago
trollbridge 8 hours ago
whartung 12 hours ago
Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.
My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.
carloslfu 11 hours ago
tyre 9 hours ago
trollbridge 8 hours ago
jmward01 4 hours ago
atif089 11 hours ago
carloslfu 11 hours ago
About the specifics, I have only anecdotal evidence, but I guess this info can be found somewhere
red_hare 9 hours ago
jacquesm 9 hours ago
cosmic_cheese 8 hours ago
DOS/Windows and PC clones were by no means the best available, but they were cheap, ubiquitous, and versatile compared to alternatives that were either much better at one task but more expensive or better at everything but wildly expensive. They were "good enough" and represented a solid improvement over what many existing computer users had as well as a good entry point for new users. As such they spread like wildfire and became the standard while the expensive alternatives either became hardcore niche or vanished.
jacquesm 8 hours ago
Though to be fair it was Linux more than Windows that killed them. Dos and Windows were competition for DEC and - ironically - IBM.
c0rruptbytes 4 hours ago
carloslfu 3 hours ago
carloslfu 3 hours ago
amelius 7 hours ago
pornel 7 hours ago
siris9476 7 hours ago
ErenayDev 12 hours ago
carloslfu 12 hours ago
drcongo 12 hours ago
AI;DR
drums8787 11 hours ago
How I have come to detest certain phrases.
bogzz 11 hours ago
drcongo 11 hours ago
brailsafe 10 hours ago
thirtygeo 11 hours ago
karmakaze 12 hours ago
0x457 12 hours ago
carloslfu 12 hours ago
rzzzt 7 hours ago
0x457 7 hours ago
carloslfu 12 hours ago
kethinov 9 hours ago
jonplackett 11 hours ago
cromka 11 hours ago
carloslfu 11 hours ago
carloslfu 11 hours ago
mrob 9 hours ago
egorfine 11 hours ago
ElectricalUnion 5 hours ago
For example, a Macbook Neo (so in theory, something with around 4GiB of free RAM lying around) might eat around 900GB of writes a day while not doing much at all, because it's basically on low on RAM and swapping all the time.
Gigachad 7 hours ago
nikanj 9 hours ago
AmazingTurtle 12 hours ago
At this point I'd much rather see people collaborate on one of these implementations, benchmark against them, or upstream the useful bits into MLX/MLX-LM instead of producing yet another near-identical repo.
The local-LLM ecosystem really does not need every implementation idea rediscovered five times and wrapped in a new README. AI-assisted coding makes producing a new repo cheap; maintaining, benchmarking, and integrating one is the actually valuable part.
api 12 hours ago
That's open source since forever, unfortunately.
docheinestages 12 hours ago
carloslfu 12 hours ago
EyMaddis 12 hours ago
carloslfu 12 hours ago
genxy 12 hours ago
oceanplexian 12 hours ago
carloslfu 12 hours ago
I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was cool.
Barbing 12 hours ago
carloslfu 12 hours ago
Barbing 7 hours ago
(The comments under the parent indicate it was improperly flagged/made dead (maybe could happen just from downvoting?) so glad I hit the Vouch.)
dofm 12 hours ago
carloslfu 12 hours ago
noir_lord 12 hours ago
carloslfu 12 hours ago
dofm 11 hours ago
I do agree that, ultimately, combining your efforts with others working in this whole area is probably really worth it, but I can see how there's an ease of pushing forward on your own these days.
I do not have fast internet so I am not sure when I'll really be able to download the weights but I do have an M1 Max to try this on, so I will at some point!
carloslfu 11 hours ago
dofm 9 hours ago
carloslfu 12 hours ago
It's an experiment for myself but I am committing to maintain it. I've been an oss person for a loooong time, way before AI was a thing. Think about it as a new, from-scratch take at it, not as a re-reproduction.
xlayn 11 hours ago
kzrdude 12 hours ago
genxy 12 hours ago
brailsafe 10 hours ago
brcmthrowaway 9 hours ago
carloslfu 3 hours ago
brailsafe 15 minutes ago
mannyv 9 hours ago
Everyone comes at it from a different point of view, and some approaches work, some don't. And when people do this themselves they learn. Existing projects have their mistakes worked out already.
Maybe one of these people is going to come up with the thing that nobody else thought of because of their experience working the problem from scratch. You may not get that from someone working from an existing project, because existing projects have their approach "baked in."
What all these projects are showing so far is that it's possible to stream from disk, but that the performance isn't ideal. But I'm sure you could take this approach with smaller models and get better performance.
In addition, it's a given that when you work with large data sets performance means organizing the data to take advantage of caches, both disk and cpu. It's not clear how that would work, exactly, given that each run is a not-quite-random walk through the data. The Big Data way is to prebuild all of that as much as possible, which is probably impossible with a big model. But what about a smaller model?
carloslfu 3 hours ago