Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...
Beam: Reflection's 501B open-weight model (reflection.ai)
Ariarule 3 hours ago
extr 2 hours ago
charlieyu1 an hour ago
criemen an hour ago
one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?
dexwiz an hour ago
Cycl0ps 9 minutes ago
htrp 3 hours ago
> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.
Early access, no weights no tech details, just a sign up here for info
zelphirkalt 2 hours ago
janalsncm an hour ago
Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.
This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.
vanuatu an hour ago
mlmonkey 43 minutes ago
wronglebowski 2 hours ago
Loquebantur an hour ago
> We will release the weights, technical report, model card, and developer artifacts later this month.
NorwegianDude 2 hours ago
Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.
mirekrusin an hour ago
onlyrealcuzzo 2 hours ago
Am I missing something?
martini333 2 hours ago
dotancohen 2 hours ago
halJordan an hour ago
janalsncm an hour ago
> That only means their training regime is inferior if their predecessors did so much more with so much less
Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.
None of that means they shouldn’t release their model.
flockonus an hour ago
500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.
jstummbillig 2 hours ago
Centigonal 2 hours ago
swiftcoder 2 hours ago
It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.
I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models
htrp 2 hours ago
mirekrusin 2 hours ago
atlasunshrugged an hour ago
ipsum2 an hour ago
dominotw 15 minutes ago
jauntywundrkind 3 minutes ago
mirekrusin 2 hours ago
aizk an hour ago
vanuatu an hour ago
seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals
michaelkdev an hour ago
drubs 2 hours ago
jeremyjh an hour ago
drubs an hour ago
TheArcane 37 minutes ago
segmondy 38 minutes ago
eaf7e281 an hour ago
It's great to see a company that acknowledges it still needs improvement instead of making false claims.
aeetes 2 hours ago
brumbelow 2 hours ago
hypfer 2 hours ago
It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.
hnedeotes an hour ago
zopper 2 hours ago
keeganpoppen 2 hours ago
vcryan 37 minutes ago
wg0 3 hours ago
Where do I get the data?
I mean, this many models. They have to start somewhere.
Hamuko 3 hours ago
kbwal7 2 hours ago
altcognito 3 hours ago
petu 2 hours ago
e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb
lucrbvi 2 hours ago
I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.
konfusinomicon 2 hours ago
ttul 2 hours ago
sharktheone 3 hours ago
pstuart 2 hours ago