Gemini 4 Argon (blog.google)
taylorfinley 6 hours ago
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
spankalee 5 hours ago
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
mapontosevenths 5 hours ago
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
drusepth 5 hours ago
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
esafak 5 hours ago
walthamstow 5 hours ago
KeplerBoy 5 hours ago
macNchz 4 hours ago
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Vacyyyy 4 hours ago
macNchz 2 hours ago
marcus_holmes 25 minutes ago
Still works better than compaction.
honr 4 hours ago
tobias2014 3 hours ago
piyh 3 hours ago
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
throwuxiytayq an hour ago
arizen 5 hours ago
sorrybutidontha 5 hours ago
sarjann 5 hours ago
KeplerBoy 5 hours ago
levelZero 5 hours ago
SomaticPirate 5 hours ago
smartbit 4 hours ago
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
LoganDark 4 hours ago
dleslie 2 hours ago
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
mapontosevenths 30 minutes ago
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
eloisant 5 hours ago
mapontosevenths 4 hours ago
tcoff91 4 hours ago
I don't want to mess with antigravity because my google account is too entrenched in my life.
shmoogy 3 hours ago
8note 2 hours ago
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
jsw97 31 minutes ago
spankalee 3 hours ago
mapontosevenths 27 minutes ago
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
moecables 4 hours ago
Conscat 3 hours ago
starfallg an hour ago
IndeanCondor 5 hours ago
seanthemon 5 hours ago
mapontosevenths 5 hours ago
ody4242 5 hours ago
alightsoul 5 hours ago
warkdarrior 5 hours ago
aspect0545 5 hours ago
FranzFerdiNaN 5 hours ago
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
hexfish 5 hours ago
qmr 5 hours ago
scarmig 5 hours ago
articulatepang 5 hours ago
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
dzhiurgis 2 hours ago
luckydata 5 hours ago
baby_souffle 3 hours ago
Why would the guy who wrote curl share it? We can all build our own now...
Why do the Linux folks need to be so selfless? We can all build our own kernel now...
otabdeveloper4 5 hours ago
dominotw 4 hours ago
taylorfinley 4 hours ago
bel8 5 hours ago
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
seanthemon 5 hours ago
warkdarrior 2 hours ago
bel8 an hour ago
Still impressive that it did so much just to answer my simple question.
amanguliani 5 hours ago
onlyrealcuzzo 4 hours ago
I'm using it to run overnight tasks, and that's it until my quota runs out.
Canceled my subscription.
amanguliani 4 hours ago
8note 2 hours ago
its a lot less chatty imo
awakeasleep 42 minutes ago
In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.
It makes openai models so good to talk to i really have trouble switching.
gottorf 5 hours ago
staticman2 5 hours ago
MILP 5 hours ago
robobo96 3 hours ago
mattjoyce 4 hours ago
nkozyra 4 hours ago
xdavidliu 3 hours ago
- AI is just a tool, like excel; it does what the human operating it tells it to
- next token prediction cannot be true understanding
- models can have no desires and goals, don't anthropomorphize it
However, "hallucination" is very much not one of them
gottorf an hour ago
WarmWash 2 hours ago
I'm assuming that Argon has at least a June 2026 date, but man, the 3 models were a mess with newer information.
yegle 5 hours ago
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
p_l 5 hours ago
Including going first for decompiling AGY binary instead of searching the web for documentation...
IshKebab 5 hours ago
illwrks 4 hours ago
martythemaniak 3 hours ago
Grimburger 3 hours ago
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
esseph 2 hours ago
It's so fucking easy.
From AMD: https://lemonade-server.ai/
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
clw8 2 hours ago
nzeid 2 hours ago
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
gcy 3 hours ago
cyanydeez 3 hours ago
solaire_oa 2 hours ago
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
cyanydeez 28 minutes ago
lardo 3 hours ago
taylorfinley 2 hours ago
danpalmer 3 hours ago
dcl an hour ago
danpalmer 2 minutes ago
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
iknowstuff 2 hours ago
plasticchris an hour ago
_heimdall an hour ago
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
jubilanti 35 minutes ago
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
mirmor23 22 minutes ago
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
nickysielicki 6 hours ago
Nobody has a moat.
aleph_minus_one 5 hours ago
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
mapontosevenths 5 hours ago
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
jaggederest 5 hours ago
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
skybrian 5 hours ago
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
pixl97 4 hours ago
My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.
But ya, lots of hardware everywhere not managed well is how you get sovereign AI.
pianopatrick 5 hours ago
pvab3 5 hours ago
CoolestBeans 5 hours ago
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.
aleph_minus_one 4 hours ago
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
3d2 4 hours ago
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
dansquizsoft 3 hours ago
3d2 4 hours ago
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
ehsankia 5 hours ago
aleph_minus_one 5 hours ago
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
koe123 4 hours ago
usef- 3 hours ago
aleph_minus_one 2 hours ago
If we theoretically found a way to shield or reverse gravity, things in aviation or space travel would be different. Or if we theoretically found a way to make cold fusion work, things would be very different. :-)
It is in my opinion not a good idea to invest in companies for which the feasibility of the business models depends on the capability of making science-fiction stories work.
usef- an hour ago
If there's a 90% chance of present AI companies failing to hit AGI, that 10% can still be worth a lot of money given the size of the prize.
Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).
diomedes 11 minutes ago
mbesto an hour ago
It's also hilarious, because OpenAI had the lead and ceded ground already.
hirako2000 5 hours ago
altruios 5 hours ago
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
culi 5 hours ago
altruios 5 hours ago
culi 2 hours ago
I don't think anyone has any clue how long it will take for them to have actual functioning EUV machines, but I highly doubt they will do it within 2 years.
aleph_minus_one 5 hours ago
Are you sure?
--
China Just Built What TSMC Said Was Impossible
https://www.youtube.com/watch?v=Pk-w279ESHg
--
China Just Built What ASML Feared Most
culi 2 hours ago
To be clear, I think China will eventually crack domestic EUV. And I also think their advances with multi-patterning LUV are remarkable. But there's just a hard physics wall of how far they could possible take it.
Right now they are producing 5nm with multi-patterning LUV but yields are at 20%! It's a massive economic loss but they are heavily subsidizing it because they have no other choice until their EUV program is achieved
danny_codes 3 hours ago
culi 2 hours ago
dgemm 2 hours ago
SwellJoe 5 hours ago
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
scottyah 5 hours ago
SwellJoe 3 hours ago
aqsnow 3 hours ago
SwellJoe 2 hours ago
I'm also skeptical that LLMs can ever invent new ideas.
I may be wrong about how soon the curve will flatten, and I may be wrong about LLMs fundamental limitations. But, I don't think it's extremely obvious that LLMs can have novel ideas or can grow into having novel ideas.
TeMPOraL 5 hours ago
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
SwellJoe 4 hours ago
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
TeMPOraL 4 hours ago
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
freecodeio 3 hours ago
just because more and more agents are doing human work, that in no way means the model somehow becomes magically more intelligent, it just means the work will stall and continue on at the same level forever
hell even if they hypothetically have an internal model that can output the entire training data set in a better format, there's no scientific evidence that the newer format has new information that is sufficient enough to train a better AI
as a matter of fact the scientific evidence is on the contrary
TeMPOraL 3 hours ago
gytt33 2 hours ago
thmoonbus 4 hours ago
At least they’re led by trustworthy and honest people or we’d need to take their claims with some dose of skepticism.
pvab3 4 hours ago
TeMPOraL 4 hours ago
Yizahi 4 hours ago
TeMPOraL 3 hours ago
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
8note 2 hours ago
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
gytt33 2 hours ago
This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close. Who has the appetite to keep throwing more given they continually need to keep raising fresh money?
woah 4 hours ago
TeMPOraL 4 hours ago
culi 5 hours ago
funnym0nk3y 5 hours ago
f311a 4 hours ago
Aboutplants 5 hours ago
RachelF 5 hours ago
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
handfuloflight 5 hours ago
Not if the genius level IQs take the market share.
fumar 5 hours ago
drewnick 3 hours ago
jppittma 5 hours ago
msy 5 hours ago
pvab3 4 hours ago
usef- 3 hours ago
People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.
vb-8448 5 hours ago
verdverm 5 hours ago
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
Rzor 5 hours ago
Iolaum 5 hours ago
esafak 5 hours ago
verdverm 5 hours ago
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
zone411 5 hours ago
piyh 3 hours ago
xnx 5 hours ago
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
IX-103 5 hours ago
LightBug1 4 hours ago
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
xnx 4 hours ago
CuriouslyC 4 hours ago
joquarky 3 hours ago
spacebanana7 4 hours ago
xnx 3 hours ago
bluGill 5 hours ago
xnx 4 hours ago
Google is already on gen 8 of its TPUs and is certainly already working on the next version or two.
woah 4 hours ago
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
koe123 4 hours ago
lossyalgo 4 hours ago
jwolfe 3 hours ago
lossyalgo 2 hours ago
xnx 3 hours ago
aforwardslash 2 hours ago
dmix 3 hours ago
zem 5 hours ago
dgellow 4 hours ago
nylonstrung 5 hours ago
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
torginus 5 hours ago
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
arizen 5 hours ago
LarsDu88 5 hours ago
jobs_throwaway 5 hours ago
aleph_minus_one 5 hours ago
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
monkeydust 4 hours ago
redanddead 4 hours ago
There’s no magic there. You get an account executive and a call with a systems architect to find out what you’re doing.
Clouds gonna cloud, this is the reason they rolled deepmind into gcp and arguably the inverse is true, the labs are trying to become clouds
topspin 4 hours ago
That feels right. It's not as if they've been missing out on great profits.
dansquizsoft 3 hours ago
tapoxi 3 hours ago
ErrantX 3 hours ago
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
giancarlostoro 2 hours ago
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
BobbyJo 3 hours ago
I'm gonna need you to look at the capex obligations they've undertaken in the last 12 months. They are definitely not being frugal. If they are behind, its not for lack of spending.
surfmike 2 hours ago
See for example the exodus of talent this year, triggered by mismanagement and politics. They still have a lot of talent but they've lost a lot.
Also, competition with Google Cloud for compute resources, less urgency and focus than the competitors, and (strange to say) not as much user LLM behavior data to feed to RL for coding, work, etc.
But I think they will keep catching up and stay relevant for a good class of LLM use cases.
merb 5 hours ago
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
aleph_minus_one 4 hours ago
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
nozzlegear 4 hours ago
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
brazukadev 3 hours ago
nozzlegear 3 hours ago
merb 3 hours ago
tfsh 4 hours ago
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
redanddead 4 hours ago
krona 4 hours ago
Meanwhile, Anthropic/OpenAI will struggle to survive the next 24 months on their current trajectory.
jeremyjh 4 hours ago
spyckie2 4 hours ago
mattm 4 hours ago
root_axis 4 hours ago
Because they're not desperate. Slow and steady wins the race, at this rate all Google has to do is wait for OpenAI and Anthropic to exhaust themselves on aggressive training, then they can casually amble along right past them.
ivanmontillam 4 hours ago
Traubenfuchs 3 hours ago
brazukadev 3 hours ago
That's actually a feature. We don't need 2 new Googles.
pavlov 3 hours ago
Slow and steady wins the race. How could they lose to something on the web, when Microsoft owns the web browser itself? Everything runs on Windows and IE. They can just relax and wait for competitors to exhaust themselves, then quickly build their own version. Isn’t that how Netscape lost. Etc.
Now Google is the new Microsoft, just like Microsoft became the new IBM.
brazukadev 3 hours ago
guelo 3 hours ago
bamboozled 3 hours ago
pavlov 3 hours ago
When the market keeps expanding, you don’t need to kill your predecessor. The shelf life for legacy enterprise computing is very long.
browningstreet 3 hours ago
overfeed 3 hours ago
They would have been right if Google's marquee product was an Office Suite.
Google was the AI company before AI companies were a thing. The comparison with Microsoft and IBM are misplaced because they failed to capture new territory; ML/AI is Google's stomping grounds. The criticism that Google is bad at consumer chatbots is true, but that's not where the real future value lays.
kursus 3 hours ago
Yizahi 4 hours ago
gandreani 3 hours ago
rstuart4133 3 hours ago
The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.
So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.
lemoncookiechip 3 hours ago
OpenAI and Anthropic are AI business, if the AI market burst tomorrow, they'd be the first to flounder.
Google just has to keep pace in the AI space, they don't have to lead. Especially since whom is leading changes like two or three times per month non-stop for four or so years now, including small (by US standards) Chinese AI labs with a fraction of the money who keep pushing the tech forward every month while being open for now.
aforwardslash 2 hours ago
Technically, they arent even a business. One is a business when the revenue model works. People in IT tend to forget that.
hackernud3s 2 hours ago
throwaway23597 2 hours ago
richardw 2 hours ago
Chinese models are barely behind the leaders. Google can catch up anytime they hit the gas. I think they’re intentionally spending less, and when this crazy race burns out they can play their cards.
aforwardslash 2 hours ago
bluecalm 5 hours ago
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
IshKebab 4 hours ago
throwitaway222 3 hours ago
UncleOxidant 3 hours ago
_fw 3 hours ago
Google should be dominating, but instead all we currently have is a disparate collection of consumer facing apps and a flash model that’s fast, clever and expensive.
Muse and Dots are doing what Google should have brought out last year, with their resources and know-how.
dieortin 2 hours ago
giancarlostoro 2 hours ago
WarmWash 2 hours ago
8note 2 hours ago
the risk of randomly getting my email banned is way too high
ivanmontillam 2 hours ago
That's why I've gone to using open models, they are getting there slowly. A bit much of hand-holding but that's fine by me. If a customer of mine decides to use Google's models, I will have them sign a disclaimer that I'm not responsible of them getting insta-banned or similar. I just can't recommend it.
selcuka 42 minutes ago
computerex 2 hours ago
cco 40 minutes ago
That's too long to be so far behind. This release looks like it puts them back in it but if they don't ship anything again for a year plus it's hard to imagine building on top of them and watching the world go by.
As others note, this is most relevant for us here, Google does not need to chase YC developers and the like. They can move more slowly, they have the size to do that. But it does suggest a lot of dysfunction given they have the world at their fingertips and couldn't seem to ship anything for a year.
giancarlostoro 2 hours ago
pedalpete 2 hours ago
Did they ever really abuse the power Google could have wielded? I could be missing something, but for the most part they seem to just get down to building and pushing technology/science forward and avoid drama rather than welcome it.
OccamsMirror 2 hours ago
WarmWash 2 hours ago
However, I don't think the alternative would have made people any happier (and frankly they probably would be justifiably even more angry)
AlexErrant 2 hours ago
https://grapheneos.social/@GrapheneOS/117282080803799576
> Google should not be gatekeeping security patches to the standard Android platform code from Android OEMs but that's what they've started doing.
The whole Android developer verification program controversy.
These 2 are the recent things that come to mind that are most adjacent to "power-abuse".
yostrovs 12 minutes ago
landdate 30 minutes ago
rayiner 2 hours ago
jasondigitized an hour ago
sfblah an hour ago
not_a_bot_4sho 29 minutes ago
(Aside: I didn't appreciate the revelation!)
motoboi an hour ago
landdate 32 minutes ago
Gemini wins and I said this since the beginning. Google wins in general. I never understood why they have not been the highest market cap companh for the last 10 years. And I despise google but its obvious.
SecretDreams 5 hours ago
esafak 5 hours ago
qgin 5 hours ago
pianopatrick 5 hours ago
pixl97 4 hours ago
Already businesses that have more compute and access to data seem to eat the world around them. If, and ya its and if, we can make something that self learns into RSI it's not looking like any business that came before this.
pianopatrick 2 hours ago
If there are multiple they will cost money to run. In that world I expect there to be a correlation between costs and quality, i.e. the highest quality AI system will likely cost more to use than a lower quality AI system because there will likely be more compute required and so on.
So in that world, the absolute top tier best in the world frontier AI will not actually be the most used system. This is for the simple reason that such a system will be more costly than a lesser tier system that can do the job just as well.
koe123 4 hours ago
Whoever builds the deathstar wins!
Heidaradar 5 hours ago
tripleee 5 hours ago
sixo 5 hours ago
randallsquared an hour ago
pvab3 5 hours ago
nater5000 4 hours ago
laybak 4 hours ago
danny_codes 3 hours ago
peddling-brink 2 hours ago
margalabargala 2 hours ago
lantry 2 hours ago
https://en.wikipedia.org/wiki/Synanthrope
https://en.wikipedia.org/wiki/Category:Species_made_extinct_...
FinnKuhn 4 hours ago
koe123 4 hours ago
fittingopposite an hour ago
dgellow 4 hours ago
pkfz 4 hours ago
FinnKuhn 4 hours ago
Gigachad 2 hours ago
This seems to match what we are seeing where Chinese models from companies with only a tiny fraction of the compute are able to be hot on the heels of the frontier models.
kushalpandya 4 hours ago
bitpush 4 hours ago
dgellow 4 hours ago
ddp26 4 hours ago
This is marketing from Google, not a competitive offering
grababner 4 hours ago
netdur 3 hours ago
cyanydeez 3 hours ago
pyaamb 3 hours ago
aurareturn 2 hours ago
dgemm 2 hours ago
_heimdall an hour ago
echelon 25 minutes ago
Or maybe everyone makes money for everything.
The world could turn into a world of plenty. Or it could become a YouTube popularity contest where the MrBeasts get to eat and nobody else is interesting enough to sell themselves.
This is an absolutely crazy time to be alive and most people still don't see it.
babelfish 6 hours ago
Gemini not beating the "can't release a model" allegations
modeless 6 hours ago
ionwake 5 hours ago
ok bro thx
Androider 5 hours ago
vlyan 5 hours ago
forshaper 3 hours ago
AuthAuth 5 hours ago
XzAeRosho 5 hours ago
asdfasgasdgasdg 4 hours ago
ttul 4 hours ago
wasabi991011 4 hours ago
What is your reason to believe this is not a bug specific to a small set of Pro users?
bobtheborg 4 hours ago
My free gmail account is not.
fragmede 3 hours ago
HotHotLava 3 hours ago
rahimnathwani 4 hours ago
3.5 Flash-Lite
3.8 Flash
3.1 Pro
In both a paid Google Workspace account (without the AI addon) and a free 'GSuite' account, I see: 3.6 Flash
3.6 Thinking
3.1 Pror1ch 3 hours ago
giancarlostoro 2 hours ago
Yizahi 3 hours ago
bakugo 5 hours ago
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
A_D_E_P_T 5 hours ago
brainwad 5 hours ago
IX-103 5 hours ago
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
A_D_E_P_T 4 hours ago
fooker 5 hours ago
ryandrake 4 hours ago
jstummbillig 5 hours ago
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
kevinh 5 hours ago
jstummbillig 5 hours ago
cmrdporcupine 5 hours ago
mrieck 5 hours ago
I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.
giancarlostoro 2 hours ago
Culonavirus 4 hours ago
tazjin 6 hours ago
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
ChickeNES 6 hours ago
baq 6 hours ago
timmg 6 hours ago
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
qalmakka 6 hours ago
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
Mond_ 4 hours ago
I don't think Carbon is dead, it just all depends on how easy it actually is to rewrite "all of C++" in Rust. (The jury is still out on this one, but it's not looking good.)
YuechenLi 5 hours ago
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
pornel an hour ago
> Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.
So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon's stopgap.
vovavili 6 hours ago
Maxatar 6 hours ago
gorbot 6 hours ago
boshalfoshal 6 hours ago
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
computerdork 5 hours ago
boshalfoshal 5 hours ago
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
chis 5 hours ago
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
Paracompact an hour ago
Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.
I predict that in the next decade: (1) the market will understand the difference between a "code writer" and a "spec writer," with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between "static typing" and "interactive proof assistant," perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I'm more confident in the first three.
tclancy 2 hours ago
That feels like a really strong conclusion. I’m not clear on why any safeguard isn’t a useful safeguard if you let agents write all the code.
mike_hearn 5 hours ago
They did that for Go and it seems to have worked out for them though.
lesuorac 5 hours ago
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
torginus 4 hours ago
Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).
Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.
Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.
fg137 6 hours ago
bvinc 6 hours ago
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
minimaxir 5 hours ago
culi 5 hours ago
LarsDu88 5 hours ago
bitexploder 5 hours ago
rafram 5 hours ago
nchie 4 hours ago
Karrot_Kream 2 hours ago
hbbio an hour ago
And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there's no magic so it's easy for devs and agents to reason about.
If you're curious: https://github.com/okcontract/oksolc
lossolo 4 hours ago
throwitaway222 3 hours ago
adamrezich 5 hours ago
6thbit 5 hours ago
DrBenCarson 3 hours ago
https://www.ll.mit.edu/r-d/projects/translating-all-c-rust-t...
pshc 5 hours ago
Gigachad 2 hours ago
I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.
6thbit 5 hours ago
Imagine they aren't even familiar with rust but are deeply familiar with the product.
mlmonkey an hour ago
So ... Rust still can't beat the C++ implementation :-D
Sorry, didn't mean to ignite a langwar, but it's still interesting to see.
juanre 2 hours ago
In order for the benefits of AI to be distributed, intelligence has to become a commodity.
As long as you control the skills, the learnings, and the infrastructure setup you will be fine.
copperx 2 hours ago
fractorial an hour ago
And, no, wrapping claude -p or any other “allowed” use wherein you don’t control the agent loop isn’t the same.
onlyrealcuzzo 20 minutes ago
small_model 2 hours ago
uvdn7 5 hours ago
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
SwellJoe 5 hours ago
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Zagitta 5 hours ago
[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p27...
izacus 4 hours ago
SwellJoe 3 hours ago
This year, maybe next, maybe a year or two after that, is probably the most C++ lines of code in production use there will ever be. Why would one choose C++ for new projects at this point? There are niches where Rust is still uncomfortable or just doesn't have the support, but not for much longer. Models are very good at Rust and good at porting to Rust. And, Rust is a good language for models because it is so strict...it helps keep them in line.
tonyhart7 5 hours ago
mattlondon 5 hours ago
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
onlyrealcuzzo 4 hours ago
exacube 17 minutes ago
randomperson321 3 hours ago
npalli 5 hours ago
gavin_gee an hour ago
mhils 4 hours ago
(Full disclosure: I am one of the coauthors)
Revanche1367 5 hours ago
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
murkt 4 hours ago
aqme28 4 hours ago
onlyrealcuzzo 4 hours ago
brokencode 4 hours ago
That's a lot already and growing quickly.
gravypod 2 hours ago
applfanboysbgon 2 hours ago
wg0 5 hours ago
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
wasabi991011 4 hours ago
So while this announcement has no details about the quantum algorithm optimization, I feel fairly confident that it will hold up.
asdfman123 4 hours ago
A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!
kridsdale1 3 hours ago
But it still takes 2 weeks to get a CL approved and past TAP.
asdfman123 3 hours ago
mridulmalpani 5 hours ago
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
5555watch 5 hours ago
mridulmalpani 5 hours ago
I am just trying to understand - why Google haven't done and have no plans for it. They have done it for Android and have Gemma models too.
sidibe an hour ago
loufe 4 hours ago
aviinuo 4 hours ago
tantalor 5 hours ago
losvedir 4 hours ago
bitexploder 3 hours ago
I suspect most people don't have enough storage space to even download a frontier model.
freecodeio 3 hours ago
tandr 3 hours ago
gopalv 6 hours ago
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
polotics 6 hours ago
janustimes 6 hours ago
So no, Google is not being punished, nor are they the people behind this technique.
bananaflag 5 hours ago
afthonos 2 hours ago
loufe 4 hours ago
lukewarm707 4 hours ago
they use a small model to make fake chain of thought and return that.
google has access to the real chain of thought.
arjunchint 6 hours ago
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
pfooti 5 hours ago
Aboutplants 5 hours ago
KeplerBoy 5 hours ago
xnx 4 hours ago
6thbit 5 hours ago
It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.
vdfs 2 hours ago
ncruces 4 hours ago
sajithdilshan 4 hours ago
bfung 2 hours ago
kylecazar an hour ago
Maybe announcing it was delayed until the agreement, and releasing it was delayed to show compliance.
summerlight 2 hours ago
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
elAhmo 6 hours ago
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
deanc 5 hours ago
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Sidio 4 hours ago
I worry Google's scale and laundry list of internal stakeholders means they will never a simple unified harness.
wasabi991011 3 hours ago
3.8 flash has been perfectly available to Pro users from the announcement day.
You are on the cheapest paid plan and complaining that you don't have access to more expensive models. (Though idk why you only have the lite version, I have access to the non-lite version even on the free gemini plan as long as I'm logged in.)
novafunc 3 hours ago
That's strange since as Plus, you still get access to latest Pro (albeit 3.1) but not the latest Flash.
I would understand if I Plus subscribers still got access to 3.7/3.8 Flash, but it just burned through the usage limits faster.
dieortin 2 hours ago
iamronaldo 6 hours ago
LucasBrandt 6 hours ago
h14h 6 hours ago
tonyhart7 5 hours ago
3371 5 hours ago
ehsankia 5 hours ago
denysvitali 5 hours ago
kingstnap 4 hours ago
There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.
jofzar an hour ago
onlyrealcuzzo 4 hours ago
Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.
asdfman123 4 hours ago
skavi 5 hours ago
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
jeffbee 5 hours ago
skavi 3 hours ago
wstrange 2 hours ago
computerdork 5 hours ago
IX-103 4 hours ago
skavi 3 hours ago
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
[0]: https://fuchsia.dev/fuchsia-src/concepts/kernel/zx_and_lk
ismailmaj 4 hours ago
GodelNumbering 5 hours ago
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
jbl0ndie 3 hours ago
Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?
SwellJoe 6 hours ago
greenchair 6 hours ago
zem 5 hours ago
jastanton 6 hours ago
wasting_time 5 hours ago
SwellJoe 5 hours ago
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
wasting_time 4 hours ago
The model is not available yet, so Google is essentially saying "trust me bro".
TeMPOraL 5 hours ago
blueaquilae 6 hours ago
hn_acc1 5 hours ago
thefourthchime 5 hours ago
formvoltron 5 hours ago
ducktoysleftout 4 hours ago
fitzn an hour ago
GalaxyNova 8 minutes ago
iamhaseeb 17 minutes ago
bottlepalm 6 hours ago
colordrops 6 hours ago
Scrapemist 6 hours ago
fer 5 hours ago
NiloCK 5 hours ago
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
rhaff 5 hours ago
jackkinsella 5 hours ago
bottlepalm 4 hours ago
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
unbrice 3 hours ago
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
bottlepalm 2 hours ago
kelvinjps10 5 hours ago
schmookeeg 5 hours ago
wg0 4 hours ago
unbrice 3 hours ago
bottlepalm 5 hours ago
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
yacthing 5 hours ago
NiloCK 5 hours ago
bottlepalm 4 hours ago
fragmede an hour ago
tiahura 5 hours ago
corford an hour ago
rsstack 6 hours ago
Rzor 5 hours ago
polotics 6 hours ago
eamsen 6 hours ago
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
mattkevan 5 hours ago
Many other models get things wrong, but Gemini is the only one to go on the defensive.
aNapierkowski 5 hours ago
RachelF 5 hours ago
Hamuko 5 hours ago
abixb 5 hours ago
bottlepalm 4 hours ago
schainks 4 hours ago
rdtsc 3 hours ago
I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".
chaostheory 3 hours ago
modzu 29 minutes ago
jjcm 6 hours ago
nurettin 6 hours ago
mydreamof 5 hours ago
xnx 4 hours ago
nonethewiser 5 hours ago
sebzim4500 5 hours ago
moostii 3 hours ago