Gemini 3.8 Flash and 3.8 Flash Cyber (blog.google)
simonw 12 hours ago
Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":
https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f
Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
pietz 12 hours ago
wayeq 12 hours ago
and probably a barely modified knock-off of some github project that it trained on
sawjet 11 hours ago
superze 11 hours ago
snet0 11 hours ago
simonw 11 hours ago
ChickeNES 10 hours ago
whateveracct 10 hours ago
slopinthebag 11 hours ago
hglaser 12 hours ago
simonw 12 hours ago
giancarlostoro 12 hours ago
I would hope the people who make one of the most used JS engines in the world are capable of making a model good at JavaScript ;)
heliosAtwork 11 hours ago
bermudi 11 hours ago
Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny".
Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.7 putting it dead last in output tokens per task in the leaderboard.
WarmWash 10 hours ago
gundmc 10 hours ago
https://artificialanalysis.ai/#cost-tabs
That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.
NicoJuicy 8 hours ago
Ps. For the last week I diverged to Luna too, still need to check 3.8 flash.
But 3.6 flash was my go-to model 3 weeks ago and before it was deepseek flash/pro for a while.
None of the claude models seemed cost effective though.
criley2 6 hours ago
Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis.
Luna high is literally 30X cheaper than Gemini 3.8 flash high.
You can limit the model viewer and they're getting better at testing multiple effort levels now: https://artificialanalysis.ai/?models=gpt-5-6-sol-medium%2Cg...
One reason is clear: Sol uses dramatically fewer output tokens than Gemini 38 flash https://artificialanalysis.ai/?models=gemini-3-8-flash%2Cgem...
simonw 11 hours ago
Since this transcript has HTML in it, I decided to upgrade that tool to also render HTML.
I set Gemini 3.8 Flash the task, using my own VERY shonky coding agent tool (llm-coding-agent) - and it did a solid job.
So now you can see the "cool thing in html" rendered within the Markdown document using code that Gemini 3.8 Flash also wrote: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Transcript where it built that is here: https://gist.github.com/simonw/3e36b98292dfdc1b3baff158faa74...
jauntywundrkind 10 hours ago
i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-line with that split-of-focus.
estetlinus 10 hours ago
User: use them both
Made me giggle.
badlucklottery 10 hours ago
I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.
kridsdale1 10 hours ago
noir_lord 10 hours ago
Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.
senordevnyc 9 hours ago
Yeah, I know, just more slop. But I do think the second agent’s eagerness to please is aligned more in your favor in that instance, so it’s likely to find most issues.
The bigger problem I’ve found is that it’ll also find all kinds of very minor edge cases that you have to pick through.
Forgeties79 9 hours ago
They argue the net is positive but clearly the “100x productivity multiplier” claims have been dashed on the shoals of reality for these groups.
This is anecdotal, but it’s across the board in my vicinity. I’m curious how common this is and if it’s just “the new normal” to adopt the nauseating Covid phrase.
briHass 6 hours ago
These types of high-level tests are frustrating beyond belief to humans due to their lack of specificity, but with the agents, they don't get annoyed investigating possible regressions from non-specific signals.
They also aren't as painful to maintain as one would think, because a regression flagging test can be traced by the agent and represented as the business rule that was violated. I've found recent models to be really excellent at discerning a true regression from an outdated test assertion, especially if they are able to trace the failing test back to the PR and work ticket that built it.
agumonkey an hour ago
noir_lord 9 hours ago
Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc.
We have the folks over here saying "AI is amazing" and the other other folks over there saying "AI is terrible".
I've largely sat it out so far and I listen to both camps (and people in the middle as well) and I keep half an eye on what they are up to (including periodically evaluating them) but my overarching impression is still "Why would we trust this when it hasn't shown it's trustworthy?"
iterateoften 8 hours ago
Obviously maybe it’s not composable like that exactly in real world but that’s the intent of agents checking agents
senordevnyc 7 hours ago
But I wouldn’t say I “trust” these agents. The degree to which I double check their work depends heavily on the consequences if it gets something wrong. Not too dissimilar from another human dev in that sense.
So for the SaaS that supports my family, there are some things I have it build where I glance at the PR for a minute or two, but if it broke something on this admin page that only I see, there’s no real downside and I’ll find out pretty quickly next time I use it. And it’s fine 95% of the time, so it doesn’t feel like the best use of my time to double-check it carefully.
But for some of the complex internal flows where a bug could be both catastrophic and difficult to even discover for awhile, I still check it very carefully.
For a little one-off vibe coded demo thing like OP shared, I wouldn’t look at the code at all, I’d just have another agent check it and fix anything it finds. Very low stakes.
Qworg 7 hours ago
riversflow 6 hours ago
dr_kiszonka 4 hours ago
jasongill 7 hours ago
cheikhcheikh 7 hours ago
cyrilng an hour ago
skybrian 8 hours ago
estearum 3 hours ago
In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?
luipugs an hour ago
sedgjh23 an hour ago
nozzlegear 32 minutes ago
trvz 10 hours ago
w4zz 9 hours ago
ericol 10 hours ago
silasdavis 9 hours ago
> Aside from reading identically forwards and backwards down to the letter
No it doesn't.
aidos 9 hours ago
tomjakubowski 7 hours ago
When I typed "Are we not pure noon, ergo, we play life; yet, we hate bad fear" into Google, I got more weird results from Gemini: it claimed, incorrectly, that it is an anagram of the "well-known philosophical statement" (?), "We are not pure nature, we are history".
https://share.google/aimode/wJosKnHig6oVYaG18
(?): the reference seems to be to Jose Ortega y Gasset's line, "El hombre no tiene naturaleza, lo que tiene es historia" -- "Man[kind] has no nature, what it has is history."
Kayou 9 hours ago
I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token. Here is the result https://coolthing-ling-3-tiny.tiiny.site (sorry for the weird hosting, first I found that worked)
(it cost me almost 0 cents and done in 49 seconds)
embedding-shape 6 hours ago
Datasets contains lots of people sharing particles simulations in various ways, with a bunch of people replying "that's so cool" and similar, so 10 years later someone asks an LLM for "cool thing" and "particle simulations" rank pretty far up when it thinks about what others have called cool.
wyrdcurt 9 hours ago
For comparison's sake, I tried something similar with a couple other cheap models I've used lately, with the prompt "Impress me. Make something cool in HTML. Ensure that it is mobile friendly." (Added the mobile condition as I was on my phone when I did it).
Mimo-2.5 created something similar, only a bit less complex than Gemini's (though, at least the FPS counter is real!), in a minute or two for about 1/3 of a cent: https://gisthost.github.io/?740c325c21e9bfbee59c4f94d9aab0af
GLM-5.3-Flash, currently my workhorse model, spent 12 minutes (ouch) thinking about the prompt. Didn't cost me anything directly because I have a GLM sub, but I did the math and it would have cost about 1.1 cents through the API. Turned out nicely in my opinion (though in reality, it still isn't really anything special): https://gisthost.github.io/?9ef050e16cec2561e6504e725a3f0bcc
Side note: thanks for setting up that Gist Host tool, it's very convenient!
---
Editing to add this bonus from Mercury-2.5-Preview, which I just learned released a couple days ago. It's much less impressive-looking than any of the above, but it cost less than 1/20th of a cent, and the response was generated effectively instantly: https://gisthost.github.io/?02f40b50aa891bf396bfaaa3a7998203
meerita 9 hours ago
dennis16384 9 hours ago
walrus01 an hour ago
Thought processs: "Oh, simonw is asking me to make something cool, I think I know what he really wants..."
jampa 13 hours ago
- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.
- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.
- Document parsing (extracting the relevant trip info from PDFs).
If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.
tziki 13 hours ago
jampa 13 hours ago
trial3 13 hours ago
dymk 12 hours ago
jampa 12 hours ago
BeetleB 12 hours ago
2. I'd wager the majority of HN commenters don't read their own comment before posting (pre-LLM days).
fc417fc802 10 hours ago
BeetleB 9 hours ago
drusepth 11 hours ago
See: why authors wait days, weeks, or even months before editing what they've written (or, if you're more interested: cognitive regression, inattentional blindness, and the effects of misdirected saccades).
dymk 8 hours ago
drusepth 3 hours ago
I read this to mean he wrote the comment, then asked Claude to fix the grammar (as many ESL speakers do). Sounds to me like he did write it.
jamiek88 5 hours ago
handzhiev 12 hours ago
owlninja 12 hours ago
aero142 12 hours ago
greenavocado 11 hours ago
sneezychl 11 hours ago
Sounds counter-intuitive at first, but Luna is overall better at sticking with what works. Sol is wicked smart but needs constraints.
handzhiev 11 hours ago
jesuslop 10 hours ago
kelvinjps10 2 hours ago
greenavocado 12 hours ago
handzhiev 11 hours ago
greenavocado 10 hours ago
trvz 10 hours ago
handzhiev 7 hours ago
qlte 3 hours ago
Subscription? -> Google One plan (http://one.google.com/)
API? -> AI Studio (https://aistudio.google.com/)
It's not really any different than the choice you'd make with OpenAI/Anthropic depending on how you plan to use it. Except as a hyperscalar, it's also offered first party from Google Cloud (like Claude via Amazon Bedrock or GPT via Microsoft Azure OpenAI Service): Google Cloud -> Gemini Enterprise AI Platform (https://cloud.google.com/ai)
But if you're using models via OpenCode or Pi or whatever, the flow chart is basically just "Go To AI Studio" unless you or your employer is already used to Google Cloud, otherwise there's no need to subject yourself to all those enterprise-y IAM dashboards and stuff. You still get free usage from AI Studio when you generate the API key without needing to add billing details so very easy to try.-0_0- 20 minutes ago
seanmcdirmid 5 hours ago
rahimnathwani 12 hours ago
Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?
mlmonkey 12 hours ago
rahimnathwani 12 hours ago
plaidfuji 12 hours ago
panarky 12 hours ago
jampa 12 hours ago
It was right on every nit, so it was surprising how well the model knows these things. If I ever release this I'll probably need the SERP API or Google Maps SDK (which I've heard is very expensive now), but for a personal trip where I will verify manually, using the LLM is okay for now.
rahimnathwani 12 hours ago
tools=[{"type": "google_search"}]
I'm curious whether in fact you were getting answers from the model weights (which is what I had assumed) or whether your API calls were resulting in web search tool calls.kridsdale1 10 hours ago
Using grounding in Gemini is indeed backed by the same canonical data source for business information (like opening hours) as Google Maps. This stuff is available in its own API for a GCP fee, but we’ve built tooling to connect it to the Gemini agentic ecosystem as well.
NiloCK 4 hours ago
The info isn't in the model weights.
Because of where I live, there are three viable airports for any given flight I might want to take, which historically has made shopping a real pain. But Gemini (and only Gemini) has greatly simplified it. Pramble plus date range plus destination and it very quickly generates potential itineraries with costs, total travel time (driving included), etc.
BlackRabbit1 12 hours ago
porridgeraisin 12 hours ago
BlackRabbit1 11 hours ago
I've been planing around with LLM-based trip planning for a very long time now as it fits my very ad hoc style of traveling very well.
But distances always had been.. lets say.. difficult.
Will test it with my upcoming trip to Greece then!
colechristensen 12 hours ago
Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.
re-thc 12 hours ago
Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.
colechristensen 12 hours ago
Every time I personally tried Gemini models up until last week they simply couldn't do the long complex tasks I'd being doing with Anthropic models for many months.
cousinbryce 10 hours ago
throwaway219450 9 hours ago
spacebanana7 8 hours ago
Gemini’s integration with maps and search is more important for Google.
dismalaf 12 hours ago
For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.
altmanaltman 12 hours ago
so far
dismalaf 11 hours ago
Also them having their own silicon means they don't have to pay the Nvidia tax and can keep costs a lot lower.
kridsdale1 10 hours ago
fc417fc802 10 hours ago
For example find a beautiful landscape shot of a place that just so happens to be accessible to tourists and ask it something along the lines of identifying the location. IME it will noticably steer the conversation towards relevant commercial offerings and offer (entirely unprompted) to help plan a trip.
Or ask it about a certain category of product with some requirements and it will initially present (relevant) options that look like paid placement to my eye. But if you ask it's happy to go on to turn up lots of alternatives and enumerate tradeoffs.
Assuming I'm correct the subtlety is on par with product placement in movies. Certainly leagues better than the internet advertising we've suffered to date.
rstuart4133 8 hours ago
As you say it was subtle, along the lines of "oh, if you are planning on going to the place you are researching, here are some helpful links to places you can stay". Subtle, in that it didn't get in the way of main result, so I didn't mind overly. Insidious, as I only noticed because I wondered why it was providing those particular links and looked them up. I can't see how you could ad-block them if I did object.
And worrying, because these unblockable sneaky ads are just a first foray coming from a company that prostitutes its own app store searches, by making the first and most obvious result utterly unrelated to to the search topic. Instead it's who paid them the most to be there. That behaviour is why everyone dumped Alta Vista when an alternative came along. Alternative Android app stores can't come soon enough.
They already skim off 15% of purchases which I'm sure makes their Android operation return a profit that makes other industries drool. Debasing their search to ad a tiny bit extra on top must by driven pure greed. Senseless, as I'm sure it will come back to bite them in the end.
dominotw 12 hours ago
this has to be stong suit of ai agents any model
robotmay 12 hours ago
My only wish is it were somewhat cheaper, as it tends to balloon pretty quickly when I'm using it in Opencode. I'm currently trying to offload a lot of work to subagents to stop the context expanding so rapidly. But on the upside, I rarely have to correct it - I've spent far less time arguing with this than with anything else so far.
newtwentysix 12 hours ago
CamilleScholtz 9 hours ago
leokennis 9 hours ago
Think 2023 style ChatGPT. Something like “to open a document on your Mac click File > Open docurrrar” - like it suddenly forgot it had to produce actual words.
Overall I enjoyed its speed and comprehensiveness. But those occurrences of nonsense just made it feel like a great car that once a month just stops in the middle of the highway.
Shayk 3 hours ago
dassh 3 hours ago
forlorn an hour ago
mattlondon 13 hours ago
https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!
Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
Gecko4072 13 hours ago
oceanplexian 13 hours ago
roosterIllusi0n 12 hours ago
nolok 10 hours ago
WarmWash 13 hours ago
scrlk 13 hours ago
ford 13 hours ago
Not sure on consumer/product use though
scrlk 12 hours ago
sotix 6 hours ago
ttul 13 hours ago
pietz 12 hours ago
It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.
re-thc 12 hours ago
It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
ttul 10 hours ago
satvikpendem 13 hours ago
onlyrealcuzzo 13 hours ago
So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.
NitpickLawyer 13 hours ago
sunaookami 13 hours ago
...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.
onlyrealcuzzo 13 hours ago
It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.
Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.
harmonic18374 12 hours ago
onlyrealcuzzo 11 hours ago
nolok 10 hours ago
markasoftware 13 hours ago
Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.
irishcoffee 13 hours ago
These folks must laugh themselves to sleep. This whole industry hoodwinked the masses. It’s impressive.
wonnage 13 hours ago
theHocineSaad 13 hours ago
With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.
Squarex 12 hours ago
pietz 12 hours ago
mattlondon 12 hours ago
Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?
asdfologist 12 hours ago
WarmWash 10 hours ago
Topfi 8 hours ago
Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.
It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.
WarmWash 5 hours ago
https://arxiv.org/html/2604.24827v1
The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.
nomel 5 hours ago
duplessitous 10 hours ago
Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.
knollimar 6 hours ago
porphyra 9 hours ago
kamranjon 12 hours ago
anthonyrstevens 12 hours ago
bertili 13 hours ago
abirch 13 hours ago
panarky 12 hours ago
Then I tell Opus to read the audit report and implement what it agrees with.
Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus.
Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor.
porridgeraisin 12 hours ago
In india, my telco gives me google ai pro for free. And agy with flash goes a long way.
prodigycorp 12 hours ago
MaxikCZ 12 hours ago
pkos98 13 hours ago
jrflo 12 hours ago
notatoad 12 hours ago
anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.
kimjune01 12 hours ago
WhitneyLand 12 hours ago
For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
simonw 13 hours ago
Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents
(I think thinking level low is a regression on 3.8 compared to 3.7.)
world2vec 13 hours ago
simonw 13 hours ago
(Next up is the comment saying that the labs are clearly training for the benchmark.)
world2vec 13 hours ago
WarmWash 12 hours ago
_puk 10 hours ago
wongarsu 13 hours ago
I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA
In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)
IshKebab 7 hours ago
If only it were true that things that are tiresome are unpopular. But witness "6 7", "first post", ... remember the "in soviet Russia" jokes on Slashdot"? It seems like there are a subset of people that simply don't get tired of tiresome things.
bitexploder 13 hours ago
anentropic 12 hours ago
onlyrealcuzzo 13 hours ago
> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
So 50x cheaper - and how much faster?
simonw 13 hours ago
dieortin 13 hours ago
simonw 12 hours ago
JacobAsmuth 11 hours ago
simonw 10 hours ago
https://twitter.com/sunjiao123sun_/status/202455551655137292...
> I’ve been developing the SVG generation capabilities for Gemini 3.1, and the complexity of the SVGs is stunning.
> This allows UX designers to transcend pixel constraints and directly output structural, production-ready code!
scosman 11 hours ago
MadameMinty 10 hours ago
isoprophlex 10 hours ago
phatfish 8 hours ago
uif124 11 hours ago
It saw the fish in the basket from some other previous attempt but completely missed the gap between the tires and the rims where the background shines through (now it knows after scraping this comment and watch the next transcript).
lern_too_spel 12 hours ago
jpadkins 12 hours ago
Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.
neuronic 7 hours ago
hughw 12 hours ago
mrdependable 12 hours ago
trentor 11 hours ago
aesthesia 11 hours ago
nonethewiser 10 hours ago
neuronic 7 hours ago
EugeneOZ 8 hours ago
anigbrowl 6 hours ago
simonw 5 hours ago
simonw 13 hours ago
Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.
Matsta 11 hours ago
We transcode everything to 480p before we send it to Gemini batch api. Works great
drusepth 10 hours ago
True multimodal support would be way better, but I have no issues pasting in full screen recordings while QA'ing games and having Claude identify and fix issues in the video.
ray_kay777 9 hours ago
WarmWash 10 hours ago
https://blog.google/innovation-and-ai/models-and-research/ge...
brap 11 hours ago
These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).
mvdtnz 11 hours ago
brap 11 hours ago
_aavaa_ 10 hours ago
watusername 9 hours ago
To preempt certain replies, yes, I know you can pay API prices and use whatever harness you want.
_aavaa_ 9 hours ago
Imanari 10 hours ago
foretop_yardarm 9 hours ago
drusepth 10 hours ago
VadimPR 8 hours ago
KeplerBoy 8 hours ago
drusepth 6 hours ago
NickHirras 2 hours ago
brap 5 hours ago
robertn702 9 hours ago
pdimitar 8 hours ago
solarkraft 7 hours ago
OpenCode has “providers” for many (many!) other services, but these are almost all unofficial and against ToS (Anthropic being famous for ban-hammering people).
cromka 7 hours ago
robertn702 5 hours ago
ryanscio 9 hours ago
Or choose Oh My PI [2] for batteries included
[1] https://github.com/earendil-works/pi [2] https://github.com/can1357/oh-my-pi
dcchambers 8 hours ago
mredigonda 8 hours ago
tobias2014 7 hours ago
simlevesque 4 hours ago
smlx 2 hours ago
ddxv 30 minutes ago
allthetime 15 minutes ago
EFLKumo 11 hours ago
cubefox 11 hours ago
EFLKumo 11 hours ago
cubefox 10 hours ago
adleyjulian 11 hours ago
SchemaLoad 5 hours ago
livinglist 5 hours ago
BeetleB 9 hours ago
> Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.
Ditto for pets.
cubefox 8 hours ago
For animals he/she does make sense, because they are male or female. An LLM is neither.
fwip 2 hours ago
literallywho an hour ago
SneakyZero 9 hours ago
mattlondon 13 hours ago
I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!
At this point it is a meme of course, but where is 3.5 Pro :)
meetpateltech 13 hours ago
hiddencost 13 hours ago
aurareturn 3 minutes ago
Sometimes it is hard for a scientist by nature to build and iterate and lead revenue generating products.
p_l 9 hours ago
neuronic 6 hours ago
a11r 13 hours ago
Jcampuzano2 13 hours ago
Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything.
I think I actually preferred it when we had models that either had reasoning enabled or didn't.
abixb 10 hours ago
One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.
Small models cataching up with their bigger siblings are fantastic news.
alephnerd 10 hours ago
A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.
kamranjon 12 hours ago
I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.
Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.
1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...
film42 12 hours ago
kamranjon 12 hours ago
film42 10 hours ago
pampas 5 hours ago
andai 13 hours ago
ipsod 13 hours ago
But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".
Flash is my go-to for prototyping, and basically anything that isn't writing production code.
ramon156 13 hours ago
ipsod 13 hours ago
Diederich 11 hours ago
1. Data. Lots of data.
2. Money. Lots of money.
3. Access to necessary hardware.
4. Business alignment/will to do it.
5. Access to talent, current and future.
This is certainly incomplete/naive. In my mind, though, Google was the clear answer.
On a more personal level, I've been deep into the Google ecosystem since I got [email protected] in 2005. (I actually paid 50 cents on ebay to get a very early invite.) There was no question in my mind that Google's AI work would deeply integrate into their whole ecosystem in very powerful and productive ways. (Yes, I can join you to discuss, at length, the various ways that Google's dominance is problematic/scary.)
Having said all that, I'm quite happy that there is, at the moment, a very rich competitive landscape. Indeed, not too long ago, with Gemini Pro 3.1 languishing, I moved most of my deeper thinking work to ChatGPT, which was, for me at least, clearly outperforming Gemini.
While I certainly didn't anticipate it, Google's strategy of making their fast/relatively inexpensive models surprisingly powerful has been a welcomed surprise.
ipsod 10 hours ago
esafak 13 hours ago
edit: I have a subscription; direct call.
dannyw 13 hours ago
momojo 11 hours ago
MrBuddyCasino 10 hours ago
realist_not 13 hours ago
worldsavior 13 hours ago
MrBuddyCasino 10 hours ago
refulgentis 13 hours ago
zuzululu 8 hours ago
experience tells me that those people simply have not used models for a long period of time specifically on coding and have run their own comparisons
to someone who uses all vendors, the differences are very palpable and drives purchase decisions.
also keep in mind Gemini and other labs have repeatedly done benchmaxxing, you must have your own benchmarks to evaluate these models.
pampas 6 hours ago
j-bu 12 hours ago
Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.
venusenvy47 12 hours ago
j-bu 12 hours ago
npn 10 hours ago
rjh29 9 hours ago
StevenWaterman 8 hours ago
make3 an hour ago
raincole 12 hours ago
Unless they have an even more powerful Gemini Pro in the oven...?
drowntoge 12 hours ago
owaiswiz 11 hours ago
anthonypasq 11 hours ago
3.0 flash -> 3.8 flash is all post training which is pretty impressive.
chrsw 7 hours ago
deaux 2 hours ago
rahidz 11 hours ago
make3 an hour ago
xnx 13 hours ago
fitsumbelay 13 hours ago
tagalog an hour ago
I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.
lysecret 10 hours ago
MrBuddyCasino 10 hours ago
weird-eye-issue an hour ago
throw10920 2 hours ago
tkgally 2 hours ago
throw10920 2 hours ago
meh2frdf 13 hours ago
onlyrealcuzzo 13 hours ago
My experience is that antigravity is awful and reckless - but that the model itself isn't.
upcoming-sesame 13 hours ago
okdood64 13 hours ago
meh2frdf 13 hours ago
meh2frdf 13 hours ago
iAMkenough 13 hours ago
sejje 10 hours ago
wongarsu 13 hours ago
That said, I do trust Opus and Fable enough to let them deploy to staging. Great for debugging. Just don't give them keys for prod
upcoming-sesame 11 hours ago
tiborsaas 13 hours ago
Ridius 12 hours ago
kyrra 12 hours ago
datlife 13 hours ago
jerkstate 12 hours ago
arctic-true 12 hours ago
jerkstate 5 hours ago
throwa356262 12 hours ago
"available to trusted defenders through our new Fairwind Program"
Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle.uif124 11 hours ago
"Valgrind is only available to trusted defenders in our new UnfairAdvantage program"
JacobAsmuth 11 hours ago
129867 10 hours ago
pampas 9 hours ago
maxnevermind 41 minutes ago
1 what is tesla cybercab plan to address legal implications of accident that will happen? who is going to be responsible for them when they happen? are they covered by tesla insurance or some other insurance? are there any official plan/statements around that?
2 what was the name of the experiment they started in san antonio tx when some cars didn't have a driver? what was the results of it? did they expand the operations? it was much smaller than waymo, is it growing? how it is related to robotaxi?
It is not able to connect the dots that I keep asking about Tesla in 2nd prompt and spit out some unrelated stuff. Really? How it can be that bad? Gemini 3.1 Pro model works fine in this case btw. I thought maybe it is about knowledge cut over date and it doesn't know about those events from 2025 but it seems it has the knowledge up to March 2025. Top 10 in Intelligence on artificialanalysis ladies and gentlemen.
adbachman 12 hours ago
Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?
leopoldj 13 hours ago