9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
Getting the most out of Opus 5.5 in Claude and Claude Code (claude.dev)
rdli an hour ago
chewchewchew an hour ago
rdli an hour ago
Tade0 25 minutes ago
rdli 18 minutes ago
(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)
atif089 3 minutes ago
I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.
GroksBarnacles 44 minutes ago
woodruffw 41 minutes ago
hazard 41 minutes ago
Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...
Klathmon 39 minutes ago
CPU time might go up while wall clock time goes down
sdthjbvuiiijbb 38 minutes ago
kgwgk 36 minutes ago
dolebirchwood 34 minutes ago
jghn 15 minutes ago
pertymcpert 12 minutes ago
whatsThisBtn4 44 minutes ago
Pros know these are lower cost models.
oidar 10 minutes ago
tamimio 38 minutes ago
slaser79 23 minutes ago
Betelbuddy 22 minutes ago
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...
karp773 6 minutes ago
Nobody in his right mind will use a Chinese clones when you have models like Opus 5.5 for peanuts.
verdverm 2 minutes ago
let a few valley elites decide how humanity can use this technology
open and transparent is the way, China is showing how
rdli 15 minutes ago
I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.
Waterluvian 12 minutes ago
alwinaugustin a minute ago
kingcauchy 7 minutes ago
It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.
It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.
hibikir 25 minutes ago
So asking it to do things on its own for a long time? Given last week, absolutely not.
istjohn 9 minutes ago
ToJans 42 minutes ago
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
magicalhippo 22 minutes ago
In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.
It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.
I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.
It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.
I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.
I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.
Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.
I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.
TomGarden 11 minutes ago
verdverm 4 minutes ago
the company is run by holier-than-thou, we know what's best... who apparently don't read claude's output and blindly trust it
the mythos "hacking" of the linux kernel, as finally told from the linux side, is eye opening
briga 25 minutes ago
epistasis 23 minutes ago
What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.
Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.
Ah well, the Chinese models will still work well, I guess.
istjohn 5 minutes ago
0. https://github.com/mattpocock/skills/blob/main/skills/produc...
alansaber an hour ago
danbrooks an hour ago
jester997 44 minutes ago
I like to watch it work though because it honestly teaches me some tricks.
whatsThisBtn4 43 minutes ago
Now we are getting downgraded models that do 100x COT because it's cheaper.
tebrun 43 minutes ago
sergiotapia 30 minutes ago
Handy-Man an hour ago
233mhz an hour ago
amelius an hour ago
yfontana 34 minutes ago