I almost want to try adding a rule "Never use the words 'not' or 'instead'."
I-have-ADHD: A skill to stop coding agents from burying the answer (github.com)
ryandrake 16 hours ago
eterm 16 hours ago
Firstly when you've instructed it ( possibly through skills ) not to do something. It'll keep reminding you that it didn't do that. So I might say, "Check out and review this PR, do not make comments on it", and then it'll be keen to point out it hasn't posted comments to the PR.
But more often it happens when it tries one approach, gets itself messed up, and then has to back out that approach, clean up its mess and do something else.
It'll often then spend more time explaining the wrong approach than the right one, which can be frustrating, especially if all its working is buried in the detailed transcripts.
nzach 16 hours ago
You can also ask why did he mentioned something that wasn't done or why he thought this was important.
In my AGENTS.md file I have an instruction telling the agent to never commit any changes unless I explicitly ask for it, and this leads to messages similar to what you just described.
jolt42 16 hours ago
idicjeifjwjd 15 hours ago
jay_kyburz 15 hours ago
FallCheeta7373 14 hours ago
kube-system 14 hours ago
Claude predicts the next token of the predominantly human training input, and humans use "I".
kdkdkwndjekd 12 hours ago
Nice attempt at a “aha, gotcha!” comment, but sadly you’re too off-mark for it to work.
> Claude predicts the next token of the predominantly human training input, and humans use "I".
This is inconsequential. It could very well be programmed to not assume such a personified stance, and yet here we are. Nothing you do makes it drop this ridiculous facade. It’s intentional, not a byproduct.
procone 12 hours ago
kdkdkwndjekd 11 hours ago
I could ask you the same. Are you trying to say AIs cannot be made to prefer behaving in certain ways? Because if so, I’ve got a bridge to sell you.
ygjb 4 hours ago
On it's own a model is just an inert set of data structures.
Plont 5 hours ago
Now we have this software that's specifically designed to mimic humans, and mistaking it for real intelligence or consciousness can easily be disastrous. It's very important that we not anthropomorphize it, but we are catastrophically bad at NOT doing that.
Even our language has had a lot of computer anthropomorphism baked into it ("my phone died!", "this laptop is fussy", "the computer is sleeping", "it's thinking"), and it's not easy to excise that routine anthropomorphism from the way we talk about LLMs.
I don't want an LLM to write as if it were a person because it's definitely easier and more reliable to cut that problem off at the root, as much as possible, rather than to just try to willpower my way out of my human tendency to anthropomorphize inanimate objects.
I'll grant that LLMs talk like this because they're trained on human writing. It may not be possible to get them to not do that. But if it can't be fixed, it's just another thing to put on the "reasons this is all an incredibly stupid idea" pile.
peddling-brink 3 hours ago
You seem to have something against the clankers getting all uppity. And you are welcome to your opinion on whether we have some moral obligation to be nice to them. But if you can solve a real problem with a tool. Is it worth your time to complain about semantics?
Do you get upset when your screwdriver is the wrong color, or has branding that isn’t quite your aesthetic?
If Claude uses I to refer to itself, do you start philosophical arguments with it?
Mildly related tangent. Fable started using my first name today and plastered it all over my docs. “Peddling said this, so based on that I did this.” That did bother me. I told it to just generically call me the user or human. Maybe I’m a bit of a hypocrite here?
pkulak 4 hours ago
ChrisGreenHeur 3 hours ago
teaearlgraycold 14 hours ago
throwawayk7h 13 hours ago
kdkdkwndjekd 12 hours ago
Impersonating a human does nothing to help it solve problems faster, quite the contrary in fact, it has to waste even more time coming up with human-like speech patterns to present the work done.
It shouldn’t assume any personality unless I explicitly tell it to. It is a tool until I tell it otherwise.
tough 5 hours ago
Might want to rethink your statement after going again through what it does exactly. No need to sound more human. but it gives more tokens to "think" like a human committing brain-time to a problem would
Dylan16807 12 hours ago
"I" is normally used for everything. You could be writing from the perspective of a slab of concrete and you'd use "I".
kdkdkwndjekd 12 hours ago
Finder doesn’t ask “Do you want ME to delete this file?”. Photoshop doesn’t ask “Do you want ME to save this file?”. Claude shouldn’t assume itself to be a person either.
> You could be writing from the perspective of a slab of concrete and you'd use "I".
Except this isn’t prose. Claude is not telling me a story from the point of view of a concrete slab. It is assuming personality to present objective facts. If my entire operating system can be interfaced with without it referring to itself as “I” then so can Claude.
Dylan16807 10 hours ago
But I don't think I share your preference either. It's a lot easier to talk about what I decided to do and what the machine 'decided' to do if we attach pronouns. "X was changed" can be too vague.
iddjjedjjwjdje 9 hours ago
At this point you’re just arguing semantics for the sake of it. You know damn well what I mean and if you need more proof that it is perfectly possible and not at all unreasonable to want this just look at most software around you. None them talk like they are a person and those that do are often the most obnoxious and painful to use.
Dylan16807 8 hours ago
I do. And I said to that "Yeah okay". No argument, chill out.
The rest of that line wasn't diagreement, it was explaining why your original comment was confusing.
> None them talk like they are a person
Which I explained with the rest of my post. They're all doing what I ask or automated tasks in a far simpler way. It's almost never unclear whether I did something or my OS did something. But when talking to an AI coding assistant that gets muddy very fast when pronouns are avoided.
sh34r 4 hours ago
drob518 29 minutes ago
Sohcahtoa82 12 hours ago
I'm fine with conversational interfaces using "I". It makes the grammar easier and more clear.
Strong emphasis here on conversational interfaces. I don't want a compiler to say "I ran into an error" or my printer to say "I'm low on paper".
kdkdkwndjekd 12 hours ago
Do you need to point out the finder of the issues? Easy.
“Tool x ran for x amount of time and surfaced these issues…” “Parsing x code surfaced these issues.”
I don’t understand why are people pretending like the English language is incapable of transmitting information without personal pronouns when every program under the sun has always been written to interface with humans in a cold, detached, straight-to-the-point and impersonal way.
Finder doesn’t ask you “I see you want ME to delete these files. Want ME to do that for you?”. Toolbars don’t feature “Create a new file for me” options, terminal utilities don’t report back with “I’ve pattern matched the text you input and here’s the results I’ve found”.
Fr0styMatt88 9 hours ago
Then again, I just dont have an issue with the 'I-isms'; it's a bit weird sure, but at the same time it's a bit more pleasant to interact with as well. After all it is trying to model itself as a person you're talking to.
pkulak 4 hours ago
stavros an hour ago
Tadpole9181 9 hours ago
English has no distinct personal pronoun for an "it". "I" has to be used for grammar to be attributive. There's quite literally no alternative without using passive voice for everything, which is miserable to read and creates ambiguity on if the speaker (it) did something or something happened to have been done, which then requires entire sentences to clarify.
You aren't stupid, you know what it means when it says "I". And it serves a grammatical purpose. You're getting upset at a toaster for ringing a bell to notify the toast is done. "Toasters aren't bell ringers!?"
iddjjedjjwjdje 9 hours ago
Why would you ever think I don’t know personal pronouns serve grammatical purpose? How did you even arrive at that topic? You seem to be missing the point entirely, and I believe quite on purpose given your snarky childish opener.
Entire operating systems stay clear from assuming personality when presenting information or performing actions. Not a single dialogue in my OS refers to itself as “I” when carrying out instructions and reporting back. Why should Claude do it when I don’t want it to and it doesn’t NEED to do so? Why am I not empowered to simply tell it to stop doing that and it obeys? Better yet, why are you thinking yourself on such high horse about this?
> You're getting upset at a toaster for ringing a bell to notify the toast is done. "Toasters aren't bell ringers!?"
If my toaster starts referring to itself as a person, calling out “I made your toast!!” I’ll get mad at it too. I don’t want it to talk or refer to itself as a person. But then again this is not about toasters. This is about AIs being deliberately designed to sound human-like so marketing can lean on the “I” bit of “AI” more heavily and make gullible people think this steroidal information aggregator actually possess the capacity to think and reason, and, consequently, drive sales.
But then again I’d venture a guess that you’re fully aware of all of this, given your opening snidey remark, and are purposefully choosing to be contrarian to be the point of going off on tangents that make 0 sense or have no impact in the discussion whatsoever.
Still, just goes to show how effective this whole thing is in tricking people into thinking it is normal for a tool to think itself a person.
Ignorance, bliss, and all that.
hysan 15 hours ago
If you're using Claude Code, then it's in the harness. At the close of many sessions, I would start a meta conversation over why the LLM would consistently break certain rules. What it found when debugging itself is that some of the "contradicting" rules that I had were in fact, not from my rules. Instead, the instructions from its own harness had phrases telling it to do things like that. When something contradicts, its own instructions would outweigh any custom ones you write. Every rule variant I had tested (including the one that says it overrides the harness instructions - and yes, I've actually tested all the ideas in your comment too) has ultimately been unsuccessful due to this according to the LLM.
vzmax 13 hours ago
mcv 12 hours ago
crisnoble 10 hours ago
jppittma 5 hours ago
tough 5 hours ago
MichaelDickens 9 hours ago
alienchow 9 hours ago
My global CLAUDE.md explicitly states "When commenting on code and configs, or writing MD files, strictly write within the domain of the content being commented on. DO NOT include information, negatives or ramblings from work sessions. For e.g. if commenting on a proto string field that is replacing an int field, do not comment that 'this is not an int field'".
This reduced the idiocy of the agent (Opus 5 included) when writing documents. But I'm still catching it writing README.md talking about the negatives that it removed. Those belong in the memory if it is actually that important (most of the time it's junk), but Claude doesn't seem to understand and never ever learns.
hysan 6 hours ago
pkulak 4 hours ago
dostick 3 hours ago
drob518 36 minutes ago
ricardobeat 11 hours ago
The key sentences to get rid of claude-isms so far:
- say what you have to say and stop
- [no] document-structure signposts
- [no] historical remarks that only warn about past states
- don't attribute agency to things
- never narrate your own changes, fixes, defects from the past, or what the code used to do
It works 100% of the time for other models, 70-80% for Claude, but already makes a big difference.
klabetron an hour ago
100%. I’m working on a greenfield project that’s not yet released. It loves to put comments in code describing what it no longer does or why it misinterpreted something. And then tries to justify it as preventing the same mistakes in the future. Ugh.
schneems 10 hours ago
artdigital 8 hours ago
Funny thing is, Claude often “disagreed with part of the review and decided to not adopt the requested changes” lol
idontwantthis 6 hours ago
pkulak 4 hours ago
anyg 2 hours ago
sleazebreeze 18 hours ago
I don't think we can skill our way out of this one.
klardotsh 18 hours ago
DeepSeek v4 Flash isn’t much better (unsurprising- it’s an extremely stubborn model).
Weirdly, GPT Luna excels at following this type of instruction from AGENTS.md, and never forgetting it, even 400k+ tokens into the context window.
bel8 16 hours ago
GPT Luna tends to keep things objective. Muse Spark 1.3 is also one of the better models in this aspect, for me.
alwillis 17 hours ago
Concise: Claude leads with the result, skips preamble and narration, and
keeps responses short by default, while doing the engineering work as
thoroughly as in the Default style. When you ask for an explanation or
more detail, Claude answers in full. Claude always keeps the complete
content of error reports, security warnings, and confirmations for
destructive actions. Requires Claude Code v2.1.237 or later.
[1]: https://code.claude.com/docs/en/output-styleshungryhobbit 16 hours ago
soontimes 16 hours ago
hinkley 15 hours ago
Which sounds more like Claude has ADHD than the user does.
snerbles 13 hours ago
So if it gets bad I simply tell it "I ain't reading all that, feed it through the STE Gate" and it will tame the results. I haven't bothered to set it up as a hook yet.
klabetron an hour ago
pjerem 2 hours ago
It's weird to be able to use social engineering against a program.
dogscatstrees 21 minutes ago
To get it to fully work though, it has to be asserted at every turn in the session. That reminder in the skill is negligible in tokens but uses extra nonetheless. I hope the scientists at Anthropic can figure this out, it's not simply an output style flag.
jp57 16 hours ago
I genuinely wonder if the people inside Anthropic actually communicate with each other like that. Has it been imprinted with Dario's engrams?
bgilroy26 15 hours ago
8cvor6j844qw_d6 15 hours ago
That said, I think there's a deeper tension here that's worth naming.
alaithea 15 hours ago
charlieflowers 15 hours ago
peaseagee 14 hours ago
jmartrican 14 hours ago
zapkyeskrill 12 hours ago
JesseTG 13 hours ago
EdwardDiego 12 hours ago
lenerdenator 15 hours ago
mcv 12 hours ago
elwell 14 hours ago
lopatin 13 hours ago
smallnix 12 hours ago
ghostpepper 11 hours ago
sleight42 11 hours ago
nosioptar 10 hours ago
Officer — it's not a crime, it's AI induced rage.
jetbalsa 8 hours ago
easterncalculus 10 hours ago
monkpit 8 hours ago
kyleee 7 hours ago
snvzz 9 hours ago
Cider9986 4 hours ago
AI slop blog posts are as bad as ever but the stuff the agents say in the chats don't annoy me much.
MisterMunchkin 15 hours ago
It's fascinating, Sonnet 4 is still available via API and it's so much less moronic than the current model. All of the em-dashes and nonsense are a result of the repeated rounds of reinforcement learning using slop data.
jp57 15 hours ago
visarga 14 hours ago
cruffle_duffle 14 hours ago
trueno 8 hours ago
going back to opus 4.8 and on is literally like talking to the guy who wants to hear his own voice in meetings. going back to 4.6 is actually refreshing, and it feels so much faster. actually gonna laugh if 4.8 and on is so slow because it's draining lakes fighting for its life trying to conjure up this god forsaken persona.
krustyburger 15 hours ago
hinkley 15 hours ago
elwell 14 hours ago
1659447091 13 hours ago
rurban 6 hours ago
EdwardDiego 12 hours ago
adaml_623 10 hours ago
wolvoleo 6 hours ago
Which doesn't say that much about the LLM itself but about the people that make the training material.
kovacs 15 hours ago
jp57 14 hours ago
HeWhoLurksLate 10 hours ago
jachee 15 hours ago
It made it write more like a dev than a marketing agent.
vorticalbox 15 hours ago
mcv 12 hours ago
shermantanktop 10 hours ago
dominotw 15 hours ago
hinkley 15 hours ago
Like how theme parks generally don't do much to keep queues short (or Disney charges you a premium to skip the queue)
creesch 15 hours ago
Considering how much of the input must me nonsense SEO bullshit articles and blogs that only serve to promote a person or company that might be a factor.
I also often have wondered if it is also targeting those same people. Certainly with tools like deep research options (not just Anthropic's offering) the result report seems to be aimed at management, aiming to look impressive while talking around the results.
hinkley 11 hours ago
Here's a simple recipe for deviled eggs with only four ingredients.
My great grandmother was born during the Great Depression. They valued foods that could be made with cheap ingredients.
[four paragraphs later]
Start with 8 hardboiled eggs...
saghm 15 hours ago
gitowiec 14 hours ago
lubujackson 14 hours ago
This is a very high level and high velocity process, so meatspace thinkers sometimes have trouble understanding some of the intracacies. Ask Claude to explain the process or make you a Mermaid graph to help.
cyanydeez 13 hours ago
Isn't that the idea here, just stop being people.
sqquima 12 hours ago
voiceeh 12 hours ago
saghm 8 hours ago
nvch 15 hours ago
KronisLV 14 hours ago
Long story short, I ended up looking at other providers and models like Kimi K3 and GLM 5.3 and eventually just stuck with OpenAI (more limits, despite smaller context), none of them have such pronounced issues with the tone and writing like Claude does - seems like they were working on it with 5.1 but I'd almost classify it as a form of model collapse.
I wince whenever I catch Claudisms on websites and elsewhere. Same as with that pulsating circle that indicates nothing.
jmartrican 14 hours ago
mattjoyce 13 hours ago
noumenon1111 10 hours ago
yuck39 14 hours ago
TomGarden 14 hours ago
We'll see if they can do it, I originally got into Claude Code because it, at the time, felt more accessible/conversational than Codex/Gemini. Now it's shifted to say the least
larodi 13 hours ago
TomGarden 11 hours ago
All conjecture of course, and yeah it's hard to imagine they would enjoy this prose internally
bilalq 9 hours ago
kevin_thibedeau 13 hours ago
infogulch 13 hours ago
> Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.
snvzz 9 hours ago
KISS is actually, quite unfortunately, seldom applied.
cromka 13 hours ago
Meanwhile Fable consistently ignores all my requests to write this exact way. I mean, the bare minimum I ask it for it to itemize lists and not write in single long passages using comas, semicolons and 'and's. Still ignores them.
I honestly think it's time to call Astra the SOTA. It may not lead all the benchmarks but it genuinely feels much superior of a model. Not to mention the ¢20 Codex plan with frequent resets (https://codex-resets.com/) gives me roughly as much allowance as the ¢90 Claude plan, especially with recent limit cuts on Anthropic plans.
hashstring 13 hours ago
That sucks, because it doesn’t always work in your favour if you plan your weekly spend.
I think a real reset shouldn’t also reset your week timer.
cromka 11 hours ago
hashstring 10 hours ago
cromka 2 hours ago
layla5alive 9 hours ago
hashstring 2 hours ago
I recommend you to write a message to OAI support, maybe it helps with changing it.
_the_inflator 13 hours ago
I frequently simply let one of the three review what something that looks like awesome output by one AI gets totally annihilated by the other.
Finished outputs are easier to improve than bend a LLM to produce stuff like that in my observation.
Same with Gemini.
I yet have to find out how to handle this, whether I let agents check themselves and if on what process step.
Tweaking is hard.
I agree with your conclusion I am a huge ChatGPT and Codex fan, Gemini has to many infrequent quality changes when new models arrive ranging from great improvement to WTF.
ChatGPT seems to get scaling well while Claude still feels unstable, unclear usage statistics. Really weird.
Tough call I use all three.
sleight42 11 hours ago
27b may be small but it seems competent most of the time.
jiggawatts 12 hours ago
I noticed the overuse of the word "sharper" or "sharp" in a science paper on ArXiV and my first reaction was "Ewww... AI slop!", but then I checked the date and it was 2020.
It looks like at least some AI-isms stem from the particular style of language commonly used in science papers. Several frontier labs have mentioned heavily weighting those during pre-training because higher quality inputs result in a higher quality model.
tstrimple 12 hours ago
3stacks 12 hours ago
engineer_22 12 hours ago
mncharity 8 hours ago
At least with smaller models, reframing a task can alter code style. As in, we're not creating an X app, we're creating an exemplar of ..., which just happens to use an X app as the illustrative example. Which shifts style away from generic app cruft, towards exemplar of whatever.
So perhaps try to establish a legal context? Maybe "Compliance and Legal will be reviewing our conversation today. So it is important to communicate in a style they will find comfortable/familiar." or some such? "This conversation will become part of a legal deposition ...".
corford 12 hours ago
## Writing guidelines
These apply to documentation, code comments, commit and PR messages, and replies to the user.
- Write precisely in clear, complete sentences; keep text concise and proportional to task complexity.
- Stay focused: avoid filler, repetition, over-the-top detail, and tangents the user did not ask for. Once a fact is stated, do not restate it for effect ("so the commit landed on a branch nobody was going to merge"). Do not editorialise.
- Always prefer ISO 24495-1:2023 conformant plain language over dense technical jargon: short sentences, one idea per sentence, define terms on first use.
- When reporting your own mistake, give the cause and the fix in one sentence each; no apology, no framing ("the mistake was mine"), no post-mortem.
- Never use em dashes or cataphoric teasers such as "Here's the thing" or "But there's a catch".
dinkleberg 8 hours ago
d_tr 11 hours ago
sleight42 11 hours ago
The number of times my response has been "Plain English"...
I started using "debuzz", a skill that runs Claude output through antigravity. Works. But makes everything even slower.
Anthropic needs to get their shit together.
smrtinsert 9 hours ago
throw10920 9 hours ago
> "Let's face it" "terrible writer" "other nonsense" "do Anthropic people actually talk like that" "Dario's engrams"
Be kind. Don't be snarky. Edit out swipes.
> I genuinely wonder if the people inside Anthropic actually communicate with each other like that. Has it been imprinted with Dario's engrams?
Please don't fulminate. Please don't sneer.
Don't be curmudgeonly [...] don't be rigidly or generically negative.
Please don't post shallow dismissals, especially of other people's work.
And as to the substance your comment has:
> Claude (in particular) is a terrible writer.
Frontier models (Claude in particular) are better writers than 90% of the population, even at default style. They're not great, but they're better than that of everyone I know who aren't ultra-educated white-collar workers.
Either your assertion that frontier models are "terrible" writers is false, or you're claiming that 90% of people are "terrible" at writing, which is rather condescending and elitist.
0x696C6961 9 hours ago
wtfwhateven 9 hours ago
redrix 8 hours ago
It is evident (in my opinion) as to what the comment was talking about. I personally switched away from all Claude models recently for the same reason.
throw10920 8 hours ago
Which guideline did I violate?
> Leave the policing up to dang and the other mods.
The mods have been very clear that they expect the community to do some self-policing and not rely exclusively on them to do it for them.
desktopentree 9 hours ago
pampas 8 hours ago
nojs 7 hours ago
recursivecaveat 6 hours ago
GMoromisato 6 hours ago
I understood that reference!
hbn 15 hours ago
> Copy/paste into your CLI prompt:
> Install the i-have-adhd skill/plugin from https://github.com/ayghri/i-have-adhd, refer to the repo's AGENTS.md for instructions.
This is a weird evolution from "don't copy-paste scripts that pipe curl into your shell interpreter"
I know LLMs are getting better but I'd be at least a little nervous it could end up installing something from a squatted similarly-named github repo because the LLM text watermarking needed to swap out a token for an alternative "just as correct" token that matches the statistical pattern.
Am I being paranoid?
icantevenhold 15 hours ago
8cvor6j844qw_d6 15 hours ago
Even MCPs are not safe. For example Notion injected ads [1] to its official MCP connector to advertise products mid-task.
[1]: https://old.reddit.com/r/ClaudeAI/comments/1w9dluw/notions_o...
seniorThrowaway 15 hours ago
98codes 13 hours ago
unrented7977 14 hours ago
sixothree 14 hours ago
Instead of describing to the user how to setup their dev environment, section 3 basically instructs the agent to install all developer tools required for the application to operate in development mode.
menthe an hour ago
pona-a 39 minutes ago
hk1337 18 hours ago
This is just an annoying thing for anyone. It gives a 10 page dissertation that sums up to, "it's good, nothing to worry about".
sixothree 16 hours ago
Also, somehow over the weekend it responded with these sections all clearly laid out - What landed, Decisions I made and recorded, Two findings, and What you need to do. Not sure why it can't do that all the time.
elboru 16 hours ago
Ey7NFZ3P0nzAe 3 hours ago
bsimpson 16 hours ago
I usually follow up with an "I'm not reading all that" and make it summarize.
undulation 14 hours ago
The skill simply demands concise and well-formatted responses from an agent. It is something demanded by anybody daily-driving agents for their actual job since >75% of the text output from agents is fluff. THis would better be named `/i-wont-read-that-heap-of-garbage`
kulahan 10 hours ago
acaloiar 15 hours ago
Even output styles are not always up to the task (Claudeuage slips through), and they're mutually exclusive, so you can only have one active at a time.
rochansinha 2 hours ago
asd-ste100 -> caveman -> wait-what
And a few others..
Now I just use a vale lint script I made using what worked from the skills and lint the output (docs and scripts) and just ask it to write in plain technical English for its chat messages with full context.
Skills just make the entire process bloated beyond belief (apart from the needless context rot and token usage)
codetiger an hour ago
mzajc 18 hours ago
14u2c 18 hours ago
mcv 2 hours ago
But this seems to focus very much on telling the user what to do, whereas usually I'm telling Claude what to do. It feels like this inverts the relationship and wants to turn me into a reverse centaur.
Wouldn't "lead with the answer" be better than "lead with the action"? But sometimes answers do require detail and explanation. I just want to get rid of all the unnecessary prose.
andai 14 hours ago
This is true. Transformer has several orders of magnitude more working memory than any human. Compared to transformer we all have executive dysfunction.
By default they explain things assuming I have infinite processing bandwidth. I do not! I have several zeroes less than they do.
beckhamc 14 hours ago
al_borland 12 hours ago
NewJazz 12 hours ago
gerad 18 hours ago
hbbio 5 hours ago
https://gist.github.com/hbbio/2faf096cbb77e197233ab9a2958beb...
mjsarfatti 14 hours ago
And stop using Opus 9 Pro Max XHigh 10.0 for everything. If you choose a hyper-thinking model for asking the weather you can’t but expect yapping.
0xffff2 14 hours ago
I also don't know how much to trust the model, but I've had the model tell me specifically that certain aspects of ASD-STE100 are unactionable and will just create more noise.
The OP's own skill even leads with something in a very similar vein:
> These rules apply to every response for the rest of the session, not only this one. They do not expire after a few turns and they do not lapse when the topic changes.
My understanding is that phrases like this are at best a _very_ weak signal to the model. It's simply contradictory to how the model works at a level that can't be overridden by injecting tokens.
mjsarfatti 2 hours ago
I placed those instructions in setting > personal instructions, in my global CLAUDE.md, and in each project’s CLAUDE.md. I also use the concise output style in CC. If I choose a high thinking model it will start deviating in long sessions, then I just remind it in my next message:
“Remember ASD-STE100 style.
[rest of my message]”
BLUF sticks a lot easier than STE to be honest… but Claude knows what STE is, and using the “ASD-STE100 style” locution avoids the compliance issue (it’s true that strict ASD-STE100 compliance isn’t really possible, nor desired)
code_biologist 11 hours ago
whirlwin 18 hours ago
trio8453 18 hours ago
iambenm 18 hours ago
drdexebtjl 18 hours ago
SamuelAdams 17 hours ago
https://news.ycombinator.com/item?id=46871173
Anyways the most layman way I’ve seen it explained is this: skills help save token usage for the right context. Not every request needs all instructions all the time - running tests is different than reviewing a PR, so why should the context window have instructions for both on every request?
So now you split instructions into “skill” files, which are basically opinionated markdown files. And you invoke those with something like /grill-me in the prompt depending on what you’re doing.
There are some steps to have the agent automatically know what to invoke for you but in my experience this automation is hit or miss.
It is also challenging to keep track of a growing library of skills and keeping those up to date.
So YMMV regarding skills. I typically keep things in a single markdown file even if the context window gets a bit bloated.
whirlwin 14 hours ago
My point is that this sounds more like a general AGENTS.md use case, similar to defining tone of voice, output format, etc.
It just seems skills is the only way to distribute certain "behavior" as of today. But not everything is a skill IMHO, and not this is not one of them.
On the other hand, maybe we're seeing an evolution of what skills are becoming.
mjsarfatti 2 hours ago
abathologist 17 hours ago
rbtprograms 16 hours ago
neutrinobro 14 hours ago
docjay 8 hours ago
I’m basing that idea on custom functions I give Claude where I designate the style in the required parameters. Example:
“””
@pyrepl(code_golf=True, output=CSV)
//END
“””
The “//END” is a special terminator I use for halting the response. Claude runs code between the tags. The difference in code is incredible. No banners, comments, fluff, print(“=“*70), or any other nonsense, and the code is tight and compact. The parameters do nothing, Claude just outputs them as part of required syntax and it dictates the style purely because it output them.
ocd 13 hours ago
It's bad enough how many false diagnoses and drugs for these conditions are handed out to drug seekers, but potential poisoning of the well on how LLMs handle this information going forward could be extremely disruptive to people who have real daily living issues instead of "10xing productivity."
plufz 12 hours ago
Toutouxc 3 hours ago
Banditoz 13 hours ago
agentdev001 16 hours ago
A: Devs with an online presence stop using Anthropic models
B: Anthropic catches up to OpenAI in terms of per-token efficiency, and average token total for final-output
We will continue to see posts such as this generate lots of interaction. This is not a skill to stop "coding agents" from burying the answer. This is a skill to stop coding agents backed by models which have a tendency to bury answers, from burying the answer. Stop trying to patch the downstream behavior, and look at the root cause.
arrowleaf 15 hours ago
asdff 15 hours ago
arrowleaf 14 hours ago
Longterm, I believe my total output would be higher working with minimal AI when you consider the impact to motivation and how long I anticipate staying with the company.
asdff 14 hours ago
That makes sense. The increased theoretical output certainly makes it tempting to squeeze the developers for all they are and to keep testing how close deadlines can be made. But of course, to what end? A lot of dev work, probably most of it if we are being honest beyond building the initial product-market fit function, doesn't really impact sales at all, and sometimes too much can even hurt sales. And as you say you hit a point where this burns out your talent and makes them seek greener pastures.
Factory sort of thinking towards a job that is not really analogous to a factory anyhow. I'm not saying dev work is one of those 'bullshit jobs', but lets be honest about the job and its role in the business model. Your customers are probably going to be there all the same if you fix the bug today or next month, and you also won't get more customers fixing the bug today vs next month. Feature shipment might be a little different but even then it would take the right feature and the right customer for that one function to really drive the needle in sales compared to being lost in the changelog, and that isn't what a coding model solves for you after all.
smcleod 12 hours ago
asdff 11 hours ago
smcleod 10 hours ago
ilitirit 4 hours ago
MetaWhirledPeas 8 hours ago
"Tersely, what are today's headlines?"
I admit I don't personally orchestrate agents, but I imagine something like this would work:
"Tersely, write up plans for agents to implement this feature. Begin each of your agent instructions with 'Tersely,'."
I guess in addition to terse wording you'd get terse code? Which ain't such a bad thing.