Show HN: Let your AI agents paint big arrows, boxes and text on your screen (github.com)
sicktriple 6 hours ago
sudo_cowsay 5 hours ago
Indeed, what a time to be able.
alanbernstein 4 hours ago
ASalazarMX 4 hours ago
fasterik 3 hours ago
arcanemachiner 2 hours ago
Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)
serf 4 hours ago
Anon1096 4 hours ago
godelski 3 hours ago
People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!
But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.
You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)
That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.
joquarky an hour ago
Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.
CamperBob2 2 hours ago
scotty79 32 minutes ago
When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.
Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."
Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.
1718627440 6 minutes ago
apropos ?
redanddead 12 minutes ago
/s
hn8726 10 hours ago
tkdb 9 hours ago
andai 8 hours ago
layer8 8 hours ago
trollbridge 8 hours ago
tkdb 7 hours ago
tottenhm 6 hours ago
causal 9 hours ago
cyanydeez 9 hours ago
jerf 8 hours ago
cyanydeez 8 hours ago
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
jack_pp 7 hours ago
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
mat_b 2 hours ago
Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.
nater5000 7 hours ago
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
Abimelex 7 hours ago
ccozan 6 hours ago
joquarky an hour ago
tempest_ 6 hours ago
The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.
ccozan 6 hours ago
zaik 5 hours ago
theropost 7 hours ago
fennecfoxy 7 hours ago
pydry 6 hours ago
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
glitchc 6 hours ago
m-s-y 6 hours ago
glitchc 6 hours ago
lstodd 5 hours ago
asdff 5 hours ago
jaggederest 5 hours ago
https://github.com/seb3773/ntfs-repair-rfc
https://github.com/Kotivskyi/screenshot-tool
I think it was more popular around the beginning of the year to mid-year
asdff 5 hours ago
rrr_oh_man 4 hours ago
cootsnuck 4 hours ago
Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.
asdff 2 hours ago
SwtCyber 8 hours ago
hannasanarion 7 hours ago
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
smugglerFlynn 7 hours ago
internet101010 an hour ago
I don't know why anyone would ever willingly want this.
tangotaylor 7 hours ago
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
socializer 7 hours ago
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
usrbinbash 9 hours ago
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
voidUpdate 9 hours ago
andyfilms1 9 hours ago
ghm2180 8 hours ago
mistersquid 7 hours ago
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
hannasanarion 7 hours ago
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
wartywhoa23 4 hours ago
cobbal 6 hours ago
arshxyz 10 hours ago
conception 10 hours ago
prmoustache 3 hours ago
She doesn't have to call you anymore.
kogus 4 hours ago
Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE
This kind of baked-in interactivity could really help in a training or disability context.
isoprophlex 10 hours ago
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
thih9 10 hours ago
ipsod 10 hours ago
isoprophlex 9 hours ago
franze 4 hours ago
ipsod 34 minutes ago
koalacola 10 hours ago
yen223 10 hours ago
isoprophlex 10 hours ago
ale42 9 hours ago
DonHopkins 9 hours ago
Maybe isoprophlex will add that to his PR!
voidUpdate 9 hours ago
It's already got the power to be obnoxious at you
jaapz 8 hours ago
htrp 10 hours ago
DonHopkins 9 hours ago
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
dihinbutt 7 hours ago
lbreakjai 10 hours ago
That would be a godsent for those of us with aging parents.
ghm2180 9 hours ago
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.
alansaber 6 hours ago
cyberjunkie 10 hours ago
priyashunt 6 hours ago
vessenes 11 hours ago
alansaber 10 hours ago
melvinroest 10 hours ago
For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.
We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?
FinnLobsien 10 hours ago
peaxkl 9 hours ago
We built something to help with that [1].
alexpotato 8 hours ago
Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:
"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."
swframe2 6 hours ago
satyanash 11 hours ago
What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?
If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.
inanutshellus 11 hours ago
"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".
Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?
ravila4 9 hours ago
dr_kiszonka 6 hours ago
(OP, nice project! Sorry for my rant.)
bel8 6 hours ago
But it's a subpar experience compared to just opening github on the browser.
ElijahLynn 2 hours ago
m-s-y 9 hours ago
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
ballofrubber1 9 hours ago
IanCal 9 hours ago
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
TekMol 11 hours ago
Does one need 4 programming languages to draw something on a mac?
sitzkrieg 11 hours ago
iandanforth 6 hours ago
ghm2180 9 hours ago
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.