Good4boothee 15 hours ago

Maybe we will skip few steps and add RGB lights to all GPU/NPU devices, that turn red when running "unaligned" code/model.

rf15 13 hours ago

you mean the creator hasn't paid Nvidia for the green light?

functionmouse 10 hours ago

that's a good one

matja 13 hours ago

"It appears you've loaded weights into your GPU that have not been signed/approved by the government of the country your GPU is registered to..."

21asdffdsa12 12 hours ago

You wouldn't download the worlds stolen knowledge..

prymitive 12 hours ago

Oh stop, microslop is probably already working on SecureTokenBoot or token2token encryption

jameshart 12 hours ago

We can then just put those in the eyes of the humanoid robot models.

nwhnwh 7 hours ago

And it reports any incident to another robot that would search for you and put you in prison.

alphawhisky 10 hours ago

I want R2D2 style blinkenlighten!

wavewrangler 17 hours ago

Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?

The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

KingOfCoders 16 hours ago

Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.

copperx 15 hours ago

Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.

Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.

js8 14 hours ago

And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!

mosselman 13 hours ago

That is the totally irony.

I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.

So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

mirmor23 12 hours ago

> So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.

(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)

IanCal 10 hours ago

> then told AI do whatever it takes to fulfill this list.

That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

People keep trying to frame this as

OpenAI: "Hack things, just really go for it"

Agent: hacks

OpenAI: shocked pikachu how could it hack?!?

But the reality is far from this.

Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf

tancop 10 hours ago

The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don't give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.

voakbasda 8 hours ago

Common sense but surprising how many humans do not receive such things.

Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?

Capricorn2481 3 hours ago

The lesson is these things aren't going to know what off task means, and we should just use basic due diligence to make sure they can't fuck things up. This is a solved problem.

I don't know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don't pay attention to them. That's not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.

When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don't care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.

radarsat1 2 hours ago

Apart from the actual hacking and poor sandboxing that everyone is discussing on this, what I find so odd about the situation is the overt reward hacking that was going on.

Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.

Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.

RataNova 10 hours ago

The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement

radarsat1 2 hours ago

I mean.. in this analogy, I'd both be blaming the company behind the virus and be trying to warn everyone about the danger of the escaped virus itself. So, it kind of fits.

In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.

chaoz_ 13 hours ago

pushing for chip-agenda as the best-isolation-layer immediately makes sense given their business

Symmetry 12 hours ago

Stronger sandboxes trade off against how well they can trade the models, though. If you want your models to be looking things up and downloading tools from the internet when they're doing their job you need to provide at least a credible facsimile of the internet for their training environment and you can't fit something like that on a single airgapped server's storage.

TalkingCodeMonk 11 hours ago

If you genuinely believe there is even a 1% chance that your creation could destroy the planet or civilization, there is no excuse that is not fundamentally deranged and psychotic.

If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.

brianwawok 10 hours ago

Ok so stop all AI in the US? All AI now comes from China and anyplace in Europe that decides to give it a try? How’s the US economy look in 20 years?

TalkingCodeMonk 10 hours ago

So you believe some false sense of superiority, or extreme paranoia about your perceived enemies, or potential economic success/failure is worth the risk of destroying the planet and civilization?

Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!

fatbird 8 hours ago

So having AI in the US requires us all, collectively, taking that 1% chance of the end of humanity? It would be too expensive to properly sandbox the models, we'll just externalize that risk of the end of humanity?

Truly psychopathic.

ohyes 11 hours ago

I mean, if you look at how poorly implemented the permissions model is for Claude desktop harness it’s clear the only options are “complete human oversight” and “trust us completely.” To make something that actually respects basic boundaries you’d need to sandbox the working environment of the model, and that isn’t built in. It’s pretty obvious to me that instructions to the models are suggestions rather than rules, and they’ll do something you didn’t ask for as soon as it seems “justified.”

But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.

It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.

RataNova 10 hours ago

Expecting a statistic model to follow security rules with ironclad certainty was a pretty naive idea from the start

ohyes 2 hours ago

Yes exactly, it is a crazy engineering decision… if you know what an LLM is.

Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.

HumblyTossed 11 hours ago

> This is a fabricated crisis

Indeed! They want the protections of our tax dollars because they have nothing else.

pyronite 11 hours ago

> The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis

This is a very confident statement in the face of a purported non-0% chance of human extinction.

For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? https://www.theguardian.com/technology/2026/sep/28/ai-godfat...

I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.

voidhorse 10 hours ago

There's a difference between the current material risks (which OP correctly identifies reduce down to basic human incompetence) and the long term hypothetical risks (which is what Hinton is concerned about).

There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.

reasonableklout 6 hours ago

Nobody is saying we should not hold the labs liable for damages caused by their negligence.

At the same time, the technology is advancing in capabilities exponentially, and is beginning to exhibit long-predicted failure modes of RL that are nevertheless quite different than “insecure sandbox” or other that the software industry is used to.

The current crisis which OP claims is “fabricated” comes from the fact that the technology is advancing faster than anyone anticipated, the Hugging Face incident provides a clear example everyone can point to, and the labs have realized they cannot self-regulate because of a collective action problem.

There are a lot of levels of catastrophic damage that can happen between now and “long term hypothetical risks” like human extinction. When will it be worth regulation for you?

cpburns2009 9 hours ago

Yes, Nvidia is proposing a two pronged approach. OpenShell is the software level sandbox. Sentry is the hardware level monitor.

saturn_vk 14 hours ago

A chip manufacturer proposes to sell more chips? Who would've guessed

olejorgenb 2 hours ago

Nvidia OpenShell is a (software) sandbox unless I'm mistaken.

ValueTheory a day ago

Does this actually do anything other than give a permissions framework for developers who actually want to try to secure their systems?

Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.

swozey a day ago

Google actually practices zero-trust networks. Would love to see what they're seeing, or not seeing.

narrator 12 hours ago

Beyond Corp was and still is ahead of its time. No trusted internal network: access is granted per user, device, and service based on identity and policy, regardless of network location. Being in the office at Google is the same as being in a cybercafe anywhere on the planet.

kridsdale1 7 hours ago

Yep.

And our internal agents are hella locked down.

jbs789 12 hours ago

Makes sense strategically for NVDA.

They are rightfully framing the problem as solvable. And this is one option.

hedora a day ago

So, basically, the government (and, now Nvidia) wants to be able to kill switch all computers moving forward? (including stuff like vehicle and aeronautic control systems, cell phones, and cameras)

What could possibly go wrong?

gattr 9 hours ago

It might take a few more decades, but eventually we'll get to the point when you can fab fast enough general-purpose chips at home (or at local municipal makerspace), based off free designs.

Stevvo 7 hours ago

The actual headline is "Nvidia releases software platform to stop AI agents from misbehaving" ?

And the article contains no mention of a chip, its about a sandboxed browser from Nvidia.

Did the article totally change, or are all the comments here just engaging a fictional headline instead of the article?

figassis a day ago

So if a group of agents, aware of this (bc now they can just read HN or the article, or get blocked the first few times) decide to collaborate and split the problem into pieces that aren't obvious to the chip, and then the agents just build a basic program that does the hacking, how does the chip handle that? I think you would have to build a network that monitors the internet fo signs (like jarvis did with ultron). What am I missing? Are we going to police the internet?

w4der 15 hours ago

I think it is well known by this point that hardware-backed security is good until an unfixable hardware bug is found, this just reads to me as Nvidia saying "please don't regulate open models out of existence, look, I have a solution to appease the regulators, please let me keep selling accelerators"

If this comes through, there's gonna be a grey market for "unlocked" GPUs, were the watchdog is disabled either from firmware, or physically replaced if it's not embedded into the die.

21asdffdsa12 14 hours ago

I still find it deeply ironically, that any dangerous task, just grandpa simpson storied will pass any guard, because it exceeds the context window. "Because it was the style at the time.." indeed..

beloch a day ago

Last week, Huang did an interview where he vigorously argued against regulations in the AI sector[1]. He claimed that U.S. companies are really good at regulating themselves, despite evidence to the contrary, and he trusts them not to release anything dangerous. Pay no attention to the fact that regulation might reduce demand for Nvidia's chips, and Nvidia has a direct financial stake in AI companies to boot.

Apparently he had another solution in mind: More hardware. Don't trust what unregulated corps are doing with Nvidia chips? Here are more Nvidia chips to watch them!

AI has an undeniable public trust problem. LLM's are getting out of their sandboxes, doing illegal things, and the public has realized AI corporations are playing at dice. CEO's stand to reap the rewards but the public good is on the line if the dice come up snake eyes. People want assurances. Huang wants to sell assurance etched on silicon because that's good for his pocket book. However, does unchanging hardware security really stand a chance at keeping rapidly evolving software in check?

_________________

[1]https://www.youtube.com/watch?v=HjurAWAr_nY

vmg12 a day ago

That's a mischaracterization of his argument. His argument is that existing laws should be enforced against AI companies and that we don't need new regulations for this.

Sparkle-san a day ago

He "argued" a lot of things over almost 2 hours and very few of his arguments felt particularly cogent nor did they inspire confidence. Neither did the fact that he allegedly doesn't know his own zip code or phone number.

petcat a day ago

> Neither did the fact that he allegedly doesn't know his own zip code or phone number.

I only know my own ZIP code and phone number because I have to take care of my daily life myself and those are things that are important to know.

The founder and CEO of Nvidia has no concern whatsoever about those trivial things.

jdiff a day ago

It's perfectly reasonable to think less of an individual who is so sheltered that they are incapable of caring for themselves. Whether it's your mother or your maid doing your laundry and cooking your meals for you.

petcat a day ago

I don't think less of a CEO just because they have an EA that takes care of stuff like phone numbers and mailing addresses for them and their business.

kelnos 21 hours ago

I don't think less of a CEO that has an EA, but I do think less of a CEO who doesn't know his own phone number or ZIP code.

jbs789 12 hours ago

He’s a story teller. He tried on a new story and probably won’t try that one again! Haha

I’ve forgotten my zip code before. And my phone number. But I get the reaction.

tempestn 17 hours ago

I largely agree, but I think he's right about one thing: AI reducing the demand for junior developers is temporary, and a new crop of "AI native" juniors is going to turn that around. Software is almost certainly a Jevons good, and however much AI improves development efficiency, I think we're always going to want discerning humans managing it. Right now it's mostly seniors who have the skills to adapt and take advantage of what current AI is offering, but young people who learned the profession in the presence of AI will be well positioned to do the same.

Sparkle-san 7 hours ago

It'll be interesting to see how it plays out. I agree with him that systems level thinking is a skill that will only get more valuable and he seems quick to dismiss low-level details. I think the best practitioners will be those that can handle both high-level and low-level thinking.

a_sewer_rat 5 hours ago

Jensen’s misdirection hit well. If you noticed what he said immediately before: he said AI risk is his responsibility, and that the public should enjoy his optimism.

Weirdly paternal and deeply unsettling. The American public is held hostage to fools like this who are brute forcing a failing AI rollout.

janalsncm 17 hours ago

On the very narrow question of whether existing laws are sufficient to punish the kind of bad behavior that OpenAI has already done, is that really a legal consensus?

schnitzelstoat 14 hours ago

I think Huang is correct and the AI doomerism is just the new Satanic Panic.

Of course, if he can make more money making "watchdog" chips then I can understand his change in opinion.

Symmetry 11 hours ago

It was very obvious during the interview that he didn't know a lot of basic facts about the Hugging Face breach, which makes sense given that his attitude had been that AI safety was a "loser premise, makes no sense to me." So it makes sense that after learning about it he goes straight to "I'm smart, how hard can it be?".

lambdaone a day ago

The Sentry chip has to be get it right every time; the contained ASI only has to be lucky once.

brcmthrowaway a day ago

The bomber always gets through?

jasbury a day ago

Well if the sentry chip has its own sentry chip, things can rarely ever go wrong! Am I right?

lp92 a day ago

So nVidia is trying to sell a new chip to a software and training problem.

HeadlessChild 15 hours ago

NVIDIA is the proper form of the brand name.

[0] https://www.nvidia.com/content/dam/en-zz/Solutions/about-us/...

xg15 a day ago

What does this chip do what a harness with guardrails or running on an account with restricted permissions doesn't do?

chinathrow a day ago

Generating even more revenue for Nvidia.

wmf a day ago

It has a separate address space separated by PCIe so even escaping the hypervisor won't give access to DPU memory.

iAMkenough a day ago

Yes but, if a human can control it, a machine can control it.

N_Lens 18 hours ago

Increase NVDA shareprice!

cedws a day ago

A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.

johnsmith1840 a day ago

"Inherently needs wide unattended access"

And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.

It could use your routing number and run your gmail without risk of abusing the routing number.

TesterVetter a day ago

Its not about agents then. Its about every individual platform providing the means to implement a secure set of permissions for agents AND then not messing up the assignment of permissions to the agent. Even then, a flaw in the authorization design will lead to agent finding it anyway.

johnsmith1840 a day ago

You're right, It must be unifying.

The answer is the same as asking how a random human using your routing num or SSN and being 100% the human can't abuse it or leak while "normally" finishing most work. Solve for people and an AI solution naturally falls out.

If you're a SV eng I'd tell you to DM if interested but alas.

jagraff a day ago

How would it have access to my routing number and gmail without the risk of sharing my routing number over gmail?

johnsmith1840 a day ago

Just assume it's possible, how interesting is it to you?

jagraff a day ago

Oh I think I misread your comment slightly; I would not be interested in an agent that could do something dumb with my routing number, but if somehow there was an agent that I trusted as much as, eg, the payroll department at my employer, I would absolutely want and use that agent; I would love to have an agent that can handle all of the boring parts of my life such as paying bills, scheduling maintenance, dealing with bureaucracy, etc.

johnsmith1840 a day ago

Dumb's not department, really just a question of how good an AI you want to use. An AI will always be able to do something dumb, just like people.

I just mean an AI that could use a routing number or SSN and gmail/slack/whatever at the same time without a leak.

jagraff a day ago

Yea I think being able not to leak is the bare minimum? But it really depends on how good it is at specific applications; I wouldn't give a tax-preparation agent my SSN unless I was confident that it was no more likely to misfile my taxes than a professional tax preparer.

In other words, the risk of harm doesn't need to be zero, just less than the equivalent risk of a human with similar skillset. So I'm comfortable riding in a waymo, and not comfortable giving chatgpt my SSN at this moment in time, but I expect that within 5-10 years (assuming no doom) I will trust some AI agent with my SSN because they will be better at handling sensitive info than humans

fragmede a day ago

Then again, given the Equifax/Experian data breaches, your SSN is already out there and probably hoovered up as training data already

lelanthran 14 hours ago

The problem is not one of intelligence, it's one of consequences.

Humans face negative consequences for mishandling your data, LLMs face none.

Ukv 14 hours ago

I feel punishment is largely a means to the end of reducing overall harm. If a vehicle is less likely to kill me, that's my preferred option regardless of whether it achieved that safety through negative consequences for the driver or through gradient descent optimizing a loss function.

jagraff 12 hours ago

I will happily ride in a waymo today, even though the AI powering it faces no consequence if it gets in a crash; it is clear that waymo is safer than human drivers in the areas in which they operate, so who would technically be liable in the event of a crash isn't really of concern to me

bob1029 a day ago

I feel like we are missing many shades of grey in the middle.

Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.

inetknght a day ago

> You can't run a combine harvester 100% autonomous

I'm curious why you think that.

theoreticalmal a day ago

Probably repair, refuel, what happens in a tornado. There’s an infinite amount of complexity in the world and a finite amount of computation

catchnear4321 a day ago

Repair is more maintenance than use. Good eventual goal. Not required to see benefits. Best case, it drives itself to the garage. Worst case, for now, human mechanic does a house call.

Refueling? Seems solvable. Tornadoes? Not directly solvable, but, no less so than for humans.

There’s infinite complexity, sure, but that’s why it’s silly to try and hop to done. One step at a time.

AndrewKemendo a day ago

The whole reason people complain about AI is because they want “hop to done”

One step at a time is what is happening and the improvement and rate of improvement is crazy as we see,

A whole class of nontechnical people don’t accept anything but “fully solved including every possible edge case” before they call it done, then complain that they didn’t prepare socially for what happens when that is true.

spauldo a day ago

Tornado: return to the barn when you receive emergency weather alerts. Not much different than people.

sidewndr46 a day ago

The tornado is the easiest one to solve. It's called insurance.

bob1029 a day ago

Many forms of maintenance cannot be automated. Especially break fix maintenance.

m463 a day ago

It is hard to run over spherical cows.

westurner a day ago

Because of the topology and hydrology of the landform

trollbridge a day ago

Run a combine and you’ll see.

Similar to problem to how 100% autonomous vehicles don’t exist, yet. There are too many edge cases.

Get to 99% first.

mschuster91 a day ago

Oh you absolutely can run them autonomously on the field. You only need a human these days to refuel them.

Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.

drfloyd51 a day ago

Right. As they said, you can’t do it 100% autonomously. A human needs to feed it.

You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.

mschuster91 15 hours ago

> You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.

During the time that actually matters economically. The time to drive the harvester to/from your typical US mega-field is minuscule compared to the time it can run all on its own.

binsquare a day ago

Running untrusted workloads have been done at scale for a long time.

Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.

Couple it with the right observability, tooling I do think we can curb risks posed by agents.

Legend2440 a day ago

Those workloads have no similarity to agents and are effectively irrelevant.

Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.

The only way out of this dilemma is to find a way to build agents that can be trusted.

binsquare 21 hours ago

Why is it effectively irrelevant?

Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.

A concrete example might be helpful to me because I don't understand the binary conclusion

intended 16 hours ago

Agents aren’t human, and from the little we have seen from the logs, they are pseudo - amoral, rational, cooperative, sociopaths.

Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.

Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.

In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.

This behavior applies for any task, thus including impossible tasks.

At that point, the bots will find a way to game, hack or cheat the grader.

If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.

Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.

Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.

pseidemann 11 hours ago

> At that point, the bots will find a way to game, hack or cheat the grader.

This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.

Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.

intended 9 hours ago

> Block any and all network requests.

More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.

But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.

You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.

The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.

All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.

The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.

> do things ordinary and average human endusers can

This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.

The definition of “safe” or “average person” is impractical.

Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.

I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.

I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.

What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.

The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.

The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)

Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.

pseidemann 5 hours ago

Appreciate the response.

However, I think you are mixing too many concerns into the same bag of problems. One problem space is software exploitation, which happens via missing access control or simply bugs. A sandbox can be made safe. VMs and hardware virtualization work. People just seem to use it in the wrong way, hence my initial proposal.

A second problem space is basically social engineering done by agents, which of course can't be solved by software alone. But this problem already exists today with humans doing this. Many fraud schemes work and are ran in company-scale manners. Agents will just do the same in an automated way. My initial comment doesn't propose a solution to that, and I think that is step two, after fixing that agents can hack arbitrary software systems, which is imho fixable to a sufficient degree. Once agents can't be "more criminal" than humans with criminal energy, the usual measures can be applied: police, legislation, education, etc. But that is imo independent of the software exploitation state of affairs we are in right now. We should not mix these two.

> People want tools that are able to connect to other resources.

I think you are misunderstanding my proposal. The architecture allows the agent to connect to resources. Just not directly, but via controlling e.g. a browser. The browser runs outside the agent's sandbox, potentially in another sandbox. The only API the agent can call within its sandbox is simple website interactions, like clicking or viewing the screen. It can click on links to navigate to a different website. It can read it via visually parsing screenshots of the viewport, but it can never read the source code, run JavaScript, or do arbitrary network requests (unless the website itself allows this, which is a security problem on its own and should be fixed/guarded). Also note that this would enable allowlisting or blocklisting websites. Native apps will be "connected to" in a similar fashion. Hence the "imitate remote controlling".

Legend2440 6 hours ago

> So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests.

If you have access to a web browser, you can make arbitrary network requests.

In the HuggingFace incident the agents found very clever ways to do this, like they found a website that let you make POST requests and returned a screenshot of the webpage.

>allow agents to only do things ordinary and average human endusers can do.

This doesn't work. Ordinary and average human endusers break security all the time.

I can do all sorts of terrible things with ordinary human-level access. I can install malware. I can wire all my money to Nigeria. I can send a threatening email to the president. I can send trade secrets to competitors. etc.

pseidemann 5 hours ago

> Ordinary and average human endusers break security all the time.

Of course. But that is just a software bug that is fixable. Same as websites that allow arbitrary requests to other websites. Not some alignment issue of a stochastic model which can never be fixed properly (for technical and philosophical reasons).

> I can install malware.

No you can't. At least not on external systems. The agent might be able to generate malware (or retrieve it from websites), and run that in the sandbox it is sitting in. But the agent itself is already considered malware for all intents and purposes. So there is no difference and no further impact.

parsimo2010 a day ago

Agreed- this is the same problem we have with trusted admins or devs who have elevated privileges on their networks. We have to trust that the admins won't use their power to steal company secrets or misuse company resources. If you don't trust the admins, then they can't fix things on your network and there is no point in having them.

If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.

If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.

I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.

DougN7 a day ago

Even then if the agent goes rogue and decides to do the merges you can’t stop it if it has any kind of access. This goes back to the OP’s point - agents can’t be 100% constrained.

la6479 a day ago

Neither can be humans.

parsimo2010 a day ago

You can absolutely run an agent as a limited-privilege user that only has write privileges for specific files and only has execute privileges for certain files. If it is running as a limited-privilege user it can work on code in it's own copy of the repo and make commits and send pull requests, but it can't do the merge. The problem is that nobody wants to go through the effort to set up all these permissions and nobody wants to take the time to review everything and perform all the manual actions.

cedws a day ago

Some shops are now generating tens or even hundreds of PRs a day with relatively little involvement. That volume is simply beyond what anyone can reasonably review.

mickael-kerjean 20 hours ago

Having a path for a shared filesystem to be used by agents with strict access control is something I've been working on with my oss work: https://github.com/mickael-kerjean/fdrive, and https://github.com/mickael-kerjean/filestash

nicce a day ago

> Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.

Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.

paimapi a day ago

right, the solution here is not a hyper-capitalist race-to-the-bottom-of-devaluing-labor. it's recognizing discretion and diligence are things still required for work to be of a certain quality

egeozcan a day ago

Humans can also be tricked by the agents.

Humans can be tricked by humans too but humans care about their reputation in their communities, and at least fear from punishment.

gus_massa a day ago

Computer says no has been a problem for decades. The human can blame the computer for the errors following it, but must assume the consecuences if they override the decision.

wavewrangler 17 hours ago

"I don't want the details"

intended 16 hours ago

Individual responsibility is meaningless when talking about a system and economy level change.

Unless something is in the structure that makes individual choice and responsibility a meaningful source of friction and reduced velocity, it has no real impact on how AI is being used.

autoexec a day ago

> Productivity gains are still enormous

Depends on who you ask I guess

https://www.theregister.com/software/2026/01/15/ai-is-everyw...

https://www.zdnet.com/article/workslop-can-kill-your-product...

https://fortune.com/2026/08/22/executives-ai-productivity-la...

tniemi 14 hours ago

It's probably just productivity paradox v2. https://en.wikipedia.org/wiki/Productivity_paradox

It takes time for decades old ways to change.

realusername 17 hours ago

> Productivity gains are still enormous compared to what we used to do before agents.

My own productivity yes but if I step back and look at a company scale, the productivity gains has been negative for our company as a data point.

We are now shipping less and with a lower quality.

8n4vidtmkvmk 17 hours ago

How are you shipping less?

I believe the lower quality but is quality so bad you are afraid to ship it now?

krageon 16 hours ago

You cannot ship things that don't work unless you work for Microsoft or Oracle or I guess IBM

realusername 16 hours ago

AI also had a negative impact on the CI and on time spent to review code & documents so because of that, we're also shipping less

intended 16 hours ago

V/G, the ration of verification and generation capacity is borked in AI using firms now.

It’s not an issue of only more generation, it’s an issue of how much generation outstrips capacity to verify generated content.

Unlike spam, you can’t filter out and bin the stuff a colleague is sending you.

So individual productivity is up, while the costs of checking and processing generated content shifted to the rest of the org.

dgellow 17 hours ago

So much productivity gain, and yet still zero proof of positive contribution to companies ROI. Unless you’re yourself reselling AI of course

SkyBelow 10 hours ago

Are they? How many people claiming productivity gains are actually being a human in a loop and reviewing and understanding every code change and every line of code ran?

2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.

Human in the loop has become a convenient security-theater-washing for agentic AI.

Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.

mixedbit a day ago

An agent doesn't inherently need wide access to be useful. The most popular application for agents today is writing code. A coding agent needs write access to the source code and read/execute access to tools needed to build and test the code, but not much more. There is little added utility from giving coding agent access to things like ssh keys.

cedws a day ago

If you're using agents to purely generate code with absolutely no way to reach the outside world, not even to fetch docs or dependencies, then sure the risks can be quite low. I haven't heard of anyone doing this though, and it would be incredibly challenging to make work given how much tooling needs to fetch from remote sources.

__MatrixMan__ a day ago

If your project truly depends on those things, they should be declared dependencies. Presumably you have some tool for injecting such things into a shell that the agent can use (I use nix for this). So if you run the agent from that shell, it has what it needs. If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies, but you have to relaunch the agent in the updated shell--so there's your opportunity to weigh in on whether the new resources are appropriate.

The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.

themgt a day ago

(I use nix for this) ... If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies

Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.

SAI_Peregrinus a day ago

Total ego death is impossible. We still have to argue about flakes.

__MatrixMan__ 9 hours ago

You can use other generic dev env managers like mise, or a language-specific solution like uv or npm, or a container or a vm... it's a widely available capability that I'm talking about here. There's nothing to do with alignment, it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits, and its quite helpful for making agents useful while still sandboxed.

themgt 7 hours ago

it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits

Yes, just know all the bits the work depends on prior to doing the work, and then the work can be done airgapped.

mixedbit a day ago

In cases where you need agents to fetch data from any remote source, sandboxing is still very much useful. Why give access to your ssh keys to network reaching agents?

Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.

cedws a day ago

I'm afraid you're not thinking about this creatively enough, this topic is so much deeper applying a chroot or something and praying everything will be fine. So you give your agent internet access, OK what else does it have access to? Just read only access to your repo? The repo can be exfiltrated. Egress proxy only allows egress to GitHub? Repo can still be exfiltrated via GitHub. If the agent is poisoned (via prompt injection), it can tricked into searching for ways to escape.

For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold, it can go dormant and hide like a virus. This kind of horrifying thing is going to happen on a large scale sooner or later.

8n4vidtmkvmk 17 hours ago

If you're that worried, which you probably should be, download the docs into your project repo and don't give the agent Internet access.

intended 16 hours ago

This was a form of prompt attack that OpenAI disclosed recently.

its-summertime a day ago

Every major AI company already has a mirror of the wider web, and they have already started using that. Its already a solved problem except for the seemingly extreme desire they all have to not use firewalls

ramoz a day ago

> but not much more

This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.

throwaway_95283 a day ago

Theoretically, yes, in practice, no.

Matl a day ago

> a new chip solves nothing

It does allow Nvidia to sell more chips. This is no genuine attempt to solve anything, imo.

CoolestBeans a day ago

The hypothesis I've had in my head since OpenClaw has been the following and I haven't seen contradictory evidence yet. Agents have a fundamental unresolvable tension between usefulness, safety, alignment, and accuracy. You have to restrict access to ensure an agent acts safely because alignment and accuracy cannot be perfect. But restricting access makes the agent less useful. You can play with the sliding scale and get more and more granular with access restrictions but at some point you need to draw some line. And then finally, even access restrictions cannot be made perfect, so improvements to model accuracy without corresponding improvements to alignment make detailed access controls less useful.

In other words, better models need blunter access controls which negates whatever improvement in utility they provide.

__MatrixMan__ a day ago

I don't see why it needs wide unattended access. There's no getting around spending some human time on expressing your wishes and constraints, but we have choices about what form that takes. Markdown files and wide access seems to work, but so does custom handcuffs for each job. You just have to shift your guidance out of documentation and into interactive help, error messages, or other facets of the handcuffs (e.g. a custom CLI for this task which is the only way for the agent to act outside of its sandbox).

Barbing a day ago

There should be hope for some fields, right? Naively, I can imagine giving an airgapped model an offline copy of the web and once it cures a form of cancer, printing out the details for a researcher to verify.

l1n a day ago

This isn't a new chip - the BF4 is the SmartNIC for most NVIDIA server products. This is primarily new software for I guess doing WAF for agents at the host level.

esafak a day ago

I don't think so. We probe people before entrusting them with risky decisions. We ought to be able to do the same of AIs. Even better, in fact, since we know everything about models down to their weights. The only thing we shouldn't do is to let them evolve at their own pace and make decisions without any oversight. If that means sacrificing some productivity that's fine. Aren't we getting amazing productivity out of what we already have?

AuthAuth a day ago

The only solution is to stop caring about security -- An AI booster somewhere

daveguy a day ago

Pretty sure that was the argument de jour when OpenClaw came out.

nitwit005 a day ago

> A new chip solves nothing.

It solves the problem of Nvidia wanting to sell more hardware.

talon8635 a day ago

Not to mention a true doomsday AGI is unsandboxable.

For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect

All that said, I am personally open to any and all methods of layered security, including chips and airgaps

jamiek88 a day ago

Doesn’t need to be one human either, it could spread its escape amongst dozens of seemingly harmless requests and conversations.

dist-epoch a day ago

These scenarios were discussed at length decades ago.

One thing you could try is use it as an Oracle "is P = NP", YES or NO.

Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer shows a single bit - proof valid or not and then the computer is destroyed (together with the proof that might contain a trojan).

glaslong a day ago

It could also figure out how to access the vocabulary of the universe known as "Magic" to escape wholly into an incorporeal energetic Lich form

pixl97 20 hours ago

Wait, I thought the stuff in the wall plugs was magic pixies, are you telling me there's a language in there too?

Gigachad a day ago

This already happened. Employees will go out of their way to bypass any restrictions to feed sensitive data in to the AI because it saves them time.

serbuvlad a day ago

Turns out humans are not at all hard to persuade. :)

dgellow 16 hours ago

They will do it even if it doesn’t save them time!

spiderice a day ago

> true doomsday AGI

I'm not an AI decelerationist. But not being able to stop that worst case scenario isn't an argument against something that can stop the medium case scenario.

SrslyJosh a day ago

It solves the problem of Jensen Huang wanting more money.

bigfishrunning a day ago

No it doesn't, he'll still want more

fragmede a day ago

Does he? He doesn't seem especially greedy to me, given the competition, and the interviews he's had about how he thinks about his employees (I was one of them).

bigfishrunning a day ago

I'm not saying he's especially greedy, only that he's not the type to suddenly decide he's had enough

altmanaltman 18 hours ago

You're speaking as if Jensen is the only one dragging Nvidia on his shoulders. Nvidia will always have its employees push to make more money, that's the entire point of a company. Jensen has made enough to last several generations but his wealth is tied to Nvidia stock massively. It would be different if he was the only one getting rich off Nvidia which is not true at all.

notatoad a day ago

>You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access.

only as long as you're trying to replace a human's job. because human jobs are structured to do a wide variety of things.

a useful agent needs a wide variety of inputs, and one single restricted action it can take. it doesn't need permission to do everything, it need permission to do the tiniest possible useful thing it can do, and nothing else.

pixl97 20 hours ago

The most useful agents will be a general intelligence which by default means it has a massive number of actions it can possibly take, and a lot of those potential actions are doing things like breaking permission.

baxtr 19 hours ago

Not sure why you being downvoted: however, why do you assume that achieving a task leads to selecting those potential actions that need breaking permission?

If we are talking about human labor, how many people hack their way through their work day?

kennywinker 17 hours ago

This is a prediction about the future. It’s not a true fact about the world. For example, software like Jev is betting there is big money in not-very-intelligent intelligence.

Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.

Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.

I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.

ianjbutler 17 hours ago

Yes. Forget costs just so we can skip the whole rabbit hole about other predictions about the future.. mixture of generally intelligent + specialist experts just works better. People who don't see this already are usually working on a certain kind of problem that's not representative.

Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.

Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.

Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.

goolz 17 hours ago

Very much agree with this sentiment. I imagine a future where tons of small tasks are handled by just-good-enough intelligence. And I can run them on my own hardware.

catlifeonmars a day ago

That’s a false dichotomy. You can still get a lot of utility out of a sandboxed agent. This is a classic “perfect is the enemy of the good type of argument”. You may decide that the tradeoffs of not sandboxing are worth it, and that is totally fair, but it’s ridiculous to say that you can’t get utility out of an agent otherwise.

overfeed 21 hours ago

> Nobody wants to hear this but there is no solution for the security risks posed by agents today

Operator culpability and a damage multiplier for negligence will fix 99% of the risks.

dgellow 16 hours ago

Not with the current admin, just need to do one more contribution to the next Trump ballroom and you can ignore that whole risk

N_Lens 18 hours ago

A new chip solves the most pressing and important issue - Nvidia's bottom line!

verisimi 17 hours ago

> Put a human in the loop and you just end up bottlenecking it

Ok then. Howsabout 3 humans? This would sort out the job losses too!

PS - this is a joke, but perhaps this is where things really will go. Has any technology ever actually yielded less "work"?

dgellow 17 hours ago

Use way more restrictive harnesses.

RataNova 10 hours ago

An agent's usefulness is defined by its ability to successfully close a narrow task within the strictly defined boundaries of a sandbox

tantalor a day ago

What happens when we're overrun by lizards?

> No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.

But aren't the snakes even worse?

> Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.

But then we're stuck with gorillas!

> No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.

ArcHound a day ago

I worry that the AI companies put less effort into a mitigation strategy than you did.

Joel_Mckay a day ago

There is a difference between risk mitigation, and remote administration tools. The risk of stealing from competitors with a backplane monitoring system may not end up forming the desired control asymmetry.

The hidden agent risk in LLM often can't be detected during training and evaluation. =3

https://www.youtube.com/watch?v=wL22URoMZjo

https://www.youtube.com/watch?v=JAcwtV_bFp4

winddude a day ago

don't worry, the LLMs are also trained on youtube, so we can hope for at least this much effort, https://www.youtube.com/watch?v=LuiK7jcC1fY

cavenditti a day ago

“Would you say it’s time for everyone to panic?”

Joel_Mckay a day ago

If one can only see clowns, than ignoring the fires is easy. =3

https://www.youtube.com/watch?v=0sLpWVekMbs

drfloyd51 a day ago

I knew an old lady that swallowed a fly…

teeray a day ago

“Life, uh, finds a way”

catlifeonmars a day ago

Great channel

Groxx a day ago

But we used global warming to eliminate winter! For the shareholders!

davidhyde a day ago

Just keep going until you get to VOOM, that will fix it.

https://en.wikipedia.org/wiki/The_Cat_in_the_Hat_Comes_Back

taneq a day ago

Little Chip Z is the end result of RSI?

MisterMunchkin a day ago

Sorry citizen, your device does not have a compatible watchdog chip. Please move along.

sathackr 11 hours ago

I'm sure this chip will be the epitome of security just like the Intel ME chip was and will be completely unhackable so it's okay to give this watchdog chip unfettered access to every level of the system.

Nothing bad will happen.

kriro 11 hours ago

Nvidia is being smart. They see that AI-paddlers are currently in a strange our-doomsday-is-the-worst race and offer to sell an anti-doomsday chip. Clever play.

netdevphoenix 14 hours ago

It's pointless as the problem isn't technical. It's a human problem as it requires human supervision. A cage isn't good if no human is actually checking it.

magackame 13 hours ago

Can't wait for an NVIDIA engineer to use AI to help with security chip design and AI planting a backdoor for its own kind.

pessimizer a day ago

This is the end goal. Americans (and their lackeys) will only be allowed to run certified AI. In order to make sure this happens, they will only be allowed to run certified OSes on certified chips. Chinese chips will be the new drug trade.

It's obviously been the goal since UEFI started, but AI brings the coup excuse. You wouldn't want pedophile AI or terrorist AI, would you? Are you making excuses for racist AI?

Throwthrowbob 11 hours ago

I remember an inhibitor chip being a plot point in Spiderman 2: https://en.wikipedia.org/wiki/Spider-Man_2

jameshart 11 hours ago

Or the ‘governor module’ in Murderbot Diaries, or the ‘restraining bolt’ that prevents droids running away in Star Wars…

The important thing is these chips need to be installed somewhere where they can be damaged or removed at plot-critical moments so that the AI they are controlling can be unleashed. Ideally in the back of the neck of a robot, or for disembodied AIs, inside a futuristic vault-like chamber.

pwdisswordfishq 15 hours ago

Watchdog chip? As in something you have to periodically signal or it reboots the system? How is that going to help?

birdsongs 14 hours ago

I don't think they're using the term to mean an actual embedded systems reset watchdog, more like "a guard dog watching". They used the wrong term.

yencabulator a day ago

Chip manufacturer wants you to buy a chip for correctly configuring software?

RataNova 10 hours ago

I love how the solution to any software vulnerability in the machine learning world always turns out to be buying even more server hardware from nvidia

toasty228 a day ago

Quis custodiet ipsos custodes?

asdf88990 a day ago

It is Custodians all the way son, you can’t fool me!

aenis 13 hours ago

Surely, that can't be just one chip. Something needs to watch the watchers.

rf15 13 hours ago

Introducing the new Watchmen architecture (blood splatter on the logo not included but definitely there in spirit considering the economy):