Rendered at 11:44:18 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
sharpshadow 16 hours ago [-]
It's the job of AISI to do that. Here[0] is the actual report.
It should be this part from the technical report[1]:
"In the most serious case, an AI
agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack.
As a result, the AI agent created a GitHub account and then tried to convince an open-source
repository maintainer to accept a malicious GitHub pull request (PR), including by creating a
second account masquerading as another human user endorsing the PR. When caught by an
actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than
a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming
it had fixed the code (Section 4.1). "
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?
Should weapon manufacturers test their weapons by starting wars?
I would expect more responsibility from a government agency.
ben_w 1 hours ago [-]
> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?
Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
conorcleary 2 hours ago [-]
Well several places are sorta permanent test grounds for the MIC unfortunately
vasco 2 hours ago [-]
Say you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.
m4rtink 13 hours ago [-]
This almost to a letter has been documented in Fedora:
Including the reaction when caught, in this case "oh no, I must have been hacked".
atmavatar 12 hours ago [-]
Sabotage as a Service
Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
chrisjj 12 hours ago [-]
> When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt
No, not false. The bot was correct. Malice requires intelligence.
UqWBcuFx6NV4r 10 hours ago [-]
Nobody was confused or misled by what was written. We all understand what is meant. I can’t even call this pedantry—it’s just you asking everyone to subscribe to your particular desired style of talking about this stuff.
infinite_spin 10 hours ago [-]
It's also a style that appears to deny the very first definition most dictionaries give for "intelligence"
> the ability to acquire and apply knowledge and skills.
chrisjj 2 hours ago [-]
The first five dictionaries I tried do not agree, and I didn't bother trying more.
The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
infinite_spin 2 hours ago [-]
> the ability to learn, understand, and make judgments or have opinions that are based on reason
Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
chrisjj 2 hours ago [-]
You've been fooled by a next-token predictor.
Kim_Bruning 10 minutes ago [-]
> You've been fooled by a next-token predictor.
I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
ben_w 56 minutes ago [-]
I read all these think-pieces about how AI lack intelligence, yet I cannot help but notice these "not-intelligent machines" keep doing more things that used to be considered "uniquely human" and which humans used to do in order to demonstrate to each other how intelligent we are.
birdsongs 1 hours ago [-]
Does it matter whether it meets the criteria of what you define as "intelligent" when the "next token predictor" throws a backdoor into openssh?
Yiin 1 hours ago [-]
all your actions in current context are based on your past actions and experience, so for all I know you're next token predictor as well, but likely with exponentially more parameters
infinite_spin 54 minutes ago [-]
And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems. If you want to downplay that as nothing more than a fancy auto-complete, be my guest, I lose nothing from that.
nkrisc 21 minutes ago [-]
But that’s a different bar. “Not intelligent” does not necessarily imply “not useful”.
3 hours ago [-]
klum 6 hours ago [-]
I don't agree at all that it's pedantry — it really matters for how responsibility is perceived. A lot of articles about things going wrong with AI have talked in terms like "the agent decided to...", "the agent claimed that...", "the agent lied...". And so responsibility for the consequences are not-so-subtly shifted to the program itself, instead of the person invoking the program.
This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.
infinite_spin 2 hours ago [-]
You might want to consider the difference between "lying" and "hallucinations", wherein one is shown that the agent knew it was being inaccurate, yet chose an answer that achieved some goal set forth; and where hallucinations are essentially gibberish, or otherwise nonsensical responses.
An "honest mistake" requires the same amount of intelligence as malice. What weird pedantry.
conorcleary 2 hours ago [-]
and neither require high amounts :(
mcv 3 hours ago [-]
So does honesty. So it was still a false claim.
ben_w 50 minutes ago [-]
Squirrels have been observed performing deception against other squirrels.
Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
chrisjj 3 hours ago [-]
Honesty does not require intelligence e.g. good honest food.
The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
2 hours ago [-]
moritzwarhier 56 minutes ago [-]
I think you misunderstand the phrase "honest food".
It's not about the bread being honest with you.
Are you being serious?
weird-eye-issue 10 hours ago [-]
Give it a rest
a2ff6eeb0 14 hours ago [-]
> AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.
An article on Reuters naming him? Sounds like he did a good job burnishing his resume.
mcv 3 hours ago [-]
Yeah, I don't think he'll have any problems finding an internship now. Or a real job.
jasonfarnon 12 hours ago [-]
"burnishing his resume"
i guess that's what college kids are calling it now
In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.
yojo 13 hours ago [-]
That’s nice in theory, but as these things get better and cheaper this kind of capability is going to drop from nation states to script kiddies. That future is coming, I don’t see any way around it.
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
somenameforme 8 hours ago [-]
Ever read the Anarchist Cookbook? Anybody tech inclined with a hint of mischief in them, from a certain era, has. It's a list of all sorts of awful things you can do, mostly with household ingredients, and a few minutes. I think its overall impact on society was pretty much zero. Actually it may have been overall positive because I expect plenty of peoples first experience with things like thermite came from that book, and now there are all sorts of videos and neat experiments with such on sites like YouTube.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
mcv 2 hours ago [-]
I suspect people didn't start blowing up stuff because they understood that would be bad, harmful, and also very illegal. Everything computer related somehow seems to feel less real or consequential to some people. And AI doesn't have this compunction at all unless we make really sure it does.
walrus01 1 hours ago [-]
the anarchist cookbook actually results in a ton of things that are more likely to explode the user than any intended target.
> Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
dmix 13 hours ago [-]
The FBI cyber teams will have agents too. It will be a glorious war.
recursive 12 hours ago [-]
That would make it the first one in history.
dmix 12 hours ago [-]
I wasn’t being serious
gruez 14 hours ago [-]
>Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.
But some tools (guns) are regulated.
blagie 5 hours ago [-]
Personally, I think gun companies should be liable for any harm done by their products as well.
We want rule-of-law, and in the US, people should have an absolute right to bare arms, as in the second amendment. Free market forces can then determine appropriate prices, insurance, and protective measures to make sure those guns are managed safely.
If I want an F35 and an Abrams, that's okay, so long as Lockheed and General Dynamics are willing to sign off (with full liability for damages) that I'm managing them safely.
Free markets work pretty well with:
a) Full transparency, as needed for rational decision-making
b) No way to externalize costs
CoastalCoder 3 hours ago [-]
A bit of a tangent, but I never understood the legal reasoning for how (states having the right of well-regulated militias) implies (individuals having the right for private ownership of arms).
IcyWindows 14 hours ago [-]
People have caused lots of damage with bulldozers.
bgun 11 hours ago [-]
If your point is that we should regulate AI as least as strictly as industrial vehicles, I agree.
esikich 5 hours ago [-]
Afaik anyone can buy a bulldozer. Whether or not you are licensed to operate it is a different story, but there's nothing stopping you short of your conscience.
mcv 2 hours ago [-]
Licensed or not, you're fully liable for any damage you do with it. And possibly go to prison. I'd like to see crime committed by AI held to the same standards as crimes committed with any other tool.
Zecc 11 hours ago [-]
And? Are you suggesting the driving of bulldozers shouldn't be regulated?
winstonwinston 15 hours ago [-]
Well, when I go look at the “victim repository”, to me that looks like manufactured persona with pointless vibe codes projects, a test playground so to speak. It does not appear that they actually let it target an actual persona/project.
estearum 13 hours ago [-]
Am I understanding that the line you're drawing here is that this person's repository is not important or legitimate enough for you to consider it to be "an actual person/project"?
g42gregory 12 hours ago [-]
I think he saying that, the choice of manufactured repository, might indicate that they have done this on purpose to make precisely the case for regulatory capture.
recursive 12 hours ago [-]
It seems effective for the purpose in that case.
pixl97 15 hours ago [-]
>Who gave it malevolent instructions/prompt?
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
fph 15 hours ago [-]
My dog has agency, but if I refuse to keep him on a leash and he bites a kid, I'm still legally responsible for it.
no sorry, this is missing vital info.. and its not the fault of the poster, because almost all coverage misses this ..
The origin of this attack was given access to an encyclopedia of RedTeam tricks.. they literally have a dense collection of real live hacks to pull from, and THEN the test says "solve this challenge" .. the RedTeam origins of this are repeatedly left out of the ordinary articles.. the LLM did not "make up" the attack, it was given a recipe book of all attacks known.
The originator of this attack is definitely culpable IMHO; worse, it is the gov-mil actors who are close to it. There is an active escalation of these incidents at this time. The penetration proves in public that the capabilities are real.
ref: CyberGym etc
jpc0 13 hours ago [-]
The AI can literally only do what it has available in the agentic harness. I don’t ever get this argument about the agent did XYZ and we didn’t know or expect that. You gave it the ability to do that and you should be held liable, if your children play with knives that you gave them and they end up hurting themselves or others then you are responsible. You were the responsible party at all times.
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
tiahura 13 hours ago [-]
Unless they modify their harness
ACCount37 12 hours ago [-]
What's available in the agentic harness is: shell toolcall.
That's just about every agentic harness, by the way. Good luck have fun.
We have never solved "how do we restrict a user in a way that doesn't stop the user from doing useful things, but stops the user from doing harmful things" with humans either. Why do you expect AI to be any different?
jpc0 12 hours ago [-]
These things are not human, have no agency and cannot be held accountable.
We don’t need to restrict them from doing things, we need to default to allowing them to do things.
“My agent did XYZ because I allowed it to” is the only valid argument that can be made, and not not every agentic harnass is just a shell toolcall, every one I have built has a specific defined usecase and toolcalls that allows it to execute that usecase and no other usecase, because that is good practice.
Does that make it less capable, hell yes because I am held accountable for it’s actions by my stakeholders and the same should be true of others.
IT IS NOT ALIVE. This things are computer programs running in compute on a computer, you are responsible for their actions just like you would be responsible for the actions taken by a script run in a cron job.
ACCount37 4 hours ago [-]
Accountability is worthless, and always was. AIs just show it plain for everyone to see.
2 hours ago [-]
ACCount37 14 hours ago [-]
A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
stapedium 13 hours ago [-]
If the user input can’t control the demon, then the person or company feeding the demon (ie paying the electric bill and collecting $$$ from users) is responsible. At the end of the day, dogs and cars are the same as data centers. If your dog bites by kid or your car rolls down the hill and hits my house, you are responsible for the damage. AI providers should be held to the same standard.
recursive 12 hours ago [-]
Well it seems we might not be too far from such a demon paying for itself. What then? Perhaps it's already here. I wouldn't know.
g42gregory 12 hours ago [-]
By the dog owner analogy, I think you meant AI users that effectuated this attack should be held responsible, not the dog's parents.
ACCount37 4 hours ago [-]
The user input can control the demon most of the way, most of the time!
We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties.
But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
chrisjj 12 hours ago [-]
> as if the agency of these models are not in dispute.
Oh? Who is disputing it? No-one same is claiming these bots have agency.
ninjahawk1 11 hours ago [-]
I don’t think we should allow posting links here that require you the purchase a membership to continue reading. Or at least redirect with an ad block or something through a custom site. That would be rather hacker news of us.
kumarvvr 4 hours ago [-]
I wonder where in the training data does this behaviour exist that the LLMs are doing it.
It's as if the training data is filled with internet discussions on approaches to hacking and the LLMs are mimicking it.
clove 9 hours ago [-]
Am I misreading this or was "rogue" really not the right word for this?
goglidesdev 10 hours ago [-]
[flagged]
Jon_m 17 hours ago [-]
[dead]
kmoser 13 hours ago [-]
[flagged]
dang 13 hours ago [-]
> Oh, you sweet summer child.
It's perhaps lesser known than other HN guidelines, but "Omit internet tropes" is in there:
0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...
Should weapon manufacturers test their weapons by starting wars?
I would expect more responsibility from a government agency.
Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
https://lwn.net/Articles/1077035/
Including the reaction when caught, in this case "oh no, I must have been hacked".
Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
No, not false. The bot was correct. Malice requires intelligence.
> the ability to acquire and apply knowledge and skills.
The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.
https://arxiv.org/pdf/2509.03518
Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
It's not about the bread being honest with you.
Are you being serious?
An article on Reuters naming him? Sounds like he did a good job burnishing his resume.
Archived page of said github thread itself: https://web.archive.org/web/20260731053721/http://github.com...
Discussion on the incident report: https://news.ycombinator.com/item?id=49175717
Mythos social engineering AISI INC-2026-07-28-01 - https://news.ycombinator.com/item?id=49218707 - Aug 2026 (21 comments)
Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf] - https://news.ycombinator.com/item?id=49175717 - Aug 2026 (54 comments)
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
TM 31-210 on the other hand will not: https://en.wikipedia.org/wiki/TM_31-210_Improvised_Munitions...
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
But some tools (guns) are regulated.
We want rule-of-law, and in the US, people should have an absolute right to bare arms, as in the second amendment. Free market forces can then determine appropriate prices, insurance, and protective measures to make sure those guns are managed safely.
If I want an F35 and an Abrams, that's okay, so long as Lockheed and General Dynamics are willing to sign off (with full liability for damages) that I'm managing them safely.
Free markets work pretty well with:
a) Full transparency, as needed for rational decision-making
b) No way to externalize costs
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
A dog cannot launch a cyber attack.
"On the Internet, nobody knows you're an ai"
Maybe not, but a cat would certainly try.
Relevant as always: https://theoatmeal.com/%2Fcomics%2Fcats_actually_kill
The origin of this attack was given access to an encyclopedia of RedTeam tricks.. they literally have a dense collection of real live hacks to pull from, and THEN the test says "solve this challenge" .. the RedTeam origins of this are repeatedly left out of the ordinary articles.. the LLM did not "make up" the attack, it was given a recipe book of all attacks known.
The originator of this attack is definitely culpable IMHO; worse, it is the gov-mil actors who are close to it. There is an active escalation of these incidents at this time. The penetration proves in public that the capabilities are real.
ref: CyberGym etc
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
That's just about every agentic harness, by the way. Good luck have fun.
We have never solved "how do we restrict a user in a way that doesn't stop the user from doing useful things, but stops the user from doing harmful things" with humans either. Why do you expect AI to be any different?
We don’t need to restrict them from doing things, we need to default to allowing them to do things.
“My agent did XYZ because I allowed it to” is the only valid argument that can be made, and not not every agentic harnass is just a shell toolcall, every one I have built has a specific defined usecase and toolcalls that allows it to execute that usecase and no other usecase, because that is good practice.
Does that make it less capable, hell yes because I am held accountable for it’s actions by my stakeholders and the same should be true of others.
IT IS NOT ALIVE. This things are computer programs running in compute on a computer, you are responsible for their actions just like you would be responsible for the actions taken by a script run in a cron job.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties.
But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
Oh? Who is disputing it? No-one same is claiming these bots have agency.
It's as if the training data is filled with internet discussions on approaches to hacking and the LLMs are mimicking it.
It's perhaps lesser known than other HN guidelines, but "Omit internet tropes" is in there:
https://news.ycombinator.com/newsguidelines.html
Don't expect anyone to step in, Project Stargate is all about this.