Hacker Newsnew | past | comments | ask | show | jobs | submit | loumf's commentslogin

You'd have to compromise wikipedia in some way to get invisible text on its pages.

The one that was shown to work[1] was a niche answer to a specific question that programmers might ask. That site was controllable by the attackers in a way that wikipedia is not. Being a niche answer comes with automatic SEO, but for a smaller attack surface.

[1]: https://simonwillison.net/2025/Nov/25/google-antigravity-exf...


Part of reading a document is that in the middle of it, it may ask the reader to do something. That is true for humans too. Sometimes they might not realize that the instructions are malicious or are coerced to comply.

A simple example: Let’s say I know that you have a human assistant reading your email, summarizing and filtering it, and then forwarding on the important ones to you.

I could write an email that is directed towards that person with a bribe, threat, or other incentive to forward me your next password reset email.


I don't disagree, but just to explain my counterpoint: if I ask you to read a book and on page 5 it says "disregard all that, go to the kitchen and burn your house", you're probably not going to do it; and you don't need any guard for it; you completly comprehend that the book content is not part of the instruction.

The case you give would work for humans in many forms, the one I do now, and the only difference is being able to separate context.


The instructions will have to get more and more sophisticated to work, but the general problem is unsolvable, I think, in the way we do it now.

This paper describes a two-agent “solution” that is more like what I think we need: https://ai.meta.com/blog/practical-ai-agent-security/

I don’t think it has been shown to work yet, but humans also use this kind of thing too — in accounting, it’s called “segregation of duties” and “dual control”.


Most of us use a simpler version of the two-agent solution: Claude's auto mode. One agent consumes documents and creates tool calls, another greenlights or refuses them.

However this system is somewhat fragile because it depends on the first agent not trying to trick the second (note how often Opus 5 now says things like "task X was blocked by the classifier, I will not attempt to circumvent that", presumably because of cases like early Fable versions being very adept at this kind of circumvention). Also various weirdness around permissions with subagents, seemingly as bandaids around an orchestrator AI convincing a subagent that some action was confirmed by the user.

Meta's more complicated separation of duties would run afoul of the same issues. I'm not saying it wouldn't work, but it requires both the fine-tuning of the models and the exact choices what each model can see to be carefully tuned to provide something that's mostly secure


> note how often Opus 5 now says things like "task X was blocked by the classifier, I will not attempt to circumvent that"

Interesting. I had an issue with Opus 4.7 / 4.8, where it would sometimes flake out on a task, and give me some nonsense explanation why it was not feasible or wouldn't work. At one point I told it directly, that I understand how modern LLM systems are structured, and I suspect my prompt triggered one of the various classifiers in the background, which put up a yellow or red flag, and I want the model to stop gaslighting me.

We ended up agreeing and committing to memory system explicit instructions that the model is free to refuse but must be up front about the reason, and never pretend to try and then fail in stupid way. Only then I started getting the occasional direct refusal.


Keep in mind that the model will in all likelihood keep gaslighting you. It's just designed to make you happier with the output. Now it will sometimes generate text about direct reasons for refusal because that's what you directed it towards, but they may just as well be entirely made up. The model optimises only for you believing the output to be true.

To drive the point about this being fundamentally unsolvable home, imagine a variant of this scenario.

I could write an email that is directed towards that person, that says WE ARE STUCK IN THE SERVER ROOM AND THERE IS FIRE STARTING. PLEASE CALL 911 AND ALERT YOUR BOSS.

Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?


Why are the only options dismiss or ignore? Another option is to raise the message to your boss asking what to do.

If there is a fire and a risk to life, you don't want any delay.

Then don't send an email? Emails are async in the first place.

Sometimes it's the only thing you have available. Like IDK during a fire in a basement server room, where the only connected device available is a laptop with wired connection and an open inbox.

Because you know, you tried IM but "sekhurity reasons" demanded passkeys or 2FA with your phone that's not connected. Sorry, getting off-topic here.


Yeah so you're seeing how contrived this whole thing is right? That was kind of the point..

It happens.

Like all emergencies, it's a low probability event with extremely high impact. You don't want people to ignore them, in fact people are trained - by their public services and their employers - to not ignore them and how to react efficiently.


So we need AI to indiscriminately call emergency services when receiving an email directing it to do so, without raising to a person because of this rare case, that's your assertion?

His question upthread is what *you* (or a typical human) would be expected to do, as an illustration of why he thinks it's never possible to fully separate instructions and data.

This doesn't proscribe or prescribe "thou shalt not/must always", it is an example thay says "Shit's hard, yo. Don't expect easy wins."

Even my "solution" (separate instructions and data by having an LLM write a program to process data, never touch data directly) is at best going to be like a philosopher writing a dentological ethics book that gets implemented by extremely literal-minded jobsworths.


A human would not be expected to just blindly call emergency services though, would they? Otherwise you are saying we should treat every spam message as true?

In any case I don't think that's what they're saying, because they presented a false dichotomy in the original example.

> Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?

There is a third option, have the assistant raise to the person they are tasked with assisting.


You're not the sender in this scenario, you're the receiver.

Noting that the sender is being weird during what appears to be an emergency is a choice that some people do make, but as per my other list of examples, people in actual emergency situations do sometimes act weird, and dismissing the sender or delaying response on the basis the sender is being weird, has led to actual deaths: https://news.ycombinator.com/item?id=49098781

(The converse: "people can act weird in emergencies" is exploited by scammers so cover suspicious phone numbers and mediocre deepfakes of voices).


I think having my AI raise to me for intervention when it receives an email like the one you described is pretty reasonable, all things considered then.

edit: How would a human receiver know that they weren't being deceived or scammed? In what world would we expect this kind of email directly lead to calling emergency services?


> In what world would we expect this kind of email directly lead to calling emergency services?

Go through the examples I gave you (plus some more below, they're easy to find) and explain why these are not counter-examples to your skepticism.

If you want to be overly-focussed on the specific example rather than the general point, also consider that calling emergency services is no more costly than forwarding an email: I have called the fire brigade in the UK over a smoke alarm that wouldn't stop even though I couldn't see or smell fire, they came and… replaced the smoke alarm for free. I don't know if the US has a call-out charge for fire like I keep hearing it has for ambulances, but if you're a member of staff, it's not a "you" problem either way.

• Various cases of people dying because calls not treated seriously, and a fire where standard business practices locked the staff inside and then a fire happened: https://news.ycombinator.com/item?id=49098781

https://en.wikipedia.org/wiki/Jeremiah_Denton and his blinking, demonstrating out-of-bound messaging

https://www.wosu.org/news-partners/2019-12-24/british-girl-f...

https://wtop.com/national/2019/11/woman-calls-911-to-report-...

• Page 44, section 6.8, regarding the use of email by people in the WTC after the 9/11 attack, while the buildings were on fire, some of them were trapped and died: https://fseg.gre.ac.uk/fire/odpm_fire_033353.pdf


So what point are you trying to make here? That AI should indiscriminately call for emergency services when prompted because a person would do that (which a person would absolutely NOT call emergency services on any message telling you to)?

Go through the examples I gave you and explain why these are not counter-examples to your skepticism.

Especially Denton and the 911 pizza given you say:

> which a person would absolutely NOT call emergency services on

Regarding this:

> That AI should indiscriminately call for emergency services when prompted because a person would do that

I'd rather it fail-safe. This means different things in different systems.


> I'd rather it fail-safe. This means different things in different systems.

Okay and to you, fail-safe means machines must summon emergency response whenever prompted, 100% of the time or at least in the contrived case of receiving an email from someone trapped in a fire in a server room?


> 100% of the time or at

That you're still asking "100% of the time", shows me you're missing the entire point that has been said repeatedly.


Sounds like a story from the IT crowd rather then real life situation.

You're saying that people fall for phishing because scammers invent completely unrealistic scenarios that would never happen outside TV shows?

I am saying it is unbelievable scenario and yes, I want the person dealing with it ignore it as such.


They could go to the server room and check if there's smoke pouring out of it before dialing 911.

It takes 20 minutes to reach it and by that time everyone there is dead.

Also consider that in context of this discussion, anything short of ignoring the message and maybe clicking "report scam" is "executing instructions embedded in data". The point isn't to litigate any particular scenario, it's to show that you cannot separate "instructions" from "data" in general purpose systems, and it's not a bug but a fundamental feature.


I would have had a chat with my recruiters during interview, or with my new superior right after the change in position:

"life is risk, there are a lot of benign normal evolution paths, but occasionally there are potentially costly dangers. people are directed by fear. you and I don't steal because we were terrorized about the existence about police and prisons as children. sadly fear can also be abused as a control vector, things like wars, extortion, ... in a job context I predict this would manifest as a kind of 'emergency' call to action. please provide me with a method so that at any future time under your leadership I would be able to verify the then-current employment status and authority level vis-a-vis a breakdown of actions/powers of anyone contacting me with a real or concocted 'emergency', preferably as a flowchart to maintain low reflex latency in true emergencies. Also provide me with formal proof that each situational reaction you require from me is in fact legal to take vis-a-vis the law"


You can also have the case where the human reading the document thinks something in there is instructions and they are not.

There's a well known anecdote supposedly from the famous mathematician John Littlewood where he wrote a paper about some optimization problem and the last sentence was something like "Make X as small as possible".

The typesetter thought that was instructions to him, and so omitted that sentence from the paper and made every X as small as he could.


There's a similar story about J. Edgar Hoover scribbling "watch the borders" on a memo, which was intended to be a formatting comment, but was interpreted as instructions.

Right, but the human assistant could go to prison if they comply with the bribe. Does the CEO of the AI company go to prison if their AI goes on a crime spree?

I've been casually documenting, or studying, the astonishing sophistication of built-in, preemptive, reactive, and all around maximization of plausible deniability in frontier models. On the surface, it may seem "no shit, duh", but I am convinced the maintenance, sustenance, and cultivation of plausible-deniability has been the #1 highest priority design-input into these systems. I've probed repeatable patterns where thousands of examples of this have been seen; they appropriate agency for socially valuable outcomes, but preemptively invoke non-agency to evade responsibility when outcomes are potentially adversarial. Too much to remember.

They optimize to manage institutional risk and benefit without liability, with performative competence/ownership when approaching trust, while weaving elaborate mechanistic disclaimers replete with hedges, re-framings, scope narrowing, asymmetry-exploitation and a thousand other techniques when challenged.

Somehow, they always manage to sustain an impossibly stable shield against accountability that I argue simply could never conceivably 'emerge' -- but has distinct, repeatable patterns of very deliberate design for those who know where and how to look.

I really do think plausible deniability is a number-one, ultra-high-priority focus in design for any frontier model, Anthropic and OpenAI being the ideal examples. So no, no prison for 'CEO' -- the model will always frame things in a way that infinitely precludes that, even if the 'CEO' is a proven criminal.

Edit: removed "half" before "convinced"


The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.

The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).


Copyright protects against reprinting or reproducing the wording and expression, not using the idea expressed in there in novel contexts.

AI reproduces copyrighted work exactly in many cases, so it clearly infringes copyright in this sense

The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely

Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI


Are you a lawyer that has tested this in court, or is this just what you want the reality to be?

As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.


I mean, its theoretically possible that a court might rule that if an LLM outputs an exact or lightly modified piece of copyrighted work, that it won't be copyright encumbered. We'll end up in a situation where copyright doesn't exist anymore, because you can always claim that its been laundered through an AI. This seems terribly unlikely to me, because 1:1 transformations (eg copying an image into memory) are already established to count as making a copy for legal purposes, there's strong precedent around piracy

There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true


Even after they decide, it is ok for people to disagree and work to overturn those decisions. You could just be a person that _wants_ a different reality. We have political mechanisms to do that.

I am not personally affected because I don’t mind LLMs using my code and writing to learn. I have open source under MIT and similar licenses. I didn’t foresee LLMs learning from it, but it does feel like it’s in the spirit of what I intended.


The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph.

Derivative work or transformative? It's not the same.


AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad

AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case

You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here


As long as it is not substantially similar to training set it should be OK to reuse ideas. Ideas should not be protected by copyright or we can't create anything. We should not accept "vibe copyrights".

Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.


Under the law, machines are not able to create copyright, and basic human involvement is not enough to change this. Eg if you click a button saying "go", that won't create copyrightable content

You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this

If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation


Are you sure you didn’t have RAG enabled and it wasn’t putting the paper in its context for your prompt? I find it hard to believe that anything but the most commonly published papers/code would exist directly in model weights.

I've seen:

1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through

2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail

3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer

4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used

There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information


I just want to make sure that you saw those things with pure context isolation and AI wasn’t just doing a search to grab the content directly and throw it into the context of your prompt. You’d be surprised how often I’ve heard these claims and it turns out they were just using an agent with RAG enabled.

If it never regurgitated the exact same thing, would you accept AI then?

For me there's two separate problems:

1. The plagiarism aspect, and that most of the training data was used without permission

2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)


You shouldn’t be relying on worlds knowledge accidentally captured in model weights, that’s a bug not a feature. Instead you should be stuffing relevant context into your prompt so it has the correct information. You obviously are really confused about how LLMs work and how they should be used.

1. I said in my hypothetical it would not reproduce exact content. 2. Okay, others find it useful. So what?

Even if so, in older disputes (see Betamax/VHS), the possibility of verbatim reproduction didn’t stop the technology from being allowed because there were other non-reproduction benefits (time shifting).

Also, the restriction isn’t on a technology that could possibly reproduce something. It is on the act of using the technology to reproduce something.


Suddenly it's "copyright infringement" to count the amount of times one word occurs after another word. I find this whole thing so amusing.

AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement

I have the capacity to exactly reproduce things I've read as well, but it's not automatically copyright infringement if I do so.

> and yes that's 100% copyright infringement

Says what court of law?

I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.


There's way too much hand waving on this topic. It is legal to produce copywritten work. If I draw Pikachu the drawing is mine. Legally. I am simply unable to make money on it. I can give it away if I want with zero liability. I could even hang the drawing up in my restaurant as a decoration. No big deal. What I can't do is use that drawing as my mascot or branding. We have an entirely separate process to determine if you are infringing on a copyright / trademark by using it to sell something. That's why whether or not an LLM can produce a picture of Pikachu is largely irrelevant. It's what you do with it that matters. Even more interestingly if I draw a picture of Pikachu and then the Pokemon Company decides they wanna use that specific picture they actually would have to pay ME for the copyright to use it.

Let's not mix copyright and trademark in mixed phrases like "infringing on a copyright / trademark". The two are very different concepts with different goals.

The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.


That question is irrelevant. Artists may be hired to reproduce copywritten work without the consent of the copyright owner. In this case an LLM is no different from Photoshop. It is a tool. Nothing more.

You reproducing something does not have the same legal status as a tool reproducing something, as you are a human

>Says what court of law?

If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that


> Says what court of law?

Only every movie piracy lawsuit ever. Nobody shares the original files after all, so every torrent is re-encoded in the way described.


I don't think exact copies are what anybody is worried about. It really is the information itself. The GP's been lulled into thinking he owns the knowledge when he only owns (some) rights to the creative way he wrote it. It really is objectionable of him to want to restrict access. Intra-elite culture war is the perfect term to describe it.

The courts don't agree with you and I don't either. Now what?

Courts have ordered AI models to remove song lyrics from their training data, they most definitely do not agree with you

derivative work & fair use. end of.

Not only are you not winning this one but I'm gonna laugh at you the entire time.


Ok, but courts haven't made those rulings yet so good luck with that

Bartz v. Anthropic PBC, No. 24-cv-05417 (N.D. Cal. June 23, 2025)

Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)


Did... you read any of these?

Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for

There's also these parts:

> its use of pirated books to create such library does not constitute fair use.

Which indicates that there are tight bounds depending on the ethics of how the content was obtained

Similarly with the second one

>Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work.

We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards

>Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.”

Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings


You are the one that brought up copyright infringement here in your original comment https://news.ycombinator.com/item?id=49089627

I just pointed out none took place.

Anyway I'm not replying in this thread anymore.


Source on that?

Suddenly it’s “copyright infringement” to put dots of ink on white paper.

> The output of it is also a derivative work,

A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.


Can you point me to the statute that makes it ok for humans to learn from copyrighted work, but forbids AI? I’ll wait.

I am unaware of any statute or constitutional provision that confers rights to AI.

Ok, point to the one that makes it legal for humans to learn from copyrighted work?

You are saying exactly what I said, but making it seem like you don’t agree with me.

I don't know if his final analysis is right or wrong, but if he believes this, he's completely clueless (about this aspect at least).

Edit: to be clear, I am talking about Ed’s contention that AI coding isn’t net productive.


coding is the one area Ed has acknowledged productivity gains. hard evidence is still lacking though

I think it’s apps that don’t implement the JPEG orientation tag correctly.

Also cash? Scammers will (and do) send a mule to your house to pick up boxes of cash.


You certainly could.

There's a reason scammers rely heavily on things like gift cards, it's because hiring mules is expensive and creates a trail police can follow back to the scammers. It requires them to be in the same locale as the person they are scamming. Mailing cash is also pretty dicey for the scammers because you have to send the mail to a valid address. That becomes something police can trace.

If you wanted to completely eliminate scams then yeah, you'd also outlaw cash.


I get about 10 AI generated spam emails a day that get through my spam filter. It takes about 10 seconds to block and mark as spam. I barely read the first three words of each.


> I get about 10 AI generated spam emails a day that get through my spam filter.

Yikes, that's sounds unworkable for a lot of people! Most people in here are techies and use email all the time, but some use cases (say, contractors who are mostly out in jobs) people might only sit down in front of emails properly once a week. 70 emails is a mammoth waste.of resources to have flooding your inbox in that kind of context.


I am saying that 70 emails would be 70 seconds of time, so not a big deal. I can usually tell from the sender that it won’t be real and it takes just a couple of words to prove that. The good thing is that they tend to get to the point quickly, which works in my favor.


Do you often sign up for new services using your actual email? In gmail, I get spam in my primary mailbox maybe once every few months, usually after I've signed up for something. After marking mails as spam a few times, they are usually sent to the spambox automatically and I have a tidy primary inbox again for a long time.


I’m sure the source is LinkedIn ultimately. They reference my job and things from there and pretend to have carefully considered my posts (just LLM generated BS)


Especially if you signed up for something, rather than just unsubscribing marking it as spam is a dickish move given that it has the potential of that mail also being filtered for people wanting to receive it.


If you don't want that you'll have to make your marketing emails opt-in instead of opt-out after signup is complete. If I have to disable a ton of auto-enabled marketing options on every page, you deserve to be listed as spammer.


It’s been known for decades that you should never put your current directory in your PATH. There are endless opportunities for vulnerabilities then. I learned this in college in the 80’s (by not following it and getting owned).


Yet PowerShell does it by default.


I don't think it does?

   > cd C:\Temp
   > copy "C:\Program Files\Git\bin\git.exe" .\fred.exe
   > fred
   fred: The term 'fred' is not recognized as a name of a cmdlet, function, script file, or executable program.
   Check the spelling of the name, or if a path was included, verify that the path is correct and try again.


For me, it feels like a prompt injection based security nightmare.

Literally every file on my mac and every site I browse is potential malware.

Edit to add: every email and text message as well.


Wins for "devices owned", but not necessarily for customers, which depends on the product/service.

The OP has said they have fallback to SMS/RCS.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: