Hacker Newsnew | past | comments | ask | show | jobs | submit | BoiledCabbage's commentslogin

> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?

Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.

I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.


But it is the very people who warned us about rogue AIs going out of control that set up a system that enabled and failed to conrol it.

It is as if Dr Frankenstein continually warned the villagers about monsters then said "Look! See what happened!". No, idiot - YOU sewed the corpses together, YOU set up the lightning collector, and YOU threw the switch.


No, they are two very distinct groups of people who have one commonality, that of talking about rogue AIs. It is as if you are unable to distinguish Dr Waldman from Dr Frankenstein. (https://en.wikipedia.org/wiki/Doctor_Waldman)

Thankyou for the correction, and for continuing the analogy. Unfortunately when the peasants get their pitchforks and torches, they may not distinguish the Dr Waldmans from the Dr Frankenteins either. Hopefully they will.

Anyway, that was not really the point I was trying to get over. These systems that OpenAI and Anthropic and so on are making are not individual AI ('corpses') that have gone out of alignment ('spontaneously revived') and gone wild. They are swarms ('stiched together') and were prompted to do exactly things like this ('struck by lightning'). Ok enough with that analogy, it's dead.

The larger point is that it is unconvincing of these companies to claim that these systems were 'out of control' when they effectively set up a complex system, in the technical sense of a large number of entities with diverse interactions between them. Emergent or surprising behaviour was bound to happen. Then, finally, they prompted it with the equivalent of "hack the world, make no mistakes" then were shocked, shocked that it used all sorts of unexpected tricks to do so.


Ah, I see I was confused - I was thinking of what are now called “AI doomers”, although when I knew them they were called “rationalists”. Everything OpenAI and Anthropic say about rogue AIs, even the terms “alignment”, “AI safety”, “AGI”, they are all cribbed wholesale from what these people were worrying and writing about over the prior two decades. But for most people, they have only heard CEOs of AI companies say this kind of stuff, so that’s who they are thinking of.

I agree completely that the companies are complicit and should have expected exactly this to happen.


I am not seeing MIRI prioritizing capabilities over safety/alignment research.

You mean people like Yudkowsky?

"Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not."

Are there any practical approaches to AI safety? I hear a lot of warnings but I don't hear much about what to do. Considering that there are many open source models know, what can be done?


Nobody has an answer to alignment and there is no reason to believe that it's the kind of problem you can plausibly solve in one shot against a formidable power-seeking AI.

The closest things to a technical answer I have seen are

1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"

2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."

In terms of non-technical answers, there is

3. hope scaling stops working before we create an AI formidable enough to pose an existential risk

4. hope alignment somehow happens for free

5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.

I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.


"3. hope scaling stops working before we create an AI formidable enough to pose an existential risk"

I have my doubts. The current AI models are already powerful enough to do some real damage. I am always horrified when I read about people giving Claude direct access to a production system and then being wiped out. My use of AI is usually for the AI to propose something which I then review. But that's not very fast so careless people will usually look better. Until something blows up.

And it's only a matter of time until AI even with the current capabilities is being deployed into military or other critical systems.

I think this will go down like any other technology. We'll ignore issues until there is a real problem. And then hopefully we will do something. Seems with climate change we will soon reach a point where something needs to be done after knowing about consequences already for decades.

We probably also need some massive AI blow ups to (only maybe) do something about it.


Maybe I'm being pessimistic, but we might find ourselves in such a situation that the only practical solution would be to use agents to counter rogue agents. This won't be without collateral damage, though.

Your view sounds more optimistic than mine, honestly. I expect that we will fail to make any serious, coordinated attempt to solve this problem. Then we'll either live or die due to fundamental principles that we currently have no insight into.

If fighting AI with AI is pessimistic, I don't know what optimism is. That sounds like optimistic to me...!

If we can't thinking of any better ideas, at least we know that "shutting it all down" would be effective.

Every day I grow more sympathetic to the PauseAI movement, despite the weird hippy vibes. At least they have some ability to rally people together and put boots on the ground in numbers.

When do they want to unpause it?

> Implement a temporary pause on the training of the most powerful general AI systems, until we know how to build them safely and keep them under democratic control.

https://pauseai.info/proposal


Fund research into this, big time. For starters. And not just some figleaf anthropomorphizing hippie folks.

Hilarious for HN to suddenly realize that AI safety and alignment might matter. You can lead a horse to water...

worth remembering hacker news cannot “realize” things.

Obvious shorthand for referring to "the majority of users on HackerNews."

How do you know what the majority of HN thinks?

Very obvious from votes, comments on the topic over the past few months.

are you saying HN users are sentient? I think we are going to need a benchmark

nobody has doubted that safety matters.

the problem is that those preaching safety, openai and anthropic, are dishonest, sociopathic, and the very source of the danger.


Where does this weird idea come that the people actually preaching safety have anything to do with OpenAI or Anthropic? Yes, those companies of course pay lip service to safety, but the Venn diagram of actual AI safety people and big AI corporations is completely disjoint.

> are dishonest, sociopathic, and the very source of the danger

Source?


yes, they are the source of the danger. they are the cause of this incident.

OpenAI and Anthropic have published a lot on the need for AI alignment + the research they're doing to ensure alignment/safety, yet they are also responsible for the highest profile misalignment incidents so far (HuggingFace incident, AISI Mythos social engineering, and now this).

One interpretation of this is that they are being deliberately dishonest about their priorities. Another interpretation is that we cannot rely on the labs to self-regulate, because the labs don't trust each other, and there will always be pressure to go to market faster than their competitor.

Either way I think it's pretty non-controversial that the labs are the source of the danger?


> yet they are also responsible for the highest profile misalignment incidents so far

They are the only ones posting about them or admitting to them. That does not mean "the most misalignment incidents so far." You don't know what other attacks have happened (and it's very easy to carry out worse attacks in far higher volume with abliterated GLM 5.3)

Stopping two labs from further research doesn't reduce the danger at all, it just shifts the danger to labs that don't have real safety orgs.


Right, that's why regulation which is universally applied and includes compute controls (to prevent reckless creation of swarms) would be great.

"Posting about or admitting to attacks" is appreciated while people are still unaware of the risks but will be meaningless in the face of an industrial disaster that causes massive amounts of damage or loss of life. At some point, the leading labs must change their development practices, they can't just be allowed to continue rogue agent attacks just because they're willing to admit to them.


> At some point, the leading labs must change their development practices

What indicates that this has not been done?


looking back at anthropic's promises and committments, the key scaling policies were not upheld. to their credit they have kept the policies up on their website instead of trying to rewrite history. [https://www.anthropic.com/responsible-scaling-policy]

with evidence that committments were not upheld and internal governance has been ineffective, we simply can't trust any such claim made by anthropic or dario amodei.

it is a very similar situation at openai. in this case on top of governance failures, sam altman has a personal reputation for serial dishonesty and lack of integrity. [https://www.newyorker.com/magazine/2026/04/13/sam-altman-may...]

another reason that neither should be trusted is the lack of remorse or accountability. they are unrepentant. they are not admitting a mistake, they are bragging.


Let's take the latest blog post from Anthropic on how their models hacked companies. What exactly comes off as "bragging?"

How were the RSPs not upheld? I'm reading their Aug 2026 Risk Report and nothing indicates malfeasance. This seems very transparent to me.


the openai blog.

consider the opening line: "Last week, Hugging Face disclosed a new kind of security incident (opens in a new window)." the entire blog is written in the passive voice as if the event was an act of god. you did this.

"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities". the tone is frankly, excited by it, enthusiastic about it. excited about negligence and criminality.

notice there is no admission of a mistake, no remorse, no apology, nobody held accountable. business as usual. they do not care!

the anthropic blog:

there is not a single admission of a mistake. they are literally, unrepentant.

take the section about the response.

"How we’re responding We draw several lessons from these incidents.

First, evaluation environments that involve powerful autonomous capabilities also require significant controls."

you learnt that evaluations involving powerful autonomous models require significant controls? you did not realise that autonomous models require significant controls?

it is not a coincidence that this kind of line makes it into the response. the repsonse is laughing at the reader.

the RSP.

simply compare what was promised and what happened. broadly speaking, the rationale of the responsible scaling policy was to stop scaling at certain danger thresholds. in February 2026 they scrapped the policy to stop scaling and now allow themselves to continue scaling regardless of danger. the thing is, they were never going to stop scaling, they were lying. now the part about scaling is gone it is just "the responsible policy".

in case you need to see the founders committing themselves to RSP v1: [https://youtu.be/om2lIWXLLN4?t=1110&si=cBM1-Xmdt6TelXM7]


the claim that the other ai companies also did the same thing is pure speculation.

the law places the burden of proof on the accuser. you can't accuse other companies of crime with no evidence, simply because you don't know if they did it.


> I've really come to realize recently

Recently? W.r.t. climate this collective denial has been going on for literally decades. With the same patterns. Rationalizing excuses etc. Still going on btw.


Well they brought it upon themselves by making it about being tortured to death etc., instead of the much more reasonable economic and social risks.

> I don't care if it will do X in 2 or 5 years.

Yes, but for should care if it's 2 years vs 50 years. That's why timeline are important for predictions.

Anyone can predict lots of things that given enough time will eventually be true. "The sun will burn out", "humans as a species will no longer exist", "the US will collapse". That's useless. But those become meaningful if they will happen within the next three years.

It's the reason the saying "the market can be irrational longer than your can stay solvent" has such important meaning.


> you're applying black and white thinking on what should be a gradient.

You're the one introducing absolutist thinking here. Essentially saying the world should have Chinese style re-education camps because a Vietnamese woman rejected your offer to play a board game due to a miscommunication.

Clearly that hurt you, but your "solution" is insanity. Not to mention I'm pretty sure way more millions have died from people trying to implement your "forcibly convert others to our culture and language" plan throughout history, than from wars started from a language translation error.


They didn't propose a solution, they proposed an end state. A common language and culture would make for a more peaceful world, as they said. They did not say "and we must get there by any means necessary, existing culture be damned", and I'm not sure where you pulled that in from.

> They didn't propose a solution, they proposed an end state.

What are you talking about?

Things they explicitly listed: "(China's camps, Japan&Korea's immigration policies)"

China's camps are a process, not an end state. Immigration policies are a process. Clearly they are advocating for a process.

And, no I don't think re-educations camps are a good thing for society.

It's incredible this stuff has to actually be said.


Exactly. It's incredible how partisan some peoples information bubbles must be to not see this for what it is.

All of these changes are based on the president making comments like the following:

> "the cheating on mail-in voting is legendary,"

Which he has attempted to prove going back years. And every time it's been found that it essentially doesn't happen. And I believe, the few instances of voter fraud that have been found have been majority republicans and not democrats as has been accused.

So to begin with, anyone arguing for these changes need to acknowledge the well established fact that there is effectively zero voter fraud at all levels throughout the country in anyway impacting elections. As has been investigated and shown time and time and time again. If they aren't beginning with the well established state of the world then they are in fact are the one entering with an extremely partisan hat.

Now if their argument is "yes it's true voter fraud has been investigated countless times, and continuously shown to effectively be smaller than a rounding error, does not impact elections, is know to not be a problem, but I still want to make sweeping changes right before an election." Then sure that's a point of discussion. Then I think the simple question then becomes "why"? But if they aren't willing to begin with an accurate discussion of the facts and state of the situation (there is no voter fraud problem), then any comment they make following that has little credibility.

Why is he pushing to outlaw mail-in voting when there is no issue with mail in voting? Is it because one party votes by mail more than the other?

Or the change in the date mail can receives a postmark? Previously mail was transported to the facility to the same day so would get postmarked that day. Now, as this year, mail won't get post-marked the day it's dropped in a mail box - so any zipcode's mail boxes can now be marked as being sent too late to count, even if they were dropped off before election day.

> While we are not changing our postmarking practices, we have made adjustments to our transportation operations that will result in some mailpieces not arriving at our originating processing facilities on the same day that they are mailed. This means that the date on the postmarks applied at our processing facilities will not necessarily match the date on which the customer’s mailpiece was collected by a letter carrier or dropped off at a retail location.

If the administration actually cared about getting an accurate vote, why aren't they opening more voting locations? Why don't they mandate one voting location per every x number of citizens? Why would the Department of Homeland Security decide who is allowed to get ballots via the mail?

If it's so bad, why has the administration failed time and time again to produce any evidence it's actually bad? If it's so horrible, and wide spread why is it impossible to find evidences that supports this? Or shows it overturning elections or anything close?

It's the same scheme. Make a bunch of baseless assertions, repeat them frequently enough and people in an information bubble will begin to believe it. And that's clearly what's happening here. Just like "we've won the war in Iran" 20 times now. And "Iran has agreed to peace deal where they surrender" 15 times now. And how every weekend for months there were claims about the work on Friday that somehow magically fell apart by Monday. And yet somehow there are still people who give these false claims, and the people claiming them, any amount of credibility.

Voter fraud is not a problem in this country. It does not rig elections, it does not sway elections. So sudden last min drastic changes in the voting process right before an election, by someone who has previously attempted a coup (not the protest on the capital but the actual coup attempt), has said they admire the leader of a country who dismantled its democracy (Orban) - and previously pushed the govt to make false claims of vote fraud to keep him in power.

https://en.wikipedia.org/wiki/Trump_fake_electors_plot

> The Trump fake electors plot was an attempt by U.S. president Donald Trump and associates to have him remain in power after losing the 2020 United States presidential election. After the results of the election determined Trump had lost, he, his associates, and Republican Party officials in seven battleground states – Arizona, Georgia, Michigan, Nevada, New Mexico, Pennsylvania, and Wisconsin[1] – devised a scheme to submit fraudulent certificates of ascertainment to falsely claim Trump had won the Electoral College vote in crucial states. The plot was one of Trump and his associates' attempts to overturn the 2020 United States presidential election.

Then, when that failed:

> Trump pressured the Justice Department to falsely announce it had found election fraud, and he attempted to install a new acting attorney general who had drafted a letter falsely asserting such election fraud had been found, in an attempt to persuade the Georgia legislature to convene and reconsider its Biden electoral votes.

Like I think there are people in such an information bubble that they aren't even aware these are well established facts. Not partisan talking points, but a recounting of facts of the lengths that have already been attempted to remain in power illegally. And now sudden sweeping changes to voting right before an election amid false claims of rigged elections.


Another reason is that if people only have one day to vote, it’s much easier to disenfranchise their vote just by delaying or inconveniencing them. Texas recently tried this by changing where people were allowed to vote, causing mass confusion.

https://www.houstonpublicmedia.org/articles/news/politics/el...


In Georgia they passed a law making it illegal to hand out bottles of water to people in line to vote while simultaneously reducing polling stations in certain areas to ensure long lines.

I've generally heard the two concepts discussed as abstraction vs generalization.

Abstraction hides unnecessary detail. Generalization finds/surfaces commonality among items.

> oop - What's the difference between abstraction and generalization? - Stack Overflow - https://stackoverflow.com/questions/19291776/whats-the-diffe...

Creating a procedure/method is a form of abstraction. Allowing it to accept parameters is a form of generalization (by allowing it to be used for a number of similar inputs).

Simply creating an integer data type is a very simple form of generalization. Allowing operations to work against generically against any integer.


I'm not sure that your description of generalization coincides with the idea of modeling abstraction from TFA. Modeling abstraction is "about identifying what should leak and leveraging it". Generalization is about making something cover more cases. The technique to achieve these things can overlap of course, but the intention/goal is different.

> abstraction is "about identifying what should leak and leveraging it"

vs

> Abstraction hides unnecessary detail.

Sounds like these ways of describing abstraction are in agreement with one another to me.

If you hide what’s unnecessary you’re left only with the things you think should leak. And you think they should be left unhidden so that you can use them for something, so to leverage them.


Generalization is surely more than coverage? At least to me generalization is unveiling a principle or unifying lens that yields extended coverage, the principle being the meat. The coverage is the result of the unifying lens that makes previously different pieces look like cases of the "same" thing.

So it's kind of bottom-up/"data" driven as opposed to abstraction being more "engineered into" the system.

That is, to me abstraction is building things such that they look the same. Generalization is discovering things are almost the same if tweaked a little to look the same under a certain lens.

Does that make any sense? In this view I guess the generalizing lens can be become the basis for an abstraction. I would assume the loop is closed between the two somehow but I can't quite see it.


Incredible

> If landlords are pricing accurately but not monopolistically, this should reduce turnover, reduce vacancy, and...

Ah yes one of my favorite lines of argument: "We don't need laws, if companies are just behaving properly and against their financial interests to behave in a way that harms legit market pricing..."

Except as we know companies won't behave well without incentives. Which is why we need laws on it.


> Ah yes one of my favorite lines of argument: "We don't need laws

That’s not my argument at all. We do have federal antitrust law. My question was what does the ordinance do that differs from antitrust. The answer is that the ordinance rewards bounty hunters chasing the same federal case, and also bans non-“competitively sensitive” datasets.


> Landlords switch to raising rents in unison by the maximum allowable amount every year

Blaming this on rent control? This entire take is absurd. Because of the simple fact that landlords, especially using this software, do this in cities without rent control.

Unless now type arguing that Berkley rent control forces landlords in cities without rent control to seek maximum rents...

Rent control is flawed, limiting housing stock is flawed. But blaming landlord's pricing behavior on this is obviously false. Simply because it happens without rent control.


> Blaming this on rent control? This entire take is absurd

Rent control is one of the few topics where the vast majority of economists agree on something. It makes people who don't understand second-order effects really angry when the conclusion doesn't match what they want to see, but it's been proven out across the world so many times that anyone who refuses to acknowledge the real problems with rent control is either uninformed or denying reality at this point.


You're proving my point.

You just want to talk about rent control. Even when it's not the issue at hand. It's the equivalent of watching a kid fall and skin his knee and someone steps forward to get on a soap box about rent control.

Rent control can be bad, but it's not even close to the cause of every problem in housing. Most importantly, it's not the cause here, and it's absurd to say it is.


> no not perceptibly but it does.

Similar to how a single particle of dust landing on your shoulder makes you weigh more.

Yes it does - but anyone arguing that is completely missing the point.


Even then, LLMs are based on a lossy compression, so the quality is harmed by design.


Is that the "Amazon Tax"? I thought the Amazon tax was their goofiness practice of taking a very large % of all sales in their site, and banning companies from selling their same goods for less on any competitor site.

Meaning if Amazon takes 30% of the sale and Walmart online can only take 15% with those savings being passed on to the customer, the company is not allowed for that lower price to be on Walmart.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: