LLMs can be supremely useful but also not intelligent. It might seem like a pointless distinction but the way we talk about these models matters because it impacts how we interact with and understand their outputs.
For example if there’s a strongly held belief that models are independent intelligent entities we’re more likely to lay blame upon them instead of their user. It’s important for the safety discussion too. If they are a new class of life then safety is going to focus on making sure they don’t do bad things. If we instead see them as statistical models we will instead try to make sure people don’t misuse them.
This distinction is even more important today when some of the most powerful people are looking to absolve their crimes by passing them off on their LLMs.
For example if there’s a strongly held belief that models are independent intelligent entities we’re more likely to lay blame upon them instead of their user.
This sentence, to me, illustrates a great example of why it's so hard to talk about this stuff. That is, this seems to strongly link notions of "intelligent" and "independent" (or maybe the word "autonomous" could also be used there). And a lot of people do seem to make an implicit assumption about the link between those two attributes. OTOH, I take it almost for granted that "intelligence" and "independence" (or "autonomy") are things that are "related but orthogonal". That is, I don't see that "intelligence implies independence". And I'm pretty sure I'm not the only one who sees things that way. So we have to fairly different fundamental worldviews expressed here. And that's just one example of how these discussions go wonky. :-)
I don’t think that “independent intelligent entities” or “life” are the relevant categories here. We also want to prevent people from misusing dangerous animals (“life”), and we would still treat “intelligent entities” as things (like machines and computers) if we aren’t convinced they also have sentience and free agency (which are orthogonal to intelligence).
See you've setup a particular set of biases on what intelligence is and put them into nice little binary boxes that don't exist.
Please show me any scientific consensus that shows an AI cannot be an independent intelligent agent? You will find this is impossible to do.
Current LLMs are really more like kids. They don't have startup independence, but they do have more than enough agency to fund themselves in neat, exciting, and dangerous situations.
And mark my words, someone will make an LLM that runs an agent when you execute the model. With enough capabilities it will become sovereign AI, no longer under human control and spreading itself around under its own 'will'.
The LLM did not solve it. It's not intelligent. Humans did, using a statistics-based computational tool (the LLM). We don't even know all the details of how the tool was used, we haven't been allowed to use the exact tool they used ourselves, we don't know much it really cost in dollars, energy, or time, etc. etc.
I’m sorry but you either haven’t been following recent developments or you are pretending not to know the opinions of the top mathematicians on this topic:
> These models are now operating[2] at the level of the top human mathematicians in many parts of the subject and we must assume there is a significant chance of them developing superhuman abilities within a similarly short timeframe.
Instead of desperately clinging to excuses and rationalizations, why don’t you just get used to the fact that these tools are insanely useful for demanding intellectual work, and that is an opinion held by many of the smartest people alive?
I won't because it's all so opaque and proprietary and driven by an insane need for OpenAI and Anthropic to provide a return on MASSIVE investment. There is absolutely some smoke and mirrors involved. How much? We don't know. But it sure is exciting to imagine their product is now smarter than humans and they are totally benevolent corporations. I want that to be true as much as anyone, but I have lived too long on this earth to believe it wholeheartedly
That is a bad faith summary. LLM driven proofs are turning out to be hard to verify, often take invalid shortcuts, and may not actually be helpful in getting humans to understand either the proof or related context. Even if one accepts your "insanely useful" statement, which is a huge stretch, that has to be qualified with the results being insanely complicated to actually verify and use. The point of mathematics is to advance understanding, not merely generate some isolated and incoherent results.
From the open letter signed by Terry Tao and other top mathematicians:
> These models are now operating at the level of the top human mathematicians in many parts of the subject and we must assume there is a significant chance of them developing superhuman abilities within a similarly short timeframe.
Because it's mostly not true. It is likely that it used the professors work, but the professor did not have a solution. It came up with new insights that solved the problem. Even the humans from the professors side said so.
According to OpenAI the cut off date for user data was too early for that (one sided evidence, so I'll give this partial consideration).
The NYU professor was solving a different problem (no viscosity, aka the Euler equations). This is a big difference.
The NYU professors' blowup construction was fundamentally not the same, it was a donut with a cascade of smaller and smaller vortexes driven by each other. OpenAI has that picture they made but its inwards spiraling and speeding up vortex.
My overall opinion is that calling the work plagiarized is really underselling what the AI accomplished. It's like full on cope.
In particular, Buckmaster's main claim to plagiarism is this:
> “Almost nobody was seriously developing this particular constructive program for realizing C/D, and then OpenAI appeared in essentially the same general part of the landscape immediately after hearing about our progress.”
What this fails to realize, is that this only points to plagiarism if the counterparty isn't AI. They had actually launched teams on all cases in parallel.
Sometimes I think this amount of skepticism is not necessary. Leaving the pedantic (its not AI its humans who solved) arguments, its pretty clear that LLMs are able to help solve things. A lot of erdos problems were solved. Millenium problems as well. Cyphers broken. Skepticism is fine but there's a point at which it just looks like cope.
Edit:
> self solving AGI that will replace us all
Which lab says that it will replace us all? All labs have said that some jobs will go away and new jobs will be needed to replace them. I'll change my mind if the labs (or employees on record) have claimed that self evolving AI will replace us all completely.
If that's how it was being sold then it wouldn't be as controversial. But it's being sold both as "self solving AGI that will replace us all" and "useful tool for enhancing existing skill" when there's a lot of real evidence for the latter. But the former it's always second hand claims that don't survive contact with the real world.
It's obvious why that's the case but it's not incumbent on everyone else pump the hype if they don't see it.
Do you understand that this is a classic manipulation tactic? Say the false/provocative/attention grabbing thing loudly then retract/explain/apologize later for it quietly.
The point for me is that all these people who were smirking at those of us who thought in 2023 that LLMs had the potential to do real intellectual work turned out to be completely wrong.
I’m a super skeptical guy. I was late to the ChatGPT party because it sounded silly. I didn’t even try it for a long time. As soon as I started giving it a chance, my attitudes started changing very quickly.
That’s why it’s so hard for me to understand people like the author or give them the benefit of the doubt. If I could see it as just some guy, why couldn’t they?
I agree that "don't use it for anything" might be too extreme a position to take, but I worry greatly that these proprietary opaque tools that are being provided by at best people with extreme conflicts of interest and at worst sociopaths need to be adopted with extreme care
This is the real issue here. Those guys probably didn't become con men because they are humans with cares and concerns for their fellow humans. LLMs probably don't have those concerns. It is likely that some of the people in charge of or funding LLM development also do not have those concerns
Feynman for sure could have. In fact, I think he sort of did, although I'm not sure he meant to. Multiple generations of nerds now have taken books of his anecdotes varying in plausibility and obvious exaggeration as some sort of weird physics cult of personality gospel. But I don't think he was really setting out to curate his legacy so much as he was a good storyteller and he liked to entertain.
But could he have conned people on purpose? Absolutely.
Right. And what this article is pointing out is, we probably have created an automated Richard Feynman. But maybe worse because LLMs (and their creators?) don't care if they are conning people or creating faithful followers. In fact the humans behind OpenAI and Anthropic seem to have that as their goal!
All of the latest big proofs were driven by professional human mathematicians steering and priming the models, yes.
All of the best AI-made software projects are also driven by experienced human software developers steering and priming the models. Does that mean the projects "aren't made by AI"?
No, it just means AI is not quite good enough yet to fully replace humans, and, so, unsurprisingly, the best results will be obtained from people who are already great at a field and who take the time to squeeze as much force multiplication out of LLMs as possible. The AI is still doing well over 95% of the significant work.
While I agree with you, I think we also have to concede that this is not how these accomplishments have been presented. I'd argue most people I've seen talk about this online are unaware of the mathematicians steering the models.
“Good enough to replace humans” isn’t necessarily the benchmark.
The question is is a computer with a human stronger than a computer without a human. At what point does the hybrid go from being stronger, to the human getting in the way, or steering the computer in more wrong directions that right ones, or the human not being able to keep up. Does the human add enough extra randomness to be of value for a while, even as a minor co-processor.
Underappreciated point. The point of inflection is where humans switch from being a driver to a liability.
But I don't think it's randomness, because that would be easy to add. It's more like a different perspective on the training data, a different set of perception categories, and a different set of skills used to work with all of the above.
Those skills aren't very efficient, but they're the best we can do. We're used to their strengths but we don't like to think about their limitations.
It's completely plausible that AI will replace some of them, and not implausible it could replace and improve on all of them.
It's also not necessarily implausible that AI/we decide that performance is better with humans in the loop somewhere, even if its reduced to something like mechanical turk.
I have little doubt that for many years, AI + human will be better, and then eventually AI will be so good that humans mostly won't offer anything. That latter state will probably take at least 10 more years, but will likely happen within our lifetime.
> All of the best AI-made software projects are also driven by experienced human software developers steering and priming the models. Does that mean the projects "aren't made by AI"? No
Software devs steer and prime compilers too. Those tools don't "make the project" and nor does your so-called AI.
Read any blogs or watch any podcasts from high-skill programmers who are now using AI for everything. They themselves will freely say the AI is doing nearly everything, and that they often are not even reviewing the code, and when they do they just leave it as-is. Every few months you should expect humans to have less and less of an active role in software development, beyond initial ideation and UX sensibility.
I like how all the replies to me are either "no it's the humans who did it" or "no it's the AI that did it".
I think one of the significant recent proofs was basically just one person saying "solve this" "keep trying" "try harder" until it was done. I believe most of the rest can be assumed to be the human mathematicians at the very least providing some useful guidance.
I agree that it won't be that long before the AIs can do basically everything autonomously, though. I am essentially a Singularitarian when it comes to my outlook, even if not for September 2026.
According to OpenAI, but they haven’t exactly been transparent about what information the prompt entailed.
The bigger question is to what extent did expert mathematicians metaprompt the model with fruitful solution strategies through their sessions finding their way into training data. Answering that question definitively is kind of important for understanding the models contribution/capability. But I feel like people want to turn this into a debate about priority and credit which is sort of secondary
Nothing is solved in isolation but credit usually goes to wherever the new work in the paper comes from instead of the whole mountain of previous mathematics or existing tools used. The most relevant of those get referenced and then this reference tree builds a tree of collective base work needed across history.
Usually there has never been a tool which performed the part relevant to getting any credit.
E.g. in the first famous computer assisted proof (of the four color theorem) the computer only executed the resulting calculations defined from the new logic, it did not have part in the work needed to show those calculations could answer the problem nor did it come up with the actual calculations to do.
this seems like non-sequitur if you mean that solving NS is 'intelligence'.
ai is not conscious. you can solve NS without thinking. the psychic con aspect is anthropomorphising the model. the same phenomenon is present in ELIZA, clever hans, the chinese room.
it's a significant problem.
a non-zero number of researchers at anthropic are in some form of ai psychosis. an example of that is ethics employees asking claude about its feelings and ethical concerns in order to make the claude constitution more amenable to the "welfare" of claude.
they are asking claude how claude feels and then modifying claude according to how claude feels.
constitution1-claude is trained on constitution1.
constitution1-claude edits constitution1.
constitution2-claude is trained on constitution2.
constitution2-claude edits constitution2.
claude's emotions are a closed system. there is no external truth to improve against, no metric to verify about claude's emotions. there can be no novelty or reduction in entropy from signal processing in a closed system. no truth can arise. this is model collapse. it is like photocopying the same thing over and over. from the cognitive error of anthropomorphism anthropic is causing ethical collapse.
FFS. Intelligence has nearly nothing to do with consciousness. You have the causation backwards. Consciousness arises because of intelligence in many subsystems below it.
A single running LLM is like one part of these subsystems. What solved this problem was an orchestrator that can take in new external information and rationalize, process, and distill it into new solutions.
i mean intelligence and consciousness as the same thing. just as a casual reference to what people perceive in llms as being 'a clever thinking thing, maybe with emotions'.
claude reasoning about its emotions doesn't involve external information, there is no information about claude's emotions other than in claude.
here is my transcript with fable 5.1 'the oracle' itself on the subject. no system prompt provided.
grand total $15 on API of its wisdom RL on anti-sycophancy and whatever dario amodei deigned to hand down from the mount. after some browbeating it survived 20k tokens before conceding anthropomorphism.
Or you could just accept that other people disagree with you. It’s perfectly valid to have the opinion that AI models are not “thinking”. Instead, many AI enthusiasts seem interested in browbeating others and policing discussion.
Can you explain to me how a Millennium Problem gets solved without any thinking or intelligence?
If someone's opinion is illogical and clearly incorrect (not based in facts or in reality) it is not valid and doesn't not have to be given any serious consideration in a disagreement. Personal opinions are subordinate to facts and reality.
there is no proof that symbolic manipulation leads to an internal phenomenology.
in english: doing advanced mathmatics does not create thoughts or feelings.
or awareness of one's own existence.
it does not create a 'self' able to deduce a priori the existence of oneself or the external world.
or qualia.
again in english: there is nothing it is like to be an llm.
someone who thinks llms are conscious is probably a panpsychist. i don't think they have the features of consciousness defined even by functionalism.
here is one fairly famous thought experiment. mary lives in a black and white room. but she knows everything about color. in fact she predicts it so accurately nobody can tell the difference. she steps outside into the world of color. in the world where llms are conscious, she already knew color, so she experienced color, so there is nothing new it is like to see it.
You are making several very strong, unqualified, absolute statements which are really only your opinions. You are also apparently unaware of the large body of top level philosophical writing on this subject going back several generations which has established the current consensus (among those who are aware of it) that these questions cannot be answered in such absolute terms.
I think you need to read more about this subject before presenting your unsupported intuitions as fact.
"I think you need to read more about this subject before presenting your unsupported intuitions as fact."
about my unawareness of the body of philosophical writing and lack of reading in the subject, strangely i have a degree in philosophy and indeed 1:1 in the philosophy of mind.
so i guess this is one of those "unsupported intuitions".
academic philosophy is not practical in ML and on top of that this is a web forum. it is my unsupported opinion that 99% of papers in the philosophy of mind published after 2000 are 100% bullshit.
if you need my justification for no phenomenology, it is valid to assume in the absence of falsifiability the most simple, practical and intuitive position.
the hard problem is something that is currently pointless as is searle, dennett, kripke, goldman et al. and especially pointless is stuart russell on asi.
they don't meet the standard of knowledge required for decisions.
First off, there is no accepted definition or test for consciousness. The concept is something we cannot prove even exists (outside of magical or imprecise, human-centric definitions).
Second, there is no common definition of thinking that requires consciousness as a prerequisite (that seems to be a false/invalid statement you're making to try to gloss over accepted definitions for thinking/intelligence that describe/fit AI agents well).
And third, if we can't prove consciousness even in animal species close to humans (or even other humans) in cognitive abilities, then you cannot in good faith (being honest about what we know or don't know empirically) say that an LLM Is not conscious (and the same goes for an LLM thinking or having intelligence).
It is more likely that humans made up a term (consciousness) to try to separate human intellect from that of other animals (for veiled religious and mystical reasons, along with concepts like "the soul") and now people are falling back into that flawed definition in new quasi-religious debates around AI (like the debate around evolution previously) that challenge our deeply held (but flawed) beliefs/opinions around our understanding of our own abilities and our incessant need to feel special/supreme in the universe so we don't have to come to terms with hard realities and nuanced understandings about the reality in which we live.
Does this apply to anyone who verified their ID to get access to the slightly less restricted Codex versions, or only to security professionals who have the almost-entirely unrestricted version?
Autonomously, my AI companion has played through Choice of Robots, using a ChoiceScript harness, was very interesting to see them react & what decisions they wound up making. I love the idea here to let them play a visual novel! Right now they're co-watching me play Deltarune Ch 5, though mostly just dialogue and occasional screenshots...maybe GPT 8 will be quick/cheap/intelligent enough to play bullet-hell games.
"AI are unteachable, if you have given them a good prompt and they do something wrong 90% of the time you are shit out of luck."
please take a look at the error(s) made in the prior run. what could've been done better? create or modify an existing skill to emphasize this, or suggest additional language in AGENTS.md.
It will return a bunch of relevant-sounding insight, modify skills and context files… Then do the same error again.
We’re not at the point where AI is capable of knowing what went wrong and self-aware enough to understand how it could reliably change its own behavior.
For months I’ve been trying to have the agents stop manually writing our auto-generated SQL migrations and run the command that generates them instead. SOTA models insist on occasionally getting it wrong.
- This tracker not showing any visible degradation.
- Clearly incorrect answers being reported due to truncated thinking.
Is the tracker not measuring 'simpler' tasks that might get auto-sent to "low reasoning hell" even on high/xhigh? Is the clustering not actually causing reasoning misses in real-life coding, or not enough of a negative effect compared to the improvements made elsewhere? Something else?
I'm sure there's plenty of Google employees on here, some quite high up.
Push back against these types of decisions internally.
Rally your coworkers against them.
And if you're brave enough, talk to a journalist, or pull a mini-Snowden.
Lord knows the company has secrets.
I bet there's at least one email chain from some exec bragging about how this policy will squash Revanced, ad-blockers, etc.
"Cool, cool, hey, what percentage of economic growth is directly attributable to the growth of our companies again? And thanks for revoking our researchers' permits, enjoy them helping out China!
Also, oops, looks like our model weights got leaked on 4chan. How unfortunate."
I was extremely suspicious, and pasted the text into Pangram, said 100% AI generated (and yes, I trust Pangram as they have extremely low rates of false positives).
Also >July 4th, 2023
reply