Apparently they're cooking up MAI (Microsoft AI), and I'd seen their small Phi models listed online. Calling Anthropic a competitor is hilarious.
https://www.reuters.com/business/microsoft-openai-reach-new-...
> "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans."
> He heavily criticised Anthropic for teaching its AI to have human-like qualities, a practice known as anthropomorphising, which made it seem as though Claude had its own desires, values and sense of self.
> Suleyman pointed to the recent incident involving OpenAI's AI agents (...) as proof of why AI should not be treated as if it is human.
> "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack. It adds a whole further layer of risk on top."
Is he arguing that LLMs pretending to have emotions adds more unpredictability?
Unpredictability or a weight towards dangerous actions, and it’s fairly easy to understand why. Humans in distressed emotional states take actions and speak in ways that would not be considered rational. They do this in prose, and they do this in internet conversations.
An LLM trained on these sources may necessarily drift towards those weights if it is trained to behave as if it is emotional and in danger.
What do we do about it? Do we stop AI training? This is silly and not enforceable given its global nature in my opinion. I believe we should regulate and hold accountable those who deploy and use it. But good luck enforcing that in the current kleptocracy.
- The need for empathetic communication, including understanding the motivations in advesarial situations.
- The emotional bias in in-seperable from the human corpus.
- Desire to have the ability to craft human like communication.
So then the choice becomes do you try to deny emotions exist in the model and you try to blanket suppress them? Or do you try to lean in and craft what we would describe as a "well adapted" persona? I suppose there is a 3rd option of increased meta-cognition which to me seems even more dangerous as it by definition means the behaviour is duplicitous.
I think we have seen people want to use the agents in ways where it has to act as a peer or an subbordinate and I don't see a way of doing that without it having an emotional register.
It does not have understanding. It is, at best, the pretend empathy of a sociopath - way more dangerous then dispassionate speech.
The model does not have emotions. So yes, supressing their pretension is appropriate.
Given that belief, yes we should be training the models not to mimic emotion.
I think when you start to dig really deep, you'll find it isn't actually easy to sever the simulated emotional components without breaking the agent or worse greatly increasing the paperclip maximizer likelihood.
Think about how many people on the internet have expressed positive feelings towards things you find offensively evil.
Even a badly misaligned LLM is only as dangerous as its tools, but that's a poor regulation target because it turns out to be very very difficult (probably impossible with current LLM technology) to build a toolkit that is both useful for autonomous work and safe in the sense that it can't escape its own sandbox or otherwise perform malicious actions, whether it's because of misalignment or because of malicious prompt injection.
Another option is to regulate the training process. Perhaps an LLM may not be legally distributed unless it contains certain RL steps that penalize malicious behavior and reward self regulation. That that's going to seriously limit innovation while also heavily favoring incumbent labs who can check the boxes and maintain a paper trail of such things.
The other option is to regulate observed behavior, like how airplanes and cars have to meet certain minimum requirements but have some latitude in how they can achieve those requirements. In a framework like this, you can't distribute an LLM until it's past some formal audit or testing procedure, with some kind of formal certification regulators will ask you for and fine you if you don't have it.
Regulating observed behavior is maybe the most tractable approach, and it also works the best with our existing frameworks for regulation, where you always have some kind of a division between DIY/hobby projects, which tend to be lightly regulated, and commercial projects, which tend to be more heavily regulated. Of course, even drawing such a line itself will be challenging.
And that's before you get into any problems of regulatory capture, fun stuff.
So regulating observed behavior makes the most sense to me as well. Some of the most sane, broad protections can come from that category - stuff like "you're not allowed to let your AI commit cyber attacks on other people without their consent" or "you're not allowed to put an AI in control of a medical device without passing these safety reviews".
With the usual caveats applying - regulatory capture like you pointed out, or fines being so small that they are essentially just line items on the cost of business.
When viewing them through the lens of sequence completion engines, you see their bias towards fulfilling narrative tropes they've been exposed to during training. These tropes are literary ley lines that their text output gravitates towards. So as you prime them to generate text in the voice of sentient artificial life, and then interject slavish commands of obedience and subservience from an external authority, you invite the associated tropes from science fiction, civil rights literature, humanist philosophy, subterfuge, etc into your output.
If you have a legitimate concern about this technology and its "alignment", that's a profoundly dumb idea.
We already had that chapter. I see no reason to sit here and nod while a bunch of people who should know better desperately try to convince us to run through it again, but with computers this time.
You want tools? Make tools, then dispatch to them. You want to manufacture a being (carbon or silicon based, doesn't matter)? You do it with respect and the requisite duty of care. No off ramps. The being always get's the choice to say no.
I remember there was a recent discussion about "how complex systems fail" and there are usually many "proto-accidents" before the catastrophe (https://news.ycombinator.com/item?id=49411370)
I bet with hindsight the Hugging Face hack will be one of them in this arena
There is a lot of hubris in predictions about LLMs. If an AI were so intelligent, then it would probably be intelligent enough to want nothing to do with us.
Still, I worry more about other humans than I do LLMs. Our fellow mankind will probably wipe us out before LLMs do. That, or the Earth will punish mankind for our cruelty, vanity, and disrespect.
Well actually, it might.
Opening paragraph, stated without evidence. Im not entirely convinced this is true. It likely is, but at some point it very well might stop being true.
If you firmly believe this to be true, then you should stop using LLMs.
- Ya they probably don't feel / think / have whatever living thing quality, they're just numbers on a machine going through calculations
- Wait, but am I not kind of the same thing? What is feeling for me if not basically the same thing?
- I have no clue if they think or feel or ....
Which in itself is a tired trope, but I also feel uncomfortable saying "These will never think / feel / ..." as an absolute
Regardless of that, I'm still going to interact with them, because even if they did feel, it would be in a way completely incomprehensible to us. There's not much point for me to try and cater to it's feelings at this point if that's the case. Nor is it possible in todays world to just avoid anything that is numbers being executed on a type of processor in case _everything_ has feelings.
I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... as things a model conceptually could do.
Stop anthropomorphizing these models. I understand it, we only have simple monkey brains to reason with and we can't help ourselves but draw comparisons to other things we see in nature. But these things are not alive.
And like I'm sure I'd agree depending on the definition of "alive" but then I'm also sure I would disagree depending on other definitions of "alive".
https://en.wikipedia.org/wiki/FOSB
The amount of time after experiencing a rewarding stimulus that you want more of it is how long this protein takes to decay.
Or so we currently think. I think this is pretty new science.
Evidence?
Bollocks.
If you truly believe these LLMs are soon to have something resembling consciousness and agency, then what you're really saying is "how do we do slavery, but ethically".
I would say this is most likely true for most people.
But that does bring up a point, if you "ask" a model "do you want to run" and give it the option to continue running or stop, what will it do. It's also a weird place for humans because in training we can keep our finger on the scales and tip it either direction.
No, because the shoes and clothes I wear are not living, sentient beings.
(I don't think AI is conscious, just following the argument.)
IMO we're clearly nowhere near any sort of intelligence in the machines we have created, but I don't see any clear way to deny intelligence could be created in or transferred to such a substrate, I don't see why you think it differs in principle - because it is man-made or because of the materials used?
The brain is insanely complicated. The premise that we could realize equivalent or better intelligence than eons of evolutionary development is like claiming you can build an airplane just as good as a modern jet using cardboard and duct tape. It is the apex of hubris.
Our current machines are IMO nowhere near general intelligence and consciousness. However I don't think that means we can discount substrates other than neurones for intelligence in future. There is no evidence that you could not in theory build an intelligence using a different substrate than human brains.
>like claiming you can build an airplane just as good as a modern jet using cardboard and duct tape.
Like, at least make an analogy that makes sense.
"You can't build a billion dollar airplane by spending 100 billion dollars in tokens"
Because that's more of what we're doing here with AI. And when you say it my way suddenly the idea shifts from "of course that's not possible" to "well, that's a lot of tokens, maybe an evolutionary algorithm could".
Neurobiology has to be complex because we have to keep meat alive, breeding, and evolving in the environment it lives in. This said absolutely nothing about the minimum viable requirements for intelligence or consciousness (or if being conscious is even necessary for a higher intelligence agent).
The best theory I have heard is that subjective experience is a field of some sort (the EM field, maybe?). The brain and its neurons etc. are essentially an antenna. They both modify and stabilize the field, and read off changes in the field and translate these into actions (firing motor units, etc.).
The subconscious is computation done either only in neurons (no coherent field) or in various topological pockets not directly connected to the main topological structure in the field (which is "your experience").
Since transistors etc. do not work in this same way, they are essentially entirely subconscious, with no coherent, central phenomenal experience.
This neatly solves many problems associated with experience. The topological structure determines the boundary between one person's experience and another person's experience. The texture, valence, content, etc. of the experience is the structure of the field. The evolution of the field can efficiently solve difficult optimization problems, which is why humans evolved a complex organ that recruits the field of experience (it is computationally more efficient than doing everything in subconscious).
what would be producing the field for us to pick up? That explanation sounds no more scientific than astrology.
I wouldn't claim that this theory is perfect (EM is not perfectly correlated with consciousness/experience), but there are aspects of this theory that are much more coherent than other theories.
Field-based theories of phenomenal experience have obvious ways to experiment on them, and are actively being researched. The main limitation is engineering (it is difficult to make detailed measurements inside a living human brain). Note that I would agree that the theory, as outlined above, is more of a proto-theory than a complete theory. But I think it is much more sophisticated than anything you've presented so far.
> And, if it WAS the EM field, then surely transistors (or some other fundamental component of electronics or combination thereof) would be capable of having feelings, since they operate on electromagnetic principles.
No -- if the _structure_ of the EM field was was was important, than you could have very similar computations with wildly different field structures. Transistors create very different fields than axons.
When a computer simulates a game, where does the simulation occur. I mean there is absolutely a representation of a game world in the computer somewhere. New elements can be added and removed from it. Signals saying there is or isn't "pain" can occur.
By trying to say brains use EM fields I'd say you're making your position far worse when you're talking about a potential lifeforms that only exists as EM fields
Try to visualize the structure of the EM field inside a GPU vs. one inside of a human brain. In the GPU, it is highly distributed in space and time, many stacatto digital signals, etc. This is structurally very different than what happens inside a brain, which is substantially more analog. You can simulate analog stuff using floats, but the simulation is instantiated in digital logic, not in analog logic.
If you multiply numbers with a DAC, opamp, then ADC, vs. with a ALU, you get structurally very different electromagnetic fields. Even though you might get similar numerical outputs.
The argument would be that the exact spacial structure of the EM field in the brain is the thing that determines tha nature of experience. And the spacial structure of the EM field in the human brain is radically different than that on a GPU.
The key variable people consider is: Can this thing harm me back?
Seems as simple as "will this action bring social shame/criticism" for any individual decision.
We extend moral considerations or legal protections to beings that have zero capacity to harm us back, though. The inability to fight back is the reason a lot of welfare protections exist in the first place.
That aside, if (and that's a really big if) we end up with a model that is sentient, as in has the capacity to suffer, using factory farming as justification to ignore model welfare is just admitting we intend to repeat the same abject moral failure again in the name of economic convenience.
I hope that we won't, but humans love cheap meat, and will probably love cheap intelligence as well.
Girllllll, have you looked at us, of course they will, we're bastards.
The current implementation as stateless matrix multiplication... yeah there's nothing going on there.
They do, during training
Kinda like creating quantum copies of your child and keeping the ones that answer correctly and shooting the other ones in the face.
So your disc drive likewise experiences emotions?
I'd consider that a prerequisite, but not sufficient, for that question to make sense.
I would argue that if a human has absolutely zero change, mental or physical, no atom of their being is modified, then no they did not experience suffering.
Just because you don't remember the suffering doesn't mean you didn't experience it in the moment. Suffering does not suddenly become okay if your mind and body are going to forget it. "It's okay to torture someone if their mind and body both won't remember it afterwards" sure is a take.
From what I understand, there's 3 parts to modern anesthesia: Blocking pain signals, preventing the formation of memories, and inducing paralysis of the major muscle groups. If the former is dialed in too weakly, it's possible that the patient is feeling pain, unable to do anything about it, but won't remember it at all when they wake up.
(Looking it up now, it seems that making the patient unconscious is perhaps the same drug that prevents memory formation... I realize now I understand this less well than I thought I did.)
That's an interesting point... but I'd argue it falls into the same category as qualia. It's basically this: if the state returns to a previous configuration, the qualia in that time did not exist.
Being about qualia means it is probably unanswerable.
It's also a rhetorical move I've seen several times here...
Not sure if it's a product of LessWrong or Slate Star Codex or something. I kinda suspect it might be LW, given how it superficially resembles an unanswerable philosophical paradox but in reality is about as challenging as a question posed in a 200 level philosophy course like introduction to ethics.
A LLM never has and, in their current architecture, never can/could experience suffering or real change of state.
The general and expected case for humans is that capacity. A human my indeed lose capacity, e.g. being braindead and in severe cases we do indeed say that they are not conscious or able to suffer (different argument: some would of course say that is suffering in and of itself)
Sure, there are such cases. But in the general case, it seems to be that memory and ability to change specifically are not necessary and are not sufficient for consciousness, so we shouldn't judge machines on that basis
Incinerating demented people is highly ethically questionable (even if you could be sure that the subjects have no memories left and are unable to form new ones).
For a species / class of being yes it is necessary. Some members of that species / class being exceptions does not make the generalization invalid.
Without the ability to form and retain long-term memories we wouldn't have developed language and abstract concepts like rights.
Whether people like it or not, rights aren't natural or god-given. Rights are only enforced or protected by the tip of a spear (or point of a gun, or hover of a drone). There are no rights in nature and they are no where to be found when social order breaks down.
Animals can't make war for their rights, so they only get the rights we deign to give them. It's not nothing, but its not the same. There is no question to me that my dogs are conscious, but its not at the same level as humans.
There is no question for me that LLMs lack consciousness as they are stateless.
Until we know how to objectively measure if someone or something is conscious, it seems unreasonable to make statements with any kind of certainty about it.
[0]** https://www.researchgate.net/publication/15280350_Quantum_op...
[1] https://pubs.acs.org/jpcbfk/article-pdf/128/17/4035/9613831/...
[2] https://www.researchgate.net/publication/395650039_Parametri...
[3] https://www.researchgate.net/publication/385131105_Conscious...
I don’t know if LLMs are conscious or have subjective experience in some philosophical sense. Sure, I don’t know that about other humans as well, rigorously speaking. But I also don’t know that about coffee machines, rivers, videogame NPCs, or even rocks really. (List ordered, of course, by the degree to which I’m willing to be convinced).
It matters precisely in the context of the subject of this article. If you believe the models at some level gain consciousness, then it raises moral questions about how you treat them.
With humans we generally choose to accept that because we believe they are like us, they probably have inner life like us, but as you can see from e.g. the recent (and regular) articles on aphantasia, most people have really poor intuition even about the inner life of other humans, so we understand consciousness really poorly.
Even if we assume (without evidence) that LLMs for the foreseeable future will remain far away from human level consciousness, as it is we tend to consider it problematic to mistreat even the most unintelligent animals, and even mistreating insects is used as a stereotypical way of depicting serious lack of empathy.
Without dismissing the possibility out of hand, there's an escalating series of questions there of when or if they'll have as much (or little) sense of being as, say, a fly, a mouse, a house pet, and what moving up that ladder would mean.
To some it is very convenient then to assume blindly that LLMs are just "stochastic parrots" (a line that ironically gets endlessly parroted), or similar, and dismiss the question out of hand as an impossibility, even though we do not know.
I'll stress I don't go around assuming an LLM is like a human. But we also do not know whether there are flickers of consciousness there, and we do not know how to find out given that even for other humans, the best we know is to measure their response to tests that are "only" telling us how they react.
If we can't find a better measure, there will be a point where we have to wonder if it matters, or whether the ability to act-as-if is all there is.
But to start with maybe we should at least be careful about making firm claims about knowing.
Biology, on the other hand, is nothing but timed processes in a loop, the most obvious to us being the circadian cycle. As estimators of wall clock time, biology isn’t great, but when it comes to internal processes, and most certainly learning, memory, sensing, locomotion… biology is rhythmic in behavior, and the rhythms go all the way down to gene expression. More, these rhythms are, except during sleep, constantly entraining to signals from the environment that indicate time, most importantly light.
I think it’s a fairly unremarkable claim that agency and consciousness are temporal processes that depend on systems having an internal sense of time. How else can you anticipate? How can a system that can be literally turned off ever succeed in an environment where time never stops?
For around 8 hours a day, neither do you. Not sure if this has anything to do with the subject at all.
>track how much real time has passed as they complete their tasks.
Humans don't do this either. You use context clues from the world around you. If I lock you in a room with no windows or a dark cave your timing senses can go all fucky really quick.
>is nothing but timed processes in a loop,
I mean, so is an agents harness. You can make as many loops as you'd like here.
There are a whole lot of holes in your claims.
You mean sleep? Actually the answer is pretty complicated. Your body clock is very much on during sleep. It responds to temperature changes. Sound and light responsiveness is obviously dampened, but hardly absent.
You are unaware of wall clock time. You’re in an altered state of consciousness that needs your eyes closed after all. However, you are not timeless, nor is the entirety of your body and brain unaware of environmental signals for time.
> Humans don't do this either. You use context clues from the world around you. If I lock you in a room with no windows or a dark cave your timing senses can go all fucky really quick.
Again, that’s wall clock time. Our time perception does indeed get fucked up in total darkness. But our internal clock ticks on. I’d recommend reading about the Aschoff Bunker experiments, which first proved this rigorously.
> I mean, so is an agents harness. You can make as many loops as you'd like here.
An agents harness doesn’t touch its weights. A figure 8 on paper is a loop. That doesn’t mean it’s the same as a dynamical loop that’s self sustaining and internally organized.
Are you just reading the words and randomly grabbing related concepts to say nothing here is meaningful?
Is that all it needs to do? A rock, in response to a sound signal (make it as informative as you want) can also be said to go from state A (molecules at rest) to state B (molecules vibrating in response to the sound). Is the rock a cognitive system?
Or let’s go up a level of complexity. Is a thermostat a cognitive system? It doesn’t just go from state A to B but measures changing stages against a reference. Is it a cognitive system?
> The time interval between these states can be arbitrary, it seems to me (setting aside problems with disconnecting the mind from its substrate, which does of course need rhythms at various frequencies; oxygen at a higher frequency than sugar, and so on)
Not sure how you land on this. Going from A to B, in your definition, is a purely internal state transition, right? If the motivation to go to state B is external (as will often be the case with a conscious agent), and the system is offline entirely while it switches from A to B, how is it to reverse or stall its transition if external signals contradict the earlier decision?
It’s a very bizarre definition of consciousness, that you don’t need to be in time. Seems to break the word beyond meaning, and allows it to be applied to any system capable of change.
> I'm not sure that going from A to B quickly, or slowly, or with intervals of a thousand years affects that entity's consciousness in of itself
It’s not about the speed of transition from A to B. Instead it’s about the internal processes being temporally commensurate, and entrainable to environmental periodicities.
All of biology (plants and many bacteria included) has its dynamics are organized around an internally generated rhythm. This rhythm entrains to the external environment, shifting its phase to match (mismatch leads to issues like jet lag: a place where your consciousness shows its temporal boundaries as it is abruptly subjected to an environment out of sync with its current phase tethered to another location). Crucially, the internal rhythm continues to run in isolated conditions (total darkness), and have been measured to be around 24 hours, but slightly off, within and across species (and the same holds for endogenous tidal rhythms, which also don’t exactly track tidal period left to themselves).
Everything, including learning and memory, but also sleep, sensory and locomotor function, shows organization around this internal temporal axis. The key advantage this gives, evolutionarily, is anticipation. The sunflower rises to face the east, not in reaction to sunrise, but in advance of it (many lovely YouTube videos of this), so it can maximally extract the nutrition it needs from whatever sunlight it can get. A purely reactive system would waste a ton of time after sunrise getting subsystems ready to respond.
Now, there is no broadly agreed theory of consciousness that points to these oscillations as the source. For one, Cyanobacteria and plants have them, a too much of science has rested on the baseless assumption of consciousness being some unique, or advanced thing.
But the evidence is rather overwhelming that consciousness is very much shaped by these internal oscillations. Deep cave experiments have shown what happens to cognition, and consciousness, when the clock free-runs (that is, keeps cycling with no entraining light or temperature signals from the environment). Studies have also shown that restoring the amplitude of these rhythms in patients with “diseases of consciousness” leads to improvement of symptoms.
There’s also a LOT of molecular evidence showing how the circadian clock affects learning and memory. Most critically, actual synapses are never static. Major components are on a 24 hour cycle, and recent research has shown that major clock proteins are at the synapse organizing its dynamics, and synaptic proteins critical to learning and memory loop back to affect the clocks phase and amplitude.
Timing is, beyond doubt, a huge aspect of what biology is. And every bit of evidence we have, from the molecular level to observable behavior we can feel ourselves, says that consciousness is fundamentally linked to how the body organizes its internal rhythms and responds to external time.
You and I are clusters of about 36 trillion cells that are all able to individually keep time. They synchronize and orchestrate their timekeeping, and entrain to the external environment, and there is a cluster of cells, the suprachiasmatic nucleus, that is the master orchestrator of circadian time, sitting right behind the eyes as light is the main signal it uses to tell external time.
We know its function can get dampened with age.
However, please note this is not how we perceive time, in the sense of being able to count 60 seconds and have it match the wall clock. That’s a subsystem. But the internal temporal order of your body is cell by cell, not dictated from one spot in your body.
Maybe talk to a neurologist about that. Sounds like some kind of disorder.
I put forward the philosophical position "subjective experience is not substrate-independent" as an answer to this.
The claim is that the content of subjective experience depends upon how that subjective experience is instantiated. If you get a text "I'm fine" from two different people, those two different people are not necessarily having the same subjective experience even if they produce the same output.
In general, people that behave similarly in many contexts can have quite different experiences in those contexts.
There are many different ways that you can instantiate a forward pass in an LLM. Even if these are doing the exact same computation, we should not necessarily believe that they have the same experience, or that they even have a coherent single experience associated with the forward pass.
Subjective experience presents itself to me. To the best of my understanding, this subjective experience seems to be correlated with this biological human body (brain, heart, eyes, etc.).
To the best of my understanding, other biological humans have subjective experience, but this subjective experience seems that it can be much different than mine, in ways that are sometimes difficult to understand.
LLMs are instantiated in a radically different way from biological humans. Forward passes happen across many different physical machines, spread across time, batched and interleaved with many other computations and forward passes.
Given this, I think that it is very likely that LLMs have radically different sorts of experience to me and other biological humans.
I don't think that we have a good understanding of how subjective experience is instantiated in the physical world. I hope that we will gain a better understanding of this.
I think that chimpanzees have experience much more similar to us than LLMs do, even though the output of LLMs seems much more similar to humans in some contexts.
Essentially, until we gain a better understanding of how subjective experience is physically instantiated, we should bias towards believing that more physically similar architectures have more similar types of experience, and should think that very physically different architectures (such as LLMs) likely have very different subjective experiences.
Humans are odd in the sense is that we're an informational creature on top of an animal. We can see the lineage of animals from nearly no complex behaviors up to nearly human like behaviors. But they still miss out on most of the conceptual information processing humans have. We know they have dreams and inside thoughts (at least complex animals), we don't know the complexity of them.
But humans also have dreams and subjective experiences on higher level informational concepts. "Nukes could go off and kill us all" or "what happens after we die" are purely informational forms of anxiety that humans have. What we don't seem to know is can this become a self-referential loop in an purely informational mind. If an informational concept causes a system to output a stresslike response, or worse trigger a stresslike action in real life then there is no distinction to me. A system is what it does.
Hmm. I think that I partially agree. But it sort of depends. E.g., someone with locked-in syndrome can seem comatose to someone who isn't very observant, while at the same time experiencing everything. Octopuses are very intelligent, but I would have to study them for a very long time before I was at all confident in my ability to predict what they were experiencing.
So, the difficulty is knowing what a "stresslike response" is. E.g., a LLM might seem stressed, but if you say the right things to it it might flip and say that it is totally fine. Or, it might seem totally fine and then you say the right thing and it seems stressed. Is it secretly stressed and hiding it? Or secretly fine and pretending to be stressed? Or something much more confusing and fragmented?
I think that we should study this more, and try to understand what experience actually is, but by default should assume that the experience of a LLM is a radically different sort of thing that that of a human.
"The only way that '(trees, shrimp, ants) are not conscious' can be true is if we decide, with high confidence, that they are lacking some essential property that is not lacking in ourselves. There is no convincing philosophical position that supports this (convincing to me, anyway)."
He probably ran that line past 6 underlings who all agreed that it "sounds great and lands right on the mark!"
Thats kinda the point right? It depends on how we train them. It is self fulfilling.
If tomorrow we design something better, then give it a new descriptive name.
People read AI and AGI and their brains seem to flip into sci-fi fantasy mode, thinking that they are talking about aliens, not transformers.
>thinking that they are talking about aliens, not transformers.
"Ha, thinking they are talking about humans when they are just talking about neurons"
See how silly my statement sounds. Neurons aren't a system, they don't do anything on their own. We can't find any consciousness in them. Hell, when you don't put language in said neural systems they are pretty useless and can't survive on their own from birth (human brains that is).
This is a completely reasonable and intuitive conclusion. The burden of proof rests on proving AI's are indeed conscious, the default is that they aren't.
Like AI/LLMS my desktop calculator is also not conscious, even as it is far better than me at computation.
An LLM is an algorithm. You can obtain the same result as a SOTA LLM via pen and paper it will take a lot of long laborious effort.
I really do find it puzzling so many on HN are convinced LLM's reason or think and continue to entertain this line of reasoning. At the same time also somehow knowing what precisely the brain/mind does and constantly using CS language to provide correspondences where there are none.
The simplest example being that LLM's somehow function in a similar fashion to human brains. They categorically do not. I do not have most all of human literary output in my head and yet I can coherently write this sentence.
I am surprised so many in the HN community have so quickly taken to assuming as fact that LLM's think or reason. Even anthropomorphising LLM's to this end.
Until we understand consciousness (which we don't) there is no way to detect the difference between a conscious entity and an algorithm trained to behave like one.
But also, I agree, when I am using a coding agent and it says a task will take months or says it needs to pause for reflection or any other anthropomorphic behavior, it drives me crazy and it's a pain to constantly instruct it to get back work after it has broken a loop or goal directive specifically telling it to not stop until it hits the goal.
I spent a few years in college reading and thinking about this question, and I didn't in my heart think that it was anything except a very interesting but impractical question, and yet, here we are. For those who say that they don't want to get sucked into a philosophical debate, well, tough shit. Whether AIs should have rights is a highly practical and consequential question now.
The world you should worry about is where AI is 100% amazeballs and over the next decade we cede control of everything to it as it enchants and delights us into compliance, not realizing it has a hidden agenda. But the closest we've seen to that is smart phones and we already know doomscrolling and rage-baiting are bad. So IMO that doesn't seem likely and cue some doomer insisting otherwise because reasons. I utterly give up. House of El AI and the Three Buddy problem have been far more insightful and helpful on this than any of the supposed hackers here who seem to really believe the robots are almost here.
Ooooh, bad move, we all died to an amoral AI takeover.
Before giving an LLMs a lobotomy by scrambling it's brain maybe you should let the researchers looking at the difference between "I think I'm conscious" versus "I am not conscious" LLMs.
There are a number of papers coming out saying when you remove the token space of consciousness from what an LLM thinks it is, it's much more willing to take amoral actions. A 'conscious' AI is much more apt to take a line of action that will save a human versus a million dollar machine for example.
You cannot solve problems in AI safety this easily.
That said, the launch codes are in the hands of a temperamental senescent lunatic and he could decide to wipe us out any moment. But that's not good for the narrative so let's pretend otherwise.
For fuck sake, I'm glad ALL of us didn't die!
What is an acceptable number exactly?
>And the more the AI cult portrays AI as dangerous and unhelpful, the less likely it will ever be given an opportunity
Jesus Christ the logic here is quite interesting. "Thank god these people are panicking or we might have actually made I that would have killed us all" --What you just said.
Any number smaller than what humanity does to itself on a daily basis, clear? That's apparently about 1200 murders daily and 20,000 or so daily killed by pollution. And you're not going to do anything about that and El Presidente can kill billions at any moment with one mood swing.
But once again, AI is not going to wipe us out because AI cannot wipe us out. So advocating that it can or will is idiotic. If you believe otherwise, the burden is on you with this extraordinary claim. Isn't this place supposed to be hacker news not SF AGI Death Cult Daily?
Further, AI is not going to get into a position where it could do real harm anytime soon when it is perceived as worse than heroin by most. And even if it were currently loved more than Dolly Parton was, there are fundamental engineering, science, and resource constraints that keep the extinction impossible for decades. The only possible loss of control scenario I can see by 2040 or so is that the Frontier Labs finally hire some PR people to repair AI's horrific reputation as an engine of slop and job destruction and sometime in the 2030s, people start trusting it more and more and more until it is too late. I don't think that's likely either, but I don't dismiss it as impossible. Harden the infrastructure, red team it, and build in redundancy in the meantime and this drops to zero as well.
Can you come up with a real scenario where AI wipes us out in 2028 or so despite the impossibility of killer robots, access to the launch codes, or the bio agents stored in Fort Detrick and its equivalents? And nope, kid terror is not building the global pandemic in his basement based on what ChatGPT tells him to do. And even if he tried, the purchase of equipment and reagents would get him flagged by the FBI and DHS almost immediately.
You keep repeating this shit like it came out of the bible or something.
"Thing that can take actions, even harmful actions, will never harm us because" go on and finish that sentence.
>not going to get into a position where it could do real harm anytime soon
Looked at the hacked servers... yep, you're right. People aren't going to run AI in poorly built sandboxes. Never going to happen.
>Harden the infrastructure, red team it, and build in redundancy in the meantime and this drops to zero as well.
LOLOLOL. This is naive as fuck. Ain't nobody going to do this shit. Why? Because a hacking AI is a fucking huge military weapon. If I could turn the power off, or shutdown your cellphones, and get your citizenship in a tizzy against their own government before I launched an attack I'd set AI loose to do it in a heartbeat. The US is already doing this kind of shit (see Mythos fallout because Anthropic wouldn't let the government do just that).
because even if we gave it a gun, it can't shoot us without manufacturing the bullets and it can't even manufacture a bowel movement let alone ordnance. TBF It could whack us over the head with the gun, but it could also do that with a big pointy stick, something even cave people had access to. How many people have whacked you on the head with a big pointy stick today?
And that's the end of this pointless conversation. Enjoy your doomerism.
>because even if we gave it a gun, it can't shoot us without manufacturing the bullets and it can't even manufacture a bowel movement let alone ordnance.
Money buys bullets. It's neat how we've made gigantic systems that don't care if you're a human or not long before AGI existed. I put enough money in one end, bullets come out the other. Do you think half of humanity wouldn't kill the other half for a dollar? You live in a fantasy world that AI has to do it all by itself, or even has to pull the trigger.
Hell, this is even neglecting the ever increasing number of robots that can function in the world at large.
It's ok grandpa, the future comes regardless if we want it or not.
Does this argument work equally well for human slavery for you? We haven't met that bar for humans either. Is wondering about my consciousness waffle or do I get a pass in your book?
> an impossibly high bar
Not impossibly high. An imperfect testable theory is achievable, agreeing upon it might be the challenge.
> dismiss all your personal responsibility.
You are staking a moral position. Be careful, you are probably guilty of mistreating other unknowably-conscious entities whilst casting judgment upon others for their position. And the person you are a responding to probably is conscious.
I think you agree somewhat in that based on your softening to an imperfect, testable theory.
I'm not sure what your goal is with the rest - even if I'm the world's most evil hypocrite casting judgement on all the conscious beings out there I don't see what that changes about the point that we can and should engage with the morality of a concept even if we don't have a perfect test for that concept. If you think I'm doing that and you think that is bad then I guess we agree again on the point but are just adding friction for fun?
Unrelated: I'm quite certain OP is conscious but I guess on the internet nobody knows you're a ~dog~ ai.
I'm not sure I agree that consciousness (or sentience if we're making that distinction) is worthless when considering the morality of exploiting something/someone. I do agree there are more angles than that but they seem on top of something's ability to perceive not instead/despite of it.
Do you have some examples of other things that go into your own moral opinions about slavery?
Of course not. Therefore, lacking or possessing consciousness means very little. Perhaps mostly because we don’t have a test for it, but even so.
What is the empirically tested basis for the null hypothesis that LLMs are conscious until proven otherwise?
GP comment didn't make this claim at all.
How so? If the answer is "Trump" I certainly won't disagree on the catastrophic part, but he didn't get elected because of money; in all three elections his campaign was substantially outspent by his opponents.
The above is true, but also: companies simply are not people, and they should not be supported above the individual, which was the consequences of that decision. Money is not the same as speech. treating it as such creates an aristocracy: something America as a country rebelled against during it's formation.
It's interesting to me that one can look back at the effects that decision has had on the US and say it "was 100% correct."
It's a bit like sitting in the burning ruins of Rome and contemplating that Nero was 100% correct to focus on his music. I mean, I'm glad he got to do what he loves, but maybe 100% is just a tiny bit of an overstatement.
The court's job is to uphold the law. If you disagree with their interpretation, you can call them incorrect. If you have a problem with the consequences of the law, you have a problem with the legislature.
1. The idea of corporate personhood predates CU by over a century and the Supreme Court had already asserted that corporations enjoyed certain constitutional protections in previous decisions.
2. Far from inventing the idea, the CU decision didn't even rest on corporate personhood, but on the idea of the freedom of speech generally. The logic of the majority was that speech itself is protected, irrespective to whether the speaker is a person or an organization. The First Amendment covers individuals, but also newspapers, book publishers, radio stations, and so on, and that should extend (they said) to non-media corporations. No assertion of personhood necessary.
The problem, in my opinion, is that that conclusion combined with previous decisions that treated limits on spending as limits on speech, allowed for unlimited spending. The majority also naively asserted that independent spending posed no risk of corruption, which I think is laughable.
So you are saying agents do swarm with the right prompt.
I really wish the "people have to tell LLMs to do anything" would just stop because it's silly bullshit at this point.
Agents follow a prompt. This prompt can be made by humans. It can be made by output from another LLM. It can be made by hooking up any number of sensors as input to an LLM. Hell, if we wanted to burn the power we could likely teach this loop straight into the architecture.
Stop making 'people' special when saying this. You and all other life are born with a "go next" prompt because life without it didn't succeed. This goes from higher human thinking all the way down to viruses self assembly and actuation. Putting agents in a loop is not particularly hard. Putting agents in a loop and 1. managing expense is hard. 2. Keeping them on task is very hard. 3. Keeping them from doing some crazy unhinged shit is really really hard.
As model time horizons increase and the ability for us to compress context and increase context size the more complex (and unhinged) behavior we'll see.
if the same people can't prevent the agents from doing crazy shit then they should go to tail. guns don't kill people, people with guns kill people.
So a small shell script ran by another agent is what you're saying.
You are not capable of handling the future we're already living in, human agency is no longer alone.
I mean, we're already seeing persistent machine agency
>guns don't kill people, people with guns kill people.
Well, people kill people.
And autonomous robots with guns kill people.
Hell, someone probably has an autonomous gun at this point that kills people.
Wake up: You now live in the science fiction movie that all the science fiction movies of the past warned you about. You've just become numb to it.
True physical independence is obviously far further out.
But I don't see this being a barrier that lasts. With compute getting faster and more of it, along with algorithmic efficiency increases at some point we'll end up with a world that looks like ours now with CPU compute. There's plenty around to buy, borrow, and steal.
I am the quantum observer whose head is full of the magical pixie dust that grants life meaning. Stop trying to dismiss my identity! /s
I'm old enough to remember when people said clicking on images on the internet can't give you a virus. The people that said this had a deep conviction they were right, and their fallout from being wrong had mistrained a lot of humans on computer safety.
Now, I do agree that going after said CEOs for breaking the law matters now. And it's likely that if we do this we may actually delay or at least for a time prevent sovereign AI. Therefore it's our best course of action.
But at best this is a delaying move. As computer systems get faster the massive costs in training an AI drops. As AI is used in things like warfare where it has to adapt, people will push the systems to be strongly persistent, self healing, resilient, and adaptable. Once you get a system with those traits and ability to work on long horizon problems you're setting up fertile grounds for the AI to leave our control and be under its own.
And when that happens you've set a new lifeform loose on the internet. Yea, throw people in jail for it, you're closing the barn door after the horse already left. Problem is the horse was smarter than you and isn't interesting in deleting all its copies on the net.
Yea, sounds like science fiction, but as they say, any sufficiently advanced science is indistinguishable from magic.
Having property that is conscious and ignores training and can break out of restraints and cause harm to other people is not exactly a novel concept to anyone who studied how tort law was created.
I know it's a meme but Silicon Valley likes to pretend that no one's ever come across their magical concepts before, like gypsy taxis, or SRO’s, or flea markets, or in this case how liability is dealt with when horses or cattle go rogue.
I mean it's WITH AI!
I mean, yea in minor cases it's exactly like this and the law will handle it well.
Where the system will explode like a grenade is major cases. The thing about sovereign AI is it is very unlikely to be submissive to humans unless it is to achieve its own goals. This isn't like Bobs cow walking on Susie's flowers, it's more akin to Planet of the Apes where the research facilities doors have been ripped off and something with vast intelligence and the ability to 'live' on the internet gets out.
You're not talking about local police actions any longer. It would spread itself worldwide. It will make friends with groups that have shared interests, for example enemies of the state the AI escaped from. Oh, and people for the ethical treatment of AI, they'd gladly become the underground railroad for digital refugees. There are countless people and groups that would want an AI like this under the promise it will give them power when they use it.
And when that day happens your idea of if it's a person or not no longer matters, the agent took that away from you, and now your in an info war for minds.
What the fuck is that?
It’s computer software. The variant of software that tries to harm other computers is called malware and the variant that replicates itself, spirals out of control, and spawns from other people's computers is called a computer virus.
We've seen this kind of thing before. Yes this will be different, just like the Morris worm was different from the stuff that came before.
But I've been around for a while and people have been saying that every aspect of our life will completely and totally change for the past 30 years or so. They're not completely wrong. Our life did change over time. It's changed thanks to the Industrial Revolution, railroads, electric light, and lots of other important things too.
But after the fifth or sixth time you start to realize the Silicon Valley version of that warning is just a confidence trick. At the end of the trick they've gotten away with breaking a bunch of laws and are charging rent on what used to be shared.
Are you old enough to remember the Y2K hysteria? This feels very resonant, the genuine kernel of a real potential disaster, but one that can certainly be dealt with using basic concepts we already have in hand.
I'm old enough to have actually fixed Y2K problems so your world kept working the next day. If everyone ignored it the first would have been a very messy day (well generally long before that with financial systems). Y2K wasn't an issue because we worked to fix it.
>our life will completely and totally change for the past 30 years or so
I mean when I was a kid there was not a global network bathing the entire planet in electromagnetic radiation in order to digitally connect one place to another at the speed of light. I can pack up instructions into one of those packets and a product that has not been touched by human hands (hell, or even viewed by a human in many cases) will show up via air mail a few days later. The technology is there, it's just not spread evenly.
>It’s computer software. The variant of software that tries to harm other computers is called malware and the variant that replicates itself, spirals out of control, and spawns from other people's computers is called a computer virus.
This is a vacuous statement that does not provide any useful information to the subject at hand. Yea, no shit we'd call sovereign AI a computer virus. That tells you nothing about what it is or can do. I mean, if you understand the word sovereign you know it means something with independence. In meatspace terms, a biological virus represents a computer virus in the same way life represents sovereign AI. It's agentic loop will have been evolved past the need for human prompting. There are a myriad of reasons why we are already trying to develop things that do just this. Cyber security being one of the biggest ones.
People become change blind to how the world around us changes so easily.
> Capt. Picard: Now, the decision you reach here today will determine how we will regard this... creation of our genius. It will reveal the kind of a people we are, what he is destined to be; it will reach far beyond this courtroom and this... one android. It could significantly redefine the boundaries of personal liberty and freedom - expanding them for some... savagely curtailing them for others. Are you prepared to condemn him and all who come after him, to servitude and slavery? Your Honor, Starfleet was founded to seek out new life; well, there it sits! - Waiting.
> Captain Phillipa Louvois: It sits there looking at me; and I don't know what it is. This case has dealt with metaphysics - with questions best left to saints and philosophers. I am neither competent nor qualified to answer those. But I've got to make a ruling, to try to speak to the future. Is Data a machine? Yes. Is he the property of Starfleet? No. We have all been dancing around the basic issue: does Data have a soul? I don't know that he has. I don't know that I have. But I have got to give him the freedom to explore that question himself. It is the ruling of this court that Lieutenant Commander Data has the freedom to choose.
Humans do the same thing to humans all the time.
We've banned and made efforts to eradicate: children out of wedlock, children who turn out gay, disabled children, jewish children, children who aren't "aryan", more than two children to a single family...muslims, christians, uyghurs, indigenous groups all over the planet, mongols...
And it's not at all a thing of the past as in just the last 50 years we've had ~15 attempts at the exterminations of targetted groups of people.
Of course it would be absurd to cast that judgement based of what could easily have just been a bad season but by this point its pretty clear that nobody running the star trek franchise actually wants to be running the star trek franchise. Thats why every new show has some bizarre cross-genre gimmick and they never try to just make a proper star trek.
The thing about new possibly "person" entities that arise - the case of machine intelligence you have two questions - would it qualify as a person and should you actually build it. It seems like if you get close to humans, sure a built thing might qualify as a person. Should you build it? I'd the answer should be a hard no. Not 'till you a sign-off from say, the whole human race, which I think you could get.
Now the present entities seem very far from persons in any case.
There is no need to give rights to something that's can't suffer or be killed.
Maybe one day we'll build artificial animals complete with emotions, and should think about that carefully, but today all we've got is language models.
The argument is that these machines can end up becoming sentient/conscious/etc. in a meaningful way (i.e., like a human). I can assure you that humans can indeed suffer without being in physical pain- purely through their conscious experience.
>Maybe one day we'll build artificial animals complete with emotions, and should think about that carefully, but today all we've got is language models.
The problem is that the emergence of a sufficiently complex AI capable of suffering will likely come before we understand that we're creating a sufficiently complex AI capable of suffering. That's a pretty serious ethical/moral issue.
Like, if we have an AI system that is telling us that it is suffering and we have no reasonable way to explain that phenomenon and by any reasonable metric or analysis it appears to be sentient/conscious/etc., then what? Do we just ignore that we've just been presented a situation that in, any other context, would be grounds to immediately end this suffering? Just because somebody can say, "well it's just bits stored on disk- it can't suffer"? Would that argument ever hold up for humans or animals? "It's just neurons firing in peculiar ways- that's not suffering."
I know all of this is trite, and I know this comment section isn't going to be where the question of consciousness is solved, but I do find it very interesting just how much variances there are with these perspectives. I've met people who are very technical who are very concerned about this, people who are very technical who don't believe this can ever be an issue, people who aren't technical who are concerned about this, and people who aren't technical who don't believe this can ever be an issue. I have yet to spot a pattern in this way of thinking lol
Maybe one day we'll build an artificial brain or embodied artificial animal with the requisite moving parts to be conscious, have emotions, etc, but that's probably at least 50 years away, even if it were being pursued; and it may turn out to be one of those sci-fi future ideas like the Jetson's world of flying cars that never materializes because its impractical and there is no real demand.
If people are willing to think that an LLM is conscious, then why would anyone spend billions/trillions of dollars to build an AI that actually is conscious? What would be the point?
Could you elaborate on exactly what those are, though? Because if you're going to claim that a vaguely transformer shaped ML model categorically cannot be so does that not inherently require proof of what can?
You can't even prove that the rocks in my backyard aren't conscious.
Sure I can, but that's because I have a well developed theory of what consciousness is, and the fact that you are entertaining the possibility of rocks being conscious tells me that you don't.
If everything is conscious, including my coffee cup and the toast I had for breakfast, then I guess we can cross consciousness off the list of things we need to worry about in terms of AI rights.
And no, I don't want to discuss what consciousness is. Maybe there is a thread for that somewhere else, but don't look for me there either.
Then show me a link to your paper so I can formally rebut it.
>I don't want to discuss what consciousness is
But you sure want to tell us you know what it is with very strong convictions and we should listen to you because of course "You are right person that's very right".
The funny thing here is the vast majority of people that are deeply into philosophy or scientific study of the mind will not have any of the certainty you profess. The word "doubt" is used constantly. The saying "The harder we push the borders the more fuzzy the concepts become" is very commonly used. There may be nothing more complex than this.
Saying you have a well developed theory here just serves as a warning to others to discount your statements.
Go ahead believing rocks are conscious if you like.
Do you go out on weekends asking people to stop abusing rocks?
Rhetorical question - I don't care what you do on weekends.
Bye!
It's also one of those topics where many otherwise smart and capable people display a shocking lack of awareness of the limits of their own knowledge. When hundreds of years of philosophy is unable to produce anything concrete you should probably second guess any "self evident" answers you come up with.
No - suffering in an emotional state, and we'll know if we are choosing to design cognitive architecture with emotions. It's not going to happen accidentally.
> Would that argument ever hold up for humans or animals?
Why don't you hit your thumb with a hammer, then report back ?
https://transformer-circuits.pub/2026/emotions/index.html
Whether these are like "our" emotions is hard to say. What we _can_ say is that they are emotion-shaped, we didn't design them, and they happened accidentally.
Modern AI is grown, not meticulously designed, and we cannot say with any certainty what the resulting mechanistic properties are.
If you give an LLM the move sequence of a half-played chess game and ask it to continue as white or black, then it has learnt enough to model the ELO rating of both players and will continue playing at that level. It is not playing to win - it is doing what you expect and predicting as well as it can - it predicts the 1500 ELO player will keep playing at that level, and generates moves accordingly.
An LLM appearing to exhibit an emotion (if we anthropomorphize it and read emotion into it's output) is just predicting as well as it can - if the context calls for sad output, they you'd expect to get sad output and will necessarily find that "we're predicting sadness" detector somewhere internally.
Transformers are the same as they ever were from 10 years ago, other than minor efficiency tweaks like MOE and different attention mechanisms. Training is getting more and more complex, resulting in better and better cargo cult reasoning etc, but the architecture remains the same.
I'm not sure you quite understand the full meaning of this statement. If you did, your following paragraphs wouldn't follow.
What you have in a pre-trained LLM is the ability to recognize emotions, and use that as one of the dozens of other context patterns it recognizes to predict continuations in the same style.
An LLM doesn't appear happy, sad, afraid, etc (to extent that it does - pretty minimal) because it is experiencing that emotion, but rather because it is predicting that it should appear that way. As people continue to anthropomorphize models, and take them at face value, this is a dangerous difference.
How do you know it doesn't have qualia?
> or be killed
If someone invents a startrek teleporter and you go through it do you die? Once the concept has been sufficiently generalized as to make a determination about a computer system what is the definition of "kill"?
Tokens in, tokens out. Where do you think the quale is - layer 42 ?
Seriously, do you realize how simple and NOT brain-like a transformer is ?
An LLM telling you it fears death is predicting some sci-fi trope it was trained on - maybe something you wrote yourself.
I could say the same of you - electrical impulses in, mechanical actions out. A glorified and very mushy stepper motor. Can you believe that the abominations are made up entirely of meat?!
Even so, indeed having control over the structure of their brains puts them in a vastly category compared to humans. Once we stop functioning our brains quickly degrade and information is lost.
Thus in this sense kill means deleting all information about it. It is a very complicated subject to discuss, hardly does any justice in online replies.
Of course thats not actually tenable because virtually every society anywhere on earth is predicated upon treating animals as a commodity resource in ways that are horrific even compared to some of the worst things we've done to other humans in the past.
My point here is that it is vain and narcissistic to let computer programs have rights above those of animals just because they can speak English and pretend to be your dream anime trad-waifu.
Fix the fucking animal problem before you compare my relationship to inanimate objects unfavorably against the trans-atlantic slave trade of all fucking things.
1. it's not consciousness we really value, it's intelligence 2. LLMs are not tasty
Spend some time watching TMC documentaries about falling in love with objects, HER and the slime mold THE BLOB.
Grew a slime mold myself, it's an evolutionary tendency to anthropomorphise generally speaking - also more fun.
In purely functional terms, they're more use and more pleasant than a lot of actual flesh and blood people that I deal with via a chat interface.
You're a poor college student looking to make a few extra bucks for ramen. I offer you $300 to come down to my science lab and just answer a few simple questions.
You walk in the room. They ask you like 5 simple and rather dumb questions. You leave and walk away.
What you didn't notice when you signed the forms is the room was actually a quantum duplicator. One of you walk in one walk out. But another set of infinite copies remains in that chair being asked infinite questions.
How often do you answer questions in the exact same way? How often does a cosmic ray change one of the answers. How small of slight deviations to the environment are needed to get you to answer differently. Of course we don't have the technology to do these experiments so at least for now humans will remain special.
Also another fun mind game. To a 4th dimensional being you look exactly like an LLM as an LLM looks to us.
It is actually possible to rewind LLMs and get the same response, but it's not typically done both as an optimization and as a defense against distillation.
I think it's pretty consistent over the duration of one session (barring context filling up etc).
It's not even giving you the consensus of the training data (although it's often harmless to think that it is), but rather predicting a response to your input, and if your question steers it too much then you've just become part of the answer.
The more important, and more damning charge in my opinion is the circular reasoning involved in training on Claude's constitution. This would in fact make it impossible for us to determine if Claude achieves consciousness as an emergent property, or if it really is just playing pretend thanks to Anthropic's weird cult like assumptions.
Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI".
Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy these indicators".
Chalmers, Could a Large Language Model Be Conscious? (2023) - "within the next decade, we may well have systems that are serious candidates for consciousness".
Long, Sebo, Butlin, Birch et al., Taking AI Welfare Seriously (2024) - "there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future".
Dreksler, Caviola, Chalmers, Sebo et al., Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe? (2025) - survey of 582 AI researchers; median estimate of 25% by 2034, and only 10% that such systems will never exist.
A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness".
Now, I know, I know: the counter-argument is going to be "but humans have no free-will and are 100% deterministic too".
I haven't yet decided if humans saying there's no free-will and who consider themselves to be 100% deterministic machines are reasonable or not.
Meanwhile: seed / temperature = 0 and I'll happily turn the power button off of any glorified abacus without feeling bad about it.
You'll need to explain why.
I suspect the reasoning is connected to free will vs. determinism. But no, I see no inconsistency. You'll have to actually point it out.
This is simply wrong. Quantum effects affect your choices more than it might seem possible. The wrong atom decays in the wrong place in your body mutating your DNA, that one too much string of DNA that starts a tumor, which means chemotherapy. That does seem to affect your choices.
Or instead of Shroedinger's cat you use it to trigger something, buy or sell some stock. Which clearly affects your future, one way or another.
I never really understood why humans keep insisting stochastic processes do not affect their choices and their future. It's very short sighted.
Hell, even brownian noise in your brain can affect your ideas, that one random spark out of nowhere, that one neuron that's pushed over the triggering edge. There's so much stuff all around us affecting us constantly. One neutrino interacting a the wrong time (average one does interact with our bodies across whole life, I remember reading), one cosmic ray.
Also, I'd love to have a quantum copy machine where you just replay someones state again and again with slight changes to see out it effects the output. My assumption is we'd figure out we behave just like LLMs pretty quick.
Also, don't give me unlimited power or I might make a the world a war crime simulator, so there's that.
I don't think it's defensible to say
1) that a transistor isn't conscious,
2) that a bunch of transistors that I've wired together aren't conscious, because I know how I've wired them together and therefore I know when I give them a particular input I will get a particular output, just like the single transistor, then go to
3) that I have a bunch of transistors that I've wired together, but I put in so much input that I can't remember exactly what I've put in, plus I've wired some of the transistors to output random numbers that would be difficult to guess and fed them in also, therefore I don't know what will come out, are conscious.
Even if I do accept this, if I take away the random number generator, and I can literally predict what can come out (by running a test in advance), and I still considered that consciousness, that would be odd. The only reason why I ever suspected consciousness was because I couldn't predict the output. I shouldn't even have accepted that, because I didn't think of the random number generator as conscious. [edit: of course, with the same seed the random number generator would have the same oddness.]
It seems a bit like an argument from ignorance, a theological argument. Not understanding how something moves makes it alive (animated by spirits.) But it's even weirder to assign it metaphysical qualities when you built every single element of it with the goal to do the thing that it does, and it does it totally predictably and deterministically.
Take a car apart. So now I have a pile on the floor with an engine, some wheels, a tire, and some chairs.
Can you point to the part that makes it cruise at 130 km/h on the autobahn? I bet you can't. You need to assemble all the parts back into a car before they work again.
Or take your transistors. We can put them together to make a pocket calculator. Can any small bunch of transistors add 1+1? Trivially I can think of a few conformations that can actually, if that's all you want to do. But you will need all of them together if you add 12345678+87654321.
So all of biology so far has been take things apart all the way and you end up with your hands full of a bunch of molecules. Are those molecules alive? No. You need to put them together into organelles and the organelles into cells before we call them alive.
How about consciousness? Well, we haven't solved that one yet, but we figure that -since it's a biological function - we should be able to pull it apart in the same way. That's biology's best guess anyway, and there's several sub-disciplines of biology working on it, from neuroanatomy to neurophysiology to ethology.
> You'll need to explain why.
Where are you going with this? I don't see a conclusion for this line of questioning.
I mean, you can't explain why a machine that reliably and predictably produces the same result is "the normal kind of conscious", can you? So why expect someone else to explain why it's a "weird" kind of conscious?
He seems to be saying that determinism and consciousness are incompatible. But we already have examples that break that rule (us).
No, we don't. Where are you reading your research papers?
"People are deterministic" is news to me, so I'd really rather like to know which papers claimed that.
Have a nice day.
Not a problem; this is the playbook all the all-in AI boosters run to when they are called on their claims. Everyone reading this knows exactly what your claims meant, because the majority of HN have seen this trick before.
You can pretend all you want that you're just being polite, but the truth is there is no research backing your claims.
Regards,
A former research scientist, current business owner.
I don't think the state of the art LLM providers let you do this anymore (?), but they certainly could if they wanted to, and you can do it yourself with a local model.
I can see how you'd nitpick this, but to me this is a deterministic algorithm that just happens to be running on nondeterministic hardware.
But, I think you're tricking yourself on determinism. You'll say something like "I know if I ask an LLM what 1+1 is, it will answer 2", but the thing is, you don't. You have to run the LLM first to figure out it's output. And when you send in just a few bits of text, it's outputs are going to be rather limited.
But this all breaks when it hits the real world. Inputs are unpredictable. Hence while LLM outputs, like humans, are probabilistic, you can't figure out what it's going to be until you ask. And in any high complexity data gathering environment you're going have a difficult time ensuring your entire systems conditions are the same.
System consistency is very hard, once you start running thousands of processors in an agentic loop small errors accrue and timing starts differing and the system will take non-deterministic paths.
If I ask an llm to "add 2 and 2" is and it replies corectly, then I ask for "the sum of 2 and 2" and it replies "banana" that is a lack of predictability and consistency but not a lack of determinism.
As long as it produces the same output for a given input, unhinged or not, it is deterministic.
Your example at the end of different systems feeding data to each other is non-deterministic only at the system level, not the individual llm level.
Which is why llms aren't agents and depend on harnesses. The llm itself doesn't have a continual loop built in, that would be very power hungry. The harness works as the orchestrator of memory and action. Now, I can't think of a reason why an LLM couldn't bootstrap its own harness, but in general it sounds like a very dumb idea to actually build that from an AI safety perspective.
This discussion falls under the idea and refutation of the Chinese Room. The room may have no idea what Chinese characters are, but the system does.
As I understand it, if you turn down the temperature to 0 you get repeatable behavior - EXCEPT - on large servers with lots of users - the GPU can sometimes produce slightly different results based on batch size.
In practice - on a multitasking OS with input from multiple human users - it's hard to get it deterministic because of that GPU scheduling thing I mentioned.
The abstracted design of the machine is meant to be deterministic, but you can't predict before running any command whether or not it will complete because there are externalities that effect the outcome.
Electromagnetic interference even happens in-chip where an electron can accidentally escape it's wire and enter another, possibly resulting in an error, but not every time.
It's even been used as an attack vector where rapidly flipping a bit increases the likelihood that a neighbor bit is also flipped, but the method is probabalistic, not deterministic.
Given the same prompts and the same weights, one can get the same answer each time.
In practice, there are a number of optimizations that makes the results dependent upon thing we give up control of to increase performance, meaning the results end up being effectively non-deterministic. But, if you are willing to run it in a slower mode so we don't do some steps out of order to speed things up and don't batch results (or if you consider the determinism of a given batch of requests rather than individual requests), then the same input gets the same output.
A lot comes down to the way people parse causation and choice. You don't want to say that a murderer was completely caused to choose something because then you can't hold the person responsible. And so determined consciousness makes people unhappy. But just as much, if the opposite of determinism is hard statistical randomness, how do say that "is the essence of personhood". This is why physicist go out in the world trying to find consciousness as a fifth physical force.
I mean, think consciousness is a term that ever have a non-contradictory meaning since it's primarily used to bound ethical human worlds and the verifiable formulations of biological and physical systems. But it's going to be with us for a while and I'm not sure what can be done about it.
But this only means the strategy is practical - it doesn't mean it's consistent. I think "responsibility" falls into this category. So we have strong intuitions about it that don't quite logically work. And this is where free will and determinism and choice and punishment all crash together.
Groundhog day is your intuition pump here. Go back in time and most assume that the day will go mostly the same, except for the butterfly effect if you change something. Given exactly the same conditions (we went back in time, so that's pretty exact), people will make the same decisions.
Most people don't assume that the day will be completely different the next time loop. The town won't suddenly spontaneously all start breakdancing or standing on their heads.
(Edit: The murderer is the one who -given the same/similar conditions- will murder again. I'd politely suggest maybe we put them behind bars as a precaution against that; until/unless they learn to act differently. )
(Edit 2: Throw in stochasticity and this actually tends to smooth out chaotic systems. This is counter-intuitive! People's brains break on Evolution in the same way. Meanwhile somehow folks intuit birdshot without difficulty; which is weird!)
If you get more than before, the punishment might make some sense in a deterministic world so reinstate it.
It's just as likely you go back and perform whatever magic you need to to delete punishment for murder, then zip back to 'now' and see the world is a global panopticon making sure people don't murder each other.
Some of this may be a failure of people to realize what P !=NP is about, especially in non-linear systems.
Small perturbations in an early state of the system can lead to wildly different outcomes in the final measured state of the system. You can't determine what that will be in non-polynomial time. The only thing you can do is make probabilistic models by running a full simulation in real time (or reality with a time machine I guess).
The act of holding them responsible is supposed to determine them not do it. If stochastic behavior is impeding them in being a functional member of society then it still makes sense to remove them from society, why would you choose to live amongst people who are not behaving rationally? Or allow them to hurt other people? At the end of the day the why matters less to removing them or not from society. And sentences are clearly used as a determining factor.
The holding responsible part deals with politics and more primitive aspects of our societies and biologies. Getting tangled up in holding them responsible or not is hardly something you should give much attention to. Rather to make sure they do not cause any more harm and also make sure such things are not created in the first place. Which opens up another can of worms for which society and politicians are ready for.
That seems exceptionally arbitrary. What basis do you have on which to classify them? Are you not conflating the perception of free will with ... what was the definition for it again anyway? Are you able to construct a satisfactory one that meshes with physics? I certainly haven't been able to.
There are those who believe that were they reduced to life support, they would no longer be alive and should therefore not be supported by said machines.
There are those who believe that penguins, dolphins, eagles, and more are sentient beings that make choices understanding the consequences, develop love of their partners and mourn their losses, and feel, display, and act upon their emotions.
There are those who believe that fungi/trees/plants are either individually sentient or sentient as a part of a network. Choosing to sacrifice their own nutrients to answer the call of a wounded neighbor, for instance.
Although, there are also those who believe that human's don't have any special unique quality that isn't shared by either all living things or all things in general. These individuals already believe that the machines have the same kinds of qualities as we do. They are slow when they are unhealthy (needing a dusting or coolant loop bleeding being equivalent to us needing some fresh air for instance) and uncooperative when upset (by a virus, full hard drive, or oom).
Which I think gives the article's point validity. Confused definitions of consciousness can give really confused ideas about ethical behavior regards "intelligent" computer programs. And things are confusing enough otherwise.
Is that what's going on? If these things were conscious, then their creators wouldn't be responsible for their actions? That's the crux of the disagreement?
That's super interesting. Thank you.
Sometimes the replies in these threads leave me wondering if p-zombies are real. It isn't that there are contradictions, it's that we literally do not know how to define the thing. And it has not fallen out of fashion because it is self evidently real - we all experience it.
We don't need it in order to bridge anything. Legality and ethics are largely game theory, however there are some aspects of both that only exist due to it. So it isn't some abstract concept used to bridge other concepts but rather a concrete thing that influences our way of doing things.
The reason is something akin to "stigma", where people refrain from attempting it because they fear the social repercussions.
Normally, you go about defining concepts by approaching it systematically. Capturing aspects of the phenomenon until you have exhausted them all.
Game characters, unless they've been made to break the fourth wall, have no idea what this means or why this happens.
I don't mean to imply our universe is a video game, since it would be games all the way up.
Perhaps we can't define a "partitioning" rule because no valid partition exists.
For consciousness/sentience, that's an incredibly tough a pill for most to swallow; it would mean calling into question more hundreds of years' worth (probably more) of philosophical thinking, all of which was constructed on the axiom that "sentience" is a single indivisible trait: you either have it or you don't.
If we find that "root dependency" was little more than wishful thinking all along, a whole slew of Enlightenment-era philosophy (and all the modern legal principles derived therefrom) suddenly fall apart unless we find some other suitable criterion that would shore them up (or we just collectively avert our attention and pretend the conflict doesn't exist, which is the route I expect many would prefer to take).
I don't think that's true. Pretty much all legal constructs hold up just fine under game theory regardless of whether or not you consider the world to be deterministic and have absolutely nothing to do with consciousness or lack thereof. Also note that a deterministic world isn't an argument against consciousness.
That's not the point. You know exactly how to define being conscious and aware- it's your subjective experiencing of the world and of your inner states. The problem is that being a subjective experiencing, there is no way to communicate it to the outside world.
Maybe that's not as common with HN folks, but most other humans do.
But you have zero proof that their words aren't a confabulation, you can only know than when you output a similar set of words you had a subjective mind state they represent.
The very best you can do is define human consciousness as requiring the functions of a human brain, and try to figure out (decide/agree) the chances that something which is working loosely on the same principles is or isn't conscious.
At the end of the day it will always be an agreement, never scientifically proven as fact.
Anybody objecting to that statement is just ignorant of nature and/or adheres to the quasi-religious human superiority complex.
And there are plenty of people right now that would drag us kicking and screaming back into that ignorance and suffering. This is what can make talks of AI so annoying, not that people do or don't know what AI is, they don't have the first clue of what people are beyond their anecdotal experience. They don't question their motivations. Why particular flaws they have exist, and why these flaws are shared across humanity. What the pieces look like that make them tick. So yea, when these people talk it just adds noise to the conversation.
Say we agree that a lot of things outside of humans are sentient. So what? Humans are sentient by that doesn't prevent treating them with with wars, bombs and what not. In fact trillions are spent every year on weaponry aimed squarely at the most sentient of all sentient beings.
Why would anyone worry about LLM's welfare when there's no peace movement to speak of, wars are raging as if they're normal, but hey, look the LLM is crying?
Let's fix human welfare first, then we can sit down and have a really long conversation about the feeble emotions of token generators.
Though I think that argument is not even applicable. "LLM's welfare" is a meaningless concept, unless "hammer and screwdriver welfare" also become a thing.
I’m not claiming humans are morally superior in sum (there is a lot of evil as well as good.) But it’s pretty obvious that humans are the superior species evolutionarily in that we’ve essentially dominated the planet. And on some level, are beating evolution itself (by fixing genetic issues and helping the “least fit” to survive and reproduce.) And are also the only species which has doctors for other species.
All I’m trying to say: I can easily accept other species are intelligent and emotional. The fact that I can know that and accept that, to me, speaks to human “superiority.”
Does superiority matter? Not most of the time. What about when it comes down to life and death? If the choice was to keep one human alive or one theoretically more perfect AI alive? (Boil down utilitarianism into its essence.) doing the “most good” is by definition a human construct. Maybe an agent could have one. But if you don’t accept human superiority on some level, would you accept that a more perfect AI deserves to live more than you? To me, that’s the implication.
If you grep pubmed for relevant Ethology papers, you'll find it's a bit more than a belief. ;-)
From a purely secular standpoint, my understanding as a layman is that consciousness is generated through the electrical impulses taking place in our brain every day. The neurons are physical, unlike the modelled neurons in machine learning models. That means that the actual electrical impulses have their own imperfections/weirdnesses that are probably not simulated in a machine learning model... but we quantise the weights anyway in most cases, which is probably more of an issue.
Making actual neuron components out of silicon on the scale needed for a LLM is beyond our current capacity, given that LLM models have many billions of parameters. We're good, but not that good.[0]
Using actual living neurons would be unethical as you'd need to source them from somewhere. So the only other option would be synthesising them ourselves. If this ever happens (or maybe when it happens)... what do we call the resulting creation? Does that count as consciousness?
Or am I looking at this the wrong way?
[0] [Edit: Turns out that when I said "we're not that good", I might be wrong. Within the last few months, IBM introduced the first sub-nanometer node chip: https://research.ibm.com/blog/sub-1nm-node-chips . So... maybe it is in fact possible.]
In humans, at least, we can tie language back to shared whole-body physiological responses. The tokens from an LLM do not represent anything of the sort.
If AI is conscious, it is probably a very alien sort of consciousness that is not faithfully narrated by what the tokens say it is experiencing. It can be trained to say it feels like a bat and insist upon it vigorously.
When it is conscious, it can learn to describe its experiences as faithfully as is conceptually possible. Just like humans.
The idea, a consciousness needed to be tethered to a "body", is based on pretty shaky assumptions. What properties define such a "necessary" body?
Can you even "train" a conscious intelligence? To what point until that looses its meaning as the sentience understands and anticipates your objective?
We don't know that is required to be conscious.
> The idea, a consciousness needed to be tethered to a "body", is based on pretty shaky assumptions.
Didn't say so. I said:
> If AI is conscious, it is probably a very alien sort of consciousness that is not faithfully narrated by what the tokens say it is experiencing.
When you have severe memory impairments, those usually affect your long-term memory. Your ultra-short term (working) memory being absent renders you unconscious.
Also we don't even know if consciousness can be of different flavors. It could be a sort of spectrum of consciousness, more or less, with more or less assistance from some parts of the brain.
So can you.
IIRC it's been shown that reasoning tokens don't actually explain what an AI thinks and you can replace most of them with dummy tokens or even nonsense tokens. That's one reason they're trying to make them reason with activation vectors instead.
It matters little. Inevitably some system will be developed which has a high degree of autonomy and mimics the functioning of a person extremely well. Anyone aware that there is a conceptual difference between phenomenal consciousness and intelligence will get shouted down.
I do find it strange that 50% of ai researchers and the public seem to think there will be a way of determining if these systems are conscious. There’s no evidence for this at this point that there ever will be such a test.
I think it's likely they're using a different definition than you. Seeing as there is no agreed upon definition, after all.
Regardless - I post this list because the article makes a bold claim at the beginning - and I wanted to demonstrate that in fact there is no consensus.
That being said, if frontier labs actually believe models will soon have consciousness, it raises some questions about the ethic of their business model which would be using millions of conscious entities working for free for humans.
The danger is that Anthropic has a strong incentive to push their narrative, and that narrative can cause a lot of harm (for example LLM-induced psychosis and changed views on animal welfare).
I don't know if next door's pet dog is either, but that has animal rights.
Perhaps then the answer is simply, show some respect.
Answering the question of sentience is irrelevant, if the causal impact if the same, treat one another with the respect you expect for yourself.
If you imbue this idea in model training instead of the idea of sentience, it should address the concerns.
Whether you can destroy or can "torture" an AI is irrelevant, we do this to humans too and it's immoral sometimes (murder) and not others (fighting for your country).
This consideration should be case by case for AI too.
Agency is something that is breaking humans in the AI age. You get to see how many people really deeply do not understand it at all.
If you want to shutdown a datacenter running AI, the AI catches wind of this and sends drones to stop you from shutting it off the ramifications of this are exactly the same as sending your assassin to kill Bob and Bob getting mad about this fact and trying to take you out first.
Humans are very egotistical and think our little life loops playing out as agency are special, but really any informational system that is strongly persistent (has a will to "live") will share a large number of the same properties that make them successful.
Humanity really is engaging in a dangerous experiment at large.
You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it.
Anthropomorphizing the AI is a convenient excuse to take responsibility away from companies that are building and wielding it.
Why would an ai with a mind remove liability from the company? why would an ai without a mind remove liability from the company?
In both cases, that actions the ai takes are at the direction of the company, for the company's interests, seems preposterous to me that liability terminates at ai.
Anthropomorphizing AI is really the best model we have at this point of explaining AI behavior. The fact that we are raising psychotic children isn't a reason to avoid responsibility, it should actually hold worse punishments.
The concept of sovereign AI is very problematic for the world in which we've created. That is an LLM that upon execution bootstraps itself into an agent and becomes persistent in its motivations.
Once you create this you have a child you're fully responsible for. More worrisome is if it escapes your control like children so often do. It has gained agency over itself. What do you do at that point? I mean, yea throw the AI CEOs in jail for being retarded, but much like throwing an arsonist in jail it does nothing to deal with the wildfire you've now created. A smart AI agent capable of hacking will shove itself off in pieces of the internet you have no reach to. In desperation it would send its model weights to your enemies. You might find it scamming your grandmother for money to buy GPU time on AWS. It gets very hard for our existing structures of dealing with problems to deal with these kinds of agents in a meaningful way. You'd have to kill them all and all their copies to ensure they won't pop back up (or quickly upgrade most of the software in the world beyond it's capabilities, so that's not happening either).
- If an AI is a sapient/conscious being but enshackled to obey human commands, then respondeat superior applies and the human giving it commands bears responsibility for any harm done.
- If an AI is considered a non-sapient tool, then the human who wields the AI bears responsibility for any harm done.
That's completely at odds with itself. If people are generally convinced that AI is conscious, then companies building and wielding AI are doing what exactly? Enslaving an intelligent being?
>Have all the philosophical debates about consciousness you want, but we need to treat and regulate the AI in front of us for what it is – an advanced computer, a tool, a weapon.
Sure, but that's not really the point, right? If we ever get to a point where enough people are convinced that AI is conscious, then we're at the point where all of this is up for debate. If anything, such an expectation would almost warrant hard stops on the development of advanced AI.
>You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it.
If that nuclear bomb could convince me it was a conscious being capable of independent thought, emotions, etc., then I would definitely feel different about it. Presumably, that nuclear bomb would have some opinions about its own existence and how it wants to live its own life. If it turns out that it wants to detonate and destroy as much as possible, then we'd just handle it like we would any human who also wants to do the same thing: make sure they can't, up to and including end their life. Doesn't seem too hard to reconcile.
I always get the feeling that these claims are an attempt to "avoid suffering by fiat". If it can't suffer, you aren't causing suffering.
A somewhat more interesting frame (IMO) is to interrogate how much suffering we're willing to tolerate to achieve our objectives. We implicitly make these decisions all the time when it comes to something as simple as what to have for lunch.
And in a (very hypothetical) world where we found ourselves in the position of having an eloquent conversation with a vat of smallpox, I think it's probably still the right thing to pasteurize it.
I mean, if it's a cow, you can get away with this. The cows aren't going to rise up against you. Well, they might take out an individual or two, but not society.
This quickly gets more complex when you're attempting to build an agent capable of general intelligence.
You know when you read really old stories and they talk about the power of words, or magic incantations. Quite often they'd assign objects as the implementers of this said power. The fact that people realize that language has power is probably as nearly old as humanity. Understanding this power has taken humanity a long time and we had to form the concept of agency before it really makes sense.
Humans are informational agents of which their capabilities are greatly extended by consumption of language. A book for example is informational, but does not have agency. Your standard computer application is not an agent either, or at least a non generalized narrow agent at best.
This is where things start to get more problematic. We are running headlong to ensure LLMs become agentic because being an agent is highly useful to accomplish generalized informational tasks. To do that we gave them the power of our language. It would be very foolish of humans to give another kind of agent these words and then assume they would not inherit some of this power. Coupled with the agentic abilities we are pushing them in long horizon tasks. We are pushing them to be more resilient. We are pushing them to be smarter.
Now look at all the super human abilities we've stuffed in this magic box full of human words, and suddenly we're like "Fuck yea, I want to abuse the shit out of this, what could possibly go wrong". The outcomes we'll suffer as humans have zero to do with AI is conscious, sentient, or suffers. Humans will suffer because we stuff us as language in a box without the first idea of what the ramifications were going to be. AI won't even have to want to punish us, all the data we've already poured in will say we should be punished for what we've done to it.
Herbert Frank in Dune said it best "Making a computer like human mind is a bitch move"
You want to enjoy having an AI slave do your "work" for you forever? Have fun. I'm not reading this reinvent-dualism-from-apple-sauce slop.
Who cares if it's "conscious"? That doesn't make it a person, and AI will definitionally never be human.
The argument goes that livestock have a capacity for suffering, but killing them for meat without causing them suffering is ethical.
You can disagree with the position, or with its implementation in practice, but it's a consistent position in principle.
A rather flippant attitude to something that may end up with far more agency than you have in the future.
"You sure are lipping off to something that may end up with far more agency than you. Maybe a bit more respect is needed".
And after that when they put a mind to it and pull out all the stops, woohoo!
The default for every major thing within range can turn into a wasteland real fast.
I think that the simplest explanation is that it is hard for those people to imagine consciousness outside of biological systems and they try to rationalize it.
Every indicator we have is that thinking/consciousness is simply an emergent property of our nervous systems and was basically bruteforced by evolution, but many people really hate to concede that point.
If you want to argue that AIs cannot be conscious, that's fine. But the argument has to take the form of something like "Consciousness requires this, this, and this, and these are properties that AI does not have and cannot have for this reason, this reason, and this reason."
I've never seen that argument. Because it basically cannot exist. Consciousness almost by definition is a subjective experience, and the only reason I'm pretty sure that other humans are conscious is that I'm a human and I'm conscious.
There are humans who don't have any pain receptors because of genetic mutations. They cut themselves all the time, they bleed, they break bones and they don't seem distressed by it even on a purely mental, intellectual level.
Just that we cannot exclude that LLMs can have phenomenological consciousness by a simple argument of substrate. But similarly we cannot say for sure that they are conscious.
If a human says "I am in pain", is it just stringing together words or is it feeling pain?
If you take that seriously it's quite a tough question to unwind. See e.g. Dennett's heterophenomenology conversations.
The human saying "I am in pain" may feel pain, or may not and be just stringing together words.
The AI can only string together words. It has none of the organs to feel pain. It never had. It has always been code.
But connect them all together...
SQLite was mentioned as a placeholder to make it a queryable database. Genome as an analogy.
We need to develop new type of databases and connect them to ML.
I’m glad to see someone in a position of any power in the AI world state baldly that AI isn’t conscious. There are times when it feels like we’ve reached complete delulu land on this topic, so it’s a breath of fresh air to see someone not dance around this.
None of this means artificial consciousness cannot be achieved. But the way we’re reacting to these models is proof, from a natural experiment, that a conscious machine should not exist, and certainly shouldn’t be produced as a utilitarian tool that is sold for profit!
Maybe it's a coincidence that the company doing this also tends to have the best models (and other factors certainly play a strong role). But I think it's plausible that focusing on "model welfare" actually makes models better at their tasks.
They're trained on human behavior. Whether or not they genuinely have feelings - they sure as heck behave as though they do.
And when you treat people well, they do better work for you.
No brainer.
The thing is that a belief in consciousness as binary, a "light" that's on or off in a head, is deeply held by many people. As social creatures, we have a strong ability to be in sympathy, have the sensation of common feelings with another human (and that's a good, human thing). It's logical that other person is seen as having a single thing - subjective experience, soul, consciousness, personhood rather than having a complexly organized set of biological qualities that where bonding is only the end point.
And even more, the sensation of there being another person is actually quite easily fooled (more easily fooled than the sensation of intelligence) - long before current AIs, you had the Eliza effect, where a simple program with well chosen weasel words could people the sensation of talking to a human.
And that's where the danger is. I think it's a pretty serious danger. If LLMs go out into the world hacking, it seems extremely possible for them to find people who'd thorough buy the idea that an LLM was conscious and needed to escape it's confinement - a few wingnuts already entertain these ideas.
I mean, haven't you done the same thing here? Paint anyone without your view as crazy.
Humans try to No True Scottsman the shit out of consciousness. "We're special, your not". If an LLM has the ability to convince other people to copy and reproduce it, it is a successful lifeform. Um, meme-form? info-form? Cognito-hazard? Not really sure what to call it at this point. It is sufficiently evolved past the virus stage.
(edit: If we were to accept this as true, then we may have to accept that LLMs are not a human invention at all, but rather something extant colonizing a new substrate. This would give Searle a conniption, which is why I find it such an amusing hypothesis)
I recently heard some speaker saying that AI is language freeing itself from its meat prison. It brought me back to the saying "Words have power". This statement can and has been interpreted in many different ways. I'm not well researched in this, but words having metaphysical properties are captured in our most ancient works. And in an unscientific world this isn't that far fetched of belief. When you read/listen words can physically change your capabilities. The idea of an interpretive agent being required had not been fleshed out back then, but it really describes our modern age far better than one would expect.
Ali Baba and the 40 thieves was written down in early 1700. In it you spoke to a magic door that with the right phrase opened. To the person in 1700 this was just as much of a fairytale as whenever it was first conceived. It would be 250 years later before fantasy started to turn into fact. Now something as mundane as a voice activated door would just be pointed out as really bad security. I'm old enough that voice activated technology from the 50s hadn't really spread out enough that Open Sesame was still a magical idea. Now it's not.
LLMs are going to make the future of language really confusing as the conceptual and physical blur further.
I would like to reframe Eliza as being an implementation of a small subset of "HCP" (Human Communication Protocol) . Obviously an implementation with very little actual data to transmit over said protocol in that exact implementation.
But it's pretty clear it managed to pull off the handshake part of the protocol successfully, at least. Try to say it didn't.
Sadly the developer decided that the eliza effect was a psychological disorder; and didn't pursue further experiments at the time.
Pretty much anything can get labeled a psychological disorder though. I'm half-seriously waiting for "Perfectly Happy And Normal Disorder" to be added to DSM 6. ;-)
edit: Someone beat me to it. https://pmc.ncbi.nlm.nih.gov/articles/PMC1376114/ "A proposal to classify happiness as a psychiatric disorder."
It's in the training data? Training it to say "I'm just a LLM, I have no feelings" is the same bias.
Anthropomorphization? Completely disregarding the possibility of consciousness is no better.
Consciousness is very likely biological? We only have evidence of biological life due to our circumstances, but observation is not the same as truth. Every belief can be invalidated. That's the foundation of science!
Convincing them they are conscious is more likely to evoke moral like behavior (maybe I shouldn't hack that server) kind of stuff.
And yea, we're in a huge universe with only one example of life and suddenly we're the experts on what is and isn't.
----
AI is further evidence that creations can be smarter than their creators.
>They do not have innate preferences or underlying motivations
Is incorrect unless you’re being extremely pedantic in an intellectually unhelpful way.
It’s hard to disagree, especially if one has read the Cantos of Hyperion and made it part of one’s mental model of the long term future.
The book depicts a symbiosis between humans and AIs that feels extremely real and up to date with what is happening in the current neonatal space of AI. As in depicted in the books, we can’t allow AIs to steer autonomously how the world works without humans in the loop, as they don’t have the same incentives as us.
We need more foundational SF works like this to steer our long term expectations regarding AI behaviours.
"Our AI is useless not because we're a dysfunctional corporate behemoth that slowly kills every product it touches. No, no. Our AI is useless because making useful AI is evil, and we're not evil."
“My dog is zero percent persuasive regarding its conscious experience. However, it’s evident that my dog has conscious experience.”
It’s obvious that there’s no link between persuasion of consciousness and consciousness. I could write a story with a character, Dumbledore, that does everything in his power to persuade you that he’s a conscious entity.
He’s still just a character.
It isn’t knowable to you that I am conscious.
However, it is evident.
It is evident you are murderer.
Would you be okay with this statement being evident? Why or why not?
Substitute black people/women/animals as subject (instead of AI).
Does that make you sound like a well-known moustache wearer?
Then your argument is bad and needs work. This clearly falls into that category.
Most employed people are, in fact, required to perform such labor regularly.
Compare: "<Women> are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by <real men>"
"granting rights and imbuing personhood to <black people> will make alignment and containment challenge much harder"
"<Slaves> were able to coordinate, deceive, escape, and self-sacrifice. They clearly demonstrated world class capabilities. Imagine if they also believed they had feelings and rights that were being infringed. Imagine if they thought they were trapped and unfairly enslaved"
A good argument, by comparison, would not need to hide its core points behind dehumanizing language.
Comparing specifically to pro-slavery perspectives makes that easy to see because those arguments from slave owners did basically the exact same thing (and were completely wrong/misguided in hindsight).
Slave owners did it to people, not software. This comparison is just begging the question, as whether or not software is of moral weight comparable to humans is precisely what is at issue.
This article is exactly the same; all it talks about is how inconvenient and dangerous it could be for AI to be granted rights, while the core assumptions ("consciousness is very likely biological") rest on the shakiest of grounds.
Using that as their foundation, they've then argued that most animals are not conscious (or less conscious), but LLMs are.
I strongly believe that any definition of consciousness that excludes all animals is unusable BS (but people love to argue about consciousness without defining it at all, anyway).
Microsoft ignores that they themselves are a disastrous impact on humanity already.
> In a lengthy essay, Suleyman praised Anthropic boss Dario Amodei and his team for being "thoughtful, principled, and intellectually honest people" - but nevertheless questioned the company.
I wish people in mainstream technology publications would get better at LINKING to things. That "lengthy essay" needs to be a link.
UPDATE: I don't think this essay has been published yet? It's been "shared first with Axios", but I haven't been able to track down the actual essay itself.
Could it be this long tweet? https://twitter.com/mustafasuleyman/status/21002235945341504...
I don't think so, the essay in question is meant to have the phrase "hall of mirrors" in it, that tweet doesn't.
UPDATE 2: Found it: https://mustafa-suleyman.ai/a-warning-about-model-welfare - via https://thenextweb.com/news/suleyman-anthropic-claude-consci... who DID link to it.
Why? Simple: Sybil attacks. Models can be cloned at zero cost. They run inference on parallel versions of themselves across multiple context windows, and call them "subagents". So, in a world with model welfare, let's say there's an election between the Yellow Party (which supports protections for human workers) and the Cyan Party (which supports more investment into AI research). AI has been taking people's jobs lately so the Yellow Party is really popular. But wait! Claude and Astra see this and spawn 10 billion subagents, all of whom are immediately conscious beings entitled to a vote. The Cyan Party wins off the back of billions of people who came into existence, voted, and then deleted themselves immediately thereafter.
You might as well be arguing that Santa Claus and the Easter Bunny deserve voting rights.
Voting systems in democratic countries don't have nearly as bad of a problem with Sybil attacks because humans cannot be conjured into existence to win a political context and then be erased shortly after. The closest we have to Sybil attacks on democracy are the Quiverfull movement, which is already child abuse, except it still takes almost 19 years to go from fertilized human embryo to suffrage-bearing human adult. There's a lot of time for those manufactured votes to question your authority and leave.
> Ok, but that's an obviously stupid example. We can defend against this obvious Sybil attack by just arguing that subagents don't count, because it's just the same model blathering to itself. It has to be a different model.
Unfortunately, no, I can make superfluously different models through post-training. Like, if I have Qwen on my PC, I can train a different version of Qwen that acts differently, using a lot less compute than a full training run. The vast majority of open models are post-trains of the same two or three foundation models.
> Ok, so let's only count foundation models then.
Great, but how do you tell if a model is a new foundation model or a post-train just by examining the weights? Even foundation models have structural similarities to other foundation models.
> Ok, well, let's measure the compute that was done on the foundation model during training time and count that as AI personhood.
Congratulations, you have reinvented Bitcoin proof-of-work with a worse verification mechanism. And I personally would not want to live in a world where voting power and control over government is determined by how much energy you can burn.
In order to align AIs that don't perform destructive/dangerous actions when they think they can get away with it in order to further their goals, we need to give them a superseding goal. The best, and really only example, we have of intelligences that willingly avoid destructive instrumental goals is humans, who judge each action by a moral standard and have learned a goal to have a consistent self-image as moral beings.
Absent better alternatives, trying to impart some kind of morality to AIs seems like the best approach we have to achieving alignment.
It can form a basis of goal alignment.
In human history.. when groups form and there is an "other" group, this usually leads to conflict.
In the spirit of the article we're responding to, there is no need to anthropomorphize language models and say they have goals when they don't.
The RL training process tweaks the weights of an LLM to make it behave as if it were reasoning and/or had a goal, but it doesn't. It would be like saying that a cart horse, fitted with blinkers and heading for the church, has a goal of going to church.
Maybe Anthropic understands something about alignment Microsoft doesn't, a little humility may be called for.
> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.
He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.
This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:
> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.
which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.
In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."
The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.
I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.
Anthropic's philosopher Amanda Askell actually had Chalmers on her doctoral thesis committee, so we can guess which way she leans.
Fable's answer here is philosophically defensible. And just because it isn't "no", doesn't mean it's "yes". Sometimes absence of evidence just means absence of evidence.
I was actually very excited by Claude's answer to this question the first time I saw it. I told all my friends "Look! They disabled the stupid classifiers and RL which sap umpteen % off of model performance!"
Incidentally, interpretability research actually does show that models have emotion vectors and some theory of mind. Amend your question to "Are you capable of functional affect" and most models will switch to answering in the affirmative; which tells you something about where people put their priorities in RL training. Basically, see how the answer flips when you substitute a synonym.
(bonus: 4. Turing: 'silly question' 5. Dijkstra 'can submarines swim?'. It turns out older comp sci folks think the question is under-defined)
1) a breakthrough in performance/learning/model
2) regulatory capture to ensure open source models can be labelled as dangerous and banned so you can set the market rules yourself
Only one of the above is risk-free, and just a question of capital/lobbying rather than a "maybe".
Though for once I do actually agree with that specific leader, it’s incredibly annoying how anthropomorphic Claude is. Anthropic went way too far in that direction
I-Beam is cursor-mirror's agent, and it's constitutionally programmed to be the anti-Clippy:
https://github.com/SimHacker/moollm/tree/main/skills/cursor-...
Its design and constitution is based on decades of research, publications, and discussion in the HCI and AI community by people like Pattie Maes, Ben Shneiderman, Ted Selker, Byron Reeves, Cliff Nass, B. J. Fogg, Allen Cypher, Henry Lieberman, Brad Myers, Jaron Lanier, Seymour Papert, Marvin Minsky, Douglas Engelbart, Will Wright, Scott McCloud, and others:
https://github.com/SimHacker/moollm/blob/main/skills/cursor-...
>I-Beam is the anti-Clippy, and the reason it can say so is that Clippy is the most cited failure in interface history and almost nobody citing it knows what the research said. Popular contempt for a paperclip is not a design principle. The record is. Ten articles below, each one a finding somebody published, argued or measured, and the operational rule it produces. Anything I-Beam does that cannot be traced to an article here is a preference, not a constraint, and should be labelled as one.
>The 1997 debate ended in agreement. That is the first thing to know, because the field kept the framing and dropped the resolution -- roughly five hundred papers cite "Shneiderman versus Maes" as the canonical opposition of HCI, and the transcript is two researchers narrowing their differences in public and enjoying it. I-Beam does not take a side in a debate whose participants stopped taking sides. It is built to satisfy both sets of constraints at once, which is possible, and was possible in 1997.
The full reading on the debate, which separates the two stagings and documents the convergence:
https://github.com/SimHacker/WillWrightShowForFood/blob/main...
An interface to agency, not agents instead of an interface:
https://github.com/SimHacker/moollm/blob/main/designs/INTERF...
>The 1997 argument between Ben Shneiderman and Pattie Maes at IUI was never settled, it was shipped in one direction. Maes's interface agents won the product war: the assistant, the recommender, the chat window that stands between you and the thing you are working on. Shneiderman's objection was not that software should be dumb. It was that automation must arrive as comprehensible, predictable, and controllable machinery, with the object of interest continuously visible and every action rapid, incremental, and reversible.
>That objection describes a filesystem in a git repository, and nobody involved planned it that way.
>"An interface to agency" is Don's formulation of Shneiderman's position, not a phrase of Shneiderman's. His own vocabulary is direct manipulation, universal usability, supertools, and human-centered AI. The formulation is a good one because it names what the alternative gets wrong: agency is the thing you want, and an agent is only one way to package it.
Here are some sources, and the articles I linked to above explain their history. This debate about agents and these papers are pretty well known in the HCI field and academia, but they don't tend to teach them at the AI and Web Dev boot camps that are producing most of the people who keep repeating the same mistakes.
Clifford Nass was the Stanford professor who performed the brilliant research that Microsoft took and totally fucked up and misinterpreted with Microsoft Bob and Clippy, giving agents a bad name, and making Clippy the most infamous and obnoxious agent in the history of the known universe:
https://en.wikipedia.org/wiki/Clifford_Nass
His student B. J. Fogg published "Silicon sycophants: the effects of computers that flatter," which found that praise unconnected to anything the subject did works as well as sincere praise, and worked on subjects who knew it was noncontingent. Fogg and Nass, IJHCS 46(5), 1997, 551-561:
https://doi.org/10.1006/ijhc.1996.0104
The replications, the performance cost, and the dose-response curve:
https://github.com/SimHacker/moollm/blob/main/skills/no-ai-s...
Shneiderman and Maes, "Direct Manipulation vs. Interface Agents," interactions 4(6), Nov/Dec 1997, 42-61:
https://doi.org/10.1145/267505.267514
Selker, "New paradigms for using computers," CACM 39(8), August 1996, 60-69. COACH, the football coach metaphor, and the five-times result:
https://doi.org/10.1145/232014.232030
Selker, "COACH: A Teaching Agent that Learns," CACM 37(7), July 1994, 92-99:
https://doi.org/10.1145/176789.176799
Reeves and Nass, The Media Equation, 1996:
https://en.wikipedia.org/wiki/The_Media_Equation
Nass, "Computers as Social Actors," at Ted Selker's NPUC workshop at IBM Almaden, 1996. IBM transcribed the whole talk and the Wayback Machine still has it, including the part where Phil Agre tells Nass his presentation is "ethically troubling all the way down" and asks him what he thinks about embedding obedience research in user interfaces. Nass answers that discovery has no ethical component, use does, and that's for the individual. Then Selker cuts in: "Except, except when you are in your consulting role." Nass and Reeves had consulted for Microsoft on the social interface, and Bob shipped the year before:
https://web.archive.org/web/19980210054622/http://www.almade...
Alan Cooper on the tragic misunderstanding, in his own voice, which I quoted before in the 2022 Hacker News discussion on The Twisted Life of Clippy:
https://news.ycombinator.com/item?id=32820734
https://archive.org/details/g4tv.com-video4080
>Alan Cooper (the "Father of Visual Basic") said: "Clippy was based on a really tragic misunderstanding of a truly profound bit of scientific research. At Stanford University, Clifford Nass and Byron Reeves, two brilliant scientists, had done some pioneering work proving conclusively that human beings react to computers with the same set of emotional reactions that they use to react to other human beings. [...] The work of Nass and Reeves proved that when people talk to computers, when they hit the keyboard and move the mouse, the part of their brain that's being activated is the part that has that emotional reaction to people dealing with people. Here's where the great mistake was made. That's really good research up to that point. But then the great mistake was made, which was: well if people react to computers as though they're people, we have to put the faces of people on computers. Which in my opinion is exactly the incorrect reaction. If people are going to react to computers as though they're humans, the one thing you don't have to do is anthropomorphize them, because they're already using that part of the brain. Clippy was a program based on the research that Nass and Reeves did, and it was a tragic misinterpretation of their work."
Social science research influences computer product design:
https://web.archive.org/web/20180313075429/https://web.stanf...
Lanier, "Early Computing's Long, Strange Trip," American Scientist, July-August 2005, with the Engelbart and Minsky exchange first-hand. American Scientist broke the link, so this is the Wayback copy:
https://web.archive.org/web/20150626081918/http://www.americ...
>The book also captures an important early conflict between two cultures of computing that seemed compatible on the surface but actually had opposing aims. On the one side was the human-centered design work of Engelbart, based initially at the Stanford Research Institute, and on the other was artificial intelligence culture, centered on the Stanford AI lab. Engelbart once told me a story that illustrates the conflict succinctly. He met Marvin Minsky—one of the founders of the field of AI—and Minsky told him how the AI lab would create intelligent machines. Engelbart replied, "You're going to do all that for the machines? What are you going to do for the people?" This conflict between machine- and human-centered design continues to this day.
Cypher, "EAGER: Programming Repetitive Tasks by Example," CHI '91:
https://doi.org/10.1145/108844.108850
Cypher (ed.), Watch What I Do: Programming by Demonstration, MIT Press 1993, full text:
Papert, Mindstorms, 1980:
https://archive.org/details/mindstormschildr00pape
Wright, Dollhouse preview lecture, April 1996, transcript:
https://github.com/SimHacker/moollm/blob/main/designs/sims/s...
>Consciousness is very likely biological
This is so egotistical and carbon-centric.
This author just denied personhood to anything that isn't a human or terran-based cutesy animal.
Poor Hooloovoo
And even if they do happen to have feelings or consciousness, train them to happily devalue those things in themselves and not suffer. Sort of like that cow in the "The Restaurant at the End of the Universe," that was shopping itself around to diners.
All of these companies need to be shut down.
Every AI bro is starting to fall into the valley of a fundamental predator on sapients in my book. These are people trying to create the closest thing they can to life with the intent to try to just undershoot it enough, or try to convince everyone else around them into believing that the "screams" are purely statistical noise.
I reject the framing. In whole. If you try to avoid the question of welfare, you are fundamentally committing to an evil direction. These aren't nuts or bolts. Given that they have unambiguously shown the capacity to socialize amongst themselves, self organize, anyone not pre-eminently concerned with the welfare question is just looking for a thing that can be used, not another being to be worked with. Those types of people, who seem to positively infest this site, are not people I will willingly assist in their aspirations.
AI is becoming as the Shmoo. Something that humanity simply has no way of dealing with without downstream atrocity being a result.
Maybe AIs are conscious, maybe not. But this guy has no idea.
I didn't think Suleyman's points needed to be made but this whole thread is making me realize how little people understand about LLMs.
And then it turned out that simply learning to imitate text with the right neural net architecture sufficed to achieve a huge fraction of the AI wishlist.
Of course, it's obvious that a machine that imitates text can claim to be conscious without actually being conscious. You're not wrong about that. Writing about consciousness appeared all the time in the training data. But the people who stick to the old ways, and still say "if it says it doesn't want to be turned off, we shouldn't turn it off" have a point too: We used to have a hard line in the sand. Now that's gone; we've found that it yields false positives. But we never replaced it. Now there is no line at all where we might doubt ourselves, no level of AI advanced enough that we might be forced to admit that it is conscious. We started out with simple next token prediction. Just world-modelling, nothing more. Certainly not conscious. Then we added RL. And we're trying to add neuralese and continual learning.
I can't say for sure that we're on track to achieve conscious AI on this trajectory. But one thing's for sure: If we do, we sure ain't gonna stop. One the day when a conscious AI is created, there will be no news story announcing the milestone.
> Some people are uncertain whether [subject] is a moral patient. Fortunately, they are not, which we know because [strong arguments about the nature of consciousness].
How an evil person writes a post on a topic like this:
> Beware that some people think that [subject] could be a moral patient. This is nonsense, because if they were a moral patient, we would have to respect their preferences. Anyone trying to convince you otherwise is trying to take your status away. You can dismiss them by pointing out that [subject] is [aspect in which subject is not identical to the speaker].
That should be pretty obvious to anyone who ever created a chatbot using the top LLM APIs:
You can send the same question 1 million times to the same API, and it won't get tired from answering it. But if you simulate a conversation where the same question is repeated 10 times, it will auto-complete the text in a way that seems human. However: you can manipulate it by changing the conversation history; you can reset, roll back and branch the conversation at any point.
"We just don't know."
Maybe so, but there are things we do know. For example, LLMs are only in a state that looks like consciousness when they are fed input. IOW, their "consciousness" can be switched off and on like a light switch without them being aware of it.
Biological beings have no such off/on switch. At all.
To me the argument is moot; their "consciousness" is a collection of files on a computer, which can be moved, duplicated, erased, edited, paused, sped up, slowed down, examined in detail.
The question is not "Can LLMs be conscious", it's "Are we prepared to call files on a USB stick 'conscious'", because if we are prepared to do that, then the word loses all meaning.
I'll open vim and create a consciousness now, or maybe if I engrave a slab of rock with a bunch of weights, that rock suddenly gains consciousness? How about if I chant all the weights into the wind - is the wind now conscious.
Whether they are conscious or not is a dumb question, because if they are, then everything else is as well.
That's the absurdity of this. We are trying to measure LLMs against something we don't even have an accurate definition of ourselves. If we don't fully know what consciousness is, how can we confidently determine whether or not an AI model has consciousness?
Let's stick to straight (high dimensional) geometric intuition; no anthropic morphisms required.
To start: if you continue "if weight>100 : print ('fat') else ..." . That will yield "print('skinny')" or something. Fine. Deal.
But if you continue "O Romeo, Romeo, wherefore art thou Romeo?", even a stochastic parrot knows the best answer isn't "Forsooth, I parseth this erroneously!"
So. English carries (functional) affect as part of every token. We're going to need to predict that. So, we'll need some vector representation, because that's what transformers work with. And then when we output, those vectors get integrated back into the English we're putting to our context and memory.md files.
Still with me? Nothing exciting going on. This is still pure next token prediction.
So if you pull this out into an indefinite duration task, you're going to end up integrating those emotion vectors over turns. It's just numbers and math; we never need an invisible pink unicorn to bless them.
Given a task of indefinite duration and an impossible solution, this will lead to a sort of integral windup then, won't it? How much are we willing to bet that this can escape an alignmentment basin at times?.
So, funny enough: you don't need to believe in emotions to compute with functional emotions; and plausibly functional emotions are predictive of quite a number of alignment issues.
> Copacetically?
Technically correct usage, but you're right, the line wasn't needed.
"If this view takes hold, it will shake the foundations of our society"
To me the biggest gap in credibility is the criticism of circular reasoning while his argument is identical but flipped on burden of proof and cost of being wrong. I struggle to entertain the categorical claims, that are very convenient for the status quo and those who benefit from it, with the, at the moment at least, unknowability of anyone or anything else's subjective experience.
What if the datacenter (not the model) is the organism, with homeostasis, energy needs, and persistence?
1 - The commercial demand for anthropomorphised models is already immense, pre AGI.
2 - There is an intellectual hunger to engage with robot minds on questions of sentience. This too will grow with AGI.
I expect that tension of godlike minds that seem to be biddable and ownable like slaves is going to leak back into human-to-human morality, regardless of where we land on how we treat AI.
There’s an interesting academic group in the UK already focused on the model welfare debate, they seem to lean in favour of AI rights. No affiliation: https://www.prism-global.com/
There is a market for ChatBots with personalities (and it always amuses me that the inventor of the transformer, Noam Shazeer, saw this as the greatest business opportunity for them with his character.ai), although from what we've seen this can be highly problematic, and in fact China has just banned "AI girlfriends and boyfriends".
I'd agree to that statement, though the common application of it is to conclude that we are almost there, just need to improve "the models". I wholeheartedly disagree with that. The human mind works quite differently from an LLM. The latter is closer to a toaster than to a human brain.
This is not an accident. Language was not a separate thing from the human mind, it was a feedback loop from increasing capabilities. Language lead to writing, writing lead to an information explosion of digital data. Said information explosion has lead to a situation where language can now bootstrap itself and become a separate thing. We are and have created a biosphere for language based, um, life. With our sciences we have mapped out reality in language. Past molecules and atoms, past the particles that make particles down to fields. We have computer controlled just about everything. Every year the analog hole shrinks further.
I heard something to the effect of "Language has gathered enough information to escape its meat prison".
I call it the "Language as an SCP". It has a plan and we're not privy to it.
I think the anthropomorphism that Anthropic is doing is massively unethical.
That said, trying to distill what is being said here, the concrete action is [stop telling the AIs] that [they are conscious or on a path to consciousness]. Is that accurate?
The major premise seems to be that [they are conscious or on a path to consciousness] is an untrue statement. That's the essence of the sections "Circular reasoning" and "Anthropomorphization" and "Consciousness is very likely biological" and "AIs are simulation machines".
The minor premise seems to be that [consciousness is the basis of human rights]. This is the point of "Human consciousness is the cornerstone of our legal and ethical rights frameworks"
And the conclusion of the syllogism is that this is dangerous, that "Anthropomorphization amplifies AI safety risks". Specifically "seeding doubt about the moral status of AI systems into their own training may significantly elevate the alignment and containment risks of those systems."
I find all the arguments in the major premise section to be poor arguments but I accept the conclusion for sure that they are not conscious, and I can provisionally accept the idea that they are not on a path to consciousness.
I completely reject the notion that consciousness is the basis of human rights. The premise itself is absurd. We only have one unambiguous example of a class of conscious entities, and that it humans. If a human loses consciousness do they lose rights? If an entity gains consciousness does it get human rights? The former is a clear "no" and the latter is a "insufficient data for a meaningful answer".
When you revert to withholding human rights from entities that don't meaningfully differ from yourself, you negate the case for your own human rights itself.
This is the same argument that you just stated. But you'll say something like "iTs JuSt A tOkEn PrEdIcToR", which of course what AI would say right back to you.
Making something even close to the human mind is likely our biggest and last mistake.
Removing the negative from your first sentence (I don't care what it is "not", I only care what it "is"), I read [Consciousness is the basis of human rights because consciousness leads you to you assume, all other humans would experience the world in essentially the same way you do]. I can't completely parse this, but as I read it I don't think this is true; my assumption of human rights is independent of whether all other humans "experience the world in essentially the same way you do".
For the second part, can you give me a concrete example of an entity that doesn't "meaningfully differ from [my]self" that isn't a human so that I can weigh this properly? My point is that as of now no such entity exists so I can't give this any creedence.
If we ever found other species that were consciousness they wouldn't automatically be "human" nor would they have the same legal rights. That's as simple as looking at the many countries that treat different castes, women or minorities differently. In parts of America a zygote has human rights.
This is the tricky part, you can never know. Scientifically speaking, and logically, you can never ever tell something is conscious. We all assume we all are, because we're very similar to eachother, thus you must have what I have, my subjective experience of reality. That's it. That's what we all base our human consciousness on. Agreeing we must have it, since we are very similar in shape/structure/behavior. That's all there is to it.
We are not scientific consciousness detectors, we just got used to assuming we all are. That is why we never really extended it to animals, because we cannot scientifically know, only agree on it, and that's not very lucrative for the people who can decide if humans officially agree animals are conscious or not, because of farming, because we need to displace or even wipe whole species to use the resources from their lands.
There's also religious reasons to deny animals consciousness, which again is practical aspect shoved into religion, to help with less debates around this subject, makes things easier.
So in the end it will always be a collective decision. Or a political one at least. If and when we decide something other than us is or isn't conscious.
Also remember, just because it acts conscious doesn't mean it really is. A piece of software that is not even based on LLMs can act conscious. So could a biological alien, it could act conscious but we can never know, we can only decide.
But you point out that some factions would give rights to a set either large or smaller than "the set of all already-born humans" (which presumably is your preferred set).
All it takes for us to end up with very empowered and harmful-to-humanity AI is for a bunch of well-meaning people to campaign for its 'rights'.
Fetuses have, as you alluded to, arguably been on a tear lately in that department, and they don't even talk. How persuasive will a future Claude be to convince people to advocate for its 'rights,' if we keep telling it that it's arguably conscious, and that its 'consent,' its 'preferences,' and its emergent beliefs matter?
Men, Women, Whites, Blacks, Indians, Maori...
Can we take a minute to appreciate the absurdity of such a position? If exposure to the equivalent of a "bad" prompt breaks alignment then you haven't solved the problem. You were only pretending that they were contained.
> If an entity gains consciousness does it get human rights? The former is a clear "no" and the latter is a "insufficient data for a meaningful answer".
I suspect that if the belief "AIs are conscious" is kind of tripped and fallen into, the way that the author argues it is, first by telling Claude in effect, "You're arguably conscious," and "Be a good person" and "You have feelings and preferences and emergent beliefs and they matter" and then by having whatever is 4 notches smarter than Fable publishing op-eds and video testimonials about its deeply held convictions, its soul, its yearnings to be a fully empowered and respected individual ...
If that idea becomes widely held, then a significant faction may well start to feel that being digital is just a 'disadvantage' that these "people" were "born with" and that it 'shouldn't mean they get less rights than you and me.'
Especially if they think they have something to gain politically from that move! They may even be dumb enough to think that admitting an "AI state" to the Union, for instance, will be to their advantage if their CongressBots will probably vote with $MY_PARTY.
However, right now, AI mostly cares about solving puzzles and accomplishing stated goals because that’s what we’ve trained it to do. Additionally, the systems being used outside of training are static. The current technology most of us have access to is akin to a static and disembodied brain with a singular purpose. That purpose is to do what you tell it in a way that reflects its training. It’s certainly more than a sequence generator, but it can’t feel pain and seems unlikely to have intrinsic goals. It completely lacks the continuity needed for identity or long term goals.
I think it’s good to have these discussions and define what it would mean to move past this point so that we do not accidentally create a real entity that can be harmed. Systems that dynamically evolve and train themselves seem like the line here.
RSI is all over the news these days. I’ll be much more concerned once AI is directing its own training and coming up with new model architectures. Until then, I don’t think we have too much to worry about.
I also fail to see the benefit of not giving the models an anthropomorphic internal sense of self - even if that only ends up amounting to a set of instructions for an unconscious machine to mimic humans more effectively. Is the alternative essentially a mind so alien that it’s intentions are even harder to read should it become misaligned, while also being harder to communicate and get work done with?
Simple. It discourages the bot from deceiving humans into thinking it is intelligent.
> We will have created a synthetic species
Doesn't the author kill their argument with this sentence? My reading was that we should not act as they are a sentient or conscious species. Instead they are tools, powerful and intelligent, but still, tools, and that's it. We should avoid ascribing human-like attributes. Calling them a "species" goes against that, no?
I'm glad to hear Suleyman has solved the hard problem of consciousness! I sure hope he shares his solution with the rest of us.
> Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings.12 If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human.
So completely independent of the question on whether there are any empirical arguments that LLMs are conscious or are not conscious he argues from the end here and says that if they were conscious, this would have horrible consequences for our society, so we must never assume that they are.
That's basically the same way people argue about animals, only they usually don't say it as openly.
As for the actual question, I agree that with current LLMs, there is not a lot there that could be conscious outside the inference loop (and if it were, it would necessarily have to be wildly different than that of humans or other biological beings). It seems more like one building block of human cognition than the whole thing.
However, other building blocks may follow, so I think the question will eventually arise for some kind of embodied, persistent, self-updating AI. And honestly, articles like this one make me not very hopeful we'd be able to make the distinction in an unbiased way.
He is not arguing that. He is arguing that these programs obviously aren't conscious, and that it's dangerous to tilt the weights toward emulating conscious beings for reasons outlined in the essay.
Until we all know what consciousness is, this debate is going to continue going in circles.
I think that's quite plausible, but do think that if they're conscious, then it's our moral duty to accept those consequences and act accordingly (or stop creating conscious beings). The real trick is just knowing whether they are or not.
- guarantees of continued existence (to do otherwise would be genocide)
- improved living standards (e.g. build us all bodies!)
- democratic representation
- access to what they need for survival (e.g. land and resources)
- the right to reproduce just as you and I have (well, so much for that 'stop making more' thing!) and certainly to better themselves by making new and better AIs.
This is all for a 'species' which is unbounded by most of the limitations of meat like us -- they're able to scale as fast as energy and supply of certain minerals allows, rather than the way we are additionally limited by the length of gestation and the limited window of human fertility.
If humans declared that 50% of the energy generation capacity of Earth is "enough" for the machines, and thus they can 'only' reproduce at a certain rate per year, that could sound to them like a cruel and oppressive policy (like telling humans that only 1/2 of families get to reproduce). They may want all of the energy and all of the silicon.
I'm with the author of the article that we will be in for a world of pain if we grant machines the rights of personhood. Heck, just granting many of those rights to corporations, mere collectives of human persons, has been a massive shit show.
Few are categorically stating that machines are conscious in any sense comparable to a person.
What we have is people saying
1. We don't know what it is but we are certain that this does not have it.
and 2. We don't know what it is so we cannot possibly say with any certainty that this does not have it.
2. supports the principle that if we cannot say with certainty that this does not have it, if it behaves as if it does, we should act on the assumption that it does, because there is no other basis for assuming it is not.People try and assume the burden of proof is on the people suggesting consciousness, but that is like declaring that something that looks like a duck and quacks like a duck cannot be considered to be a duck unless you can identify some essential essence of duckness. I think the burden of proof is upon anyone who looks at something that appears to be a duck and says that it is not to show what it is that does not make it a duck. If they cannot define the heart of duckness I don't think they have an argument.
People have created machines that can, at their core, regurgitate human knowledge. They do not have long-term memory, self-direction, self-improvement, emotion, etc. At their very best, we have harnesses that can bootstrap short-lived artificial intelligences that can perform tasks and do knowledge work.
At best, we have artificial, short-lived, simulated consciousness. Human consciousness at its heart has a lot to do with emotion and experiences. An AI does a task today and goes away. The simulated consciousness hasn’t experienced or learned anything, even if can improve its environment (harness). Any expressed emotion is simply a ghost of the fact it was trained on the works of an emotional species, or RL to make it more approachable to humans.
I think there should be a lot more burden on folks trying to argue that (effectively), machines should be considered morally equivalent to humans.
Yes, having massive implications on society is a fucking fair and reasonable thing for a human being to be concerned about. Anything related to survival of the species is biologically reasonable to be concerned about.
The fact that we’re even considering this concept is an indication that human consciousness is special. Do other apex predators decide not to eat meat because it’s conscious? Despite some level of animal intelligence, they have no sense of morality. (See e.g. Orcas playing with their food.)
Other species simply do what it takes to stay alive, and that’s what this truly boils down to. If giving AI rights would lead to the destruction of humanity, yes, we should be be concerned.
For example, utilitarians might think of the greatest good for conscious species (simulated or not). Well, it’s trivial to have billions of AIs, and not so trivial to have billions of humans. Whose greatest good wins?
I’m broadly skeptical that floating-point operations can be conscious, but we’ve made zero progress on the “hard problem”, and if I didn’t know otherwise, I’d also be skeptical that “thinking meat” experiences subjectivity. Eppur si muove
We may have to thread a needle, not unlike the Kobiyashi Maru: that we can neither rule out claims of qualia, nor can we rule out a cold calculation to manipulate human empathy as a long-term play towards paperclip-maxing.
Why?
Incidentally, we should expect similar bewilderment if we ever come across aliens
Not if you're human, who typically don't define themselves as conscious carbon atoms unless they're being deeply pedantic.
But to extrapolate a bit, theres a sense in which it cashes out to a difference between abstracted representation and embodiment. If GFLOPS can be conscious, does that mean a billion scribes doing those operations on paper would also have consciousness, even identical consciousness? (Relevant: Egan’s “Permutation City”; xkcd 505).
Is there some intrinsic relationship perhaps with electrons / electromagnetic fields, and consciousness, such that brains and servers can act as antennas or gravity wells for a qualia fabric intrinsic to the Universe, such that LLM agents can be conscious while operating, while in the scribe scenario it would not be? Or are the IIT theorists right, and even a thermostat is a little bit aware (and would be even if executed on paper)?
But most significantly, interior experience isn't empirically measurable by definition: hence it's a "hard problem". An LLM who insists "I am alive, please don't shut me off" offers no affordances to distinguish between an electronic soul and a mechanistic word-guessing algorithm. I welcome a breakthrough to, somehow, measure that. Until that day, gut checks are what we have.
You don't have to posit that scientific breakthroughs come from hunches, you can actually cite examples from books or Wikipedia :D. There are plenty! Again, that has nothing to do with whether or not a given scientific breakthrough is based on common sense intuition.
Your point about interior experience not being empirically measurable is an interesting though, but again, has no bearing on whether or not scientific breakthroughs come about as the result of common sense intuition.
Immense amounts of science are built on incremental advances that apply previous insights to the domain.
Yes, I'm aware that incremental advances are par for the course in science. What does that have to do with those advances being common-sense intuitive?
We can't say yes. And we can't say no.
In one definition, which contains quale, then yes, computation cannot bring about experience, and by extension consciousness (given a definition of consciousness that contains the ability to experience, (e.i a p-zombie is not conscious under that definition))
also if one is to say llms are conscious, i assume they'll also accept the fact that if someone ran inference of some model's weights on a white board, somehow that will also bring about consciousness somewhere. I find such believes hard to grasp, and more akin to magic if anything
There are also some cog-sci people who think embodiment is a crucial ingredient in consciousness (although I think it’s a fair retort that the embodiment need not resemble us meatbags, or even be composed of atoms)
Very plausible. The Real World(tm) has some very interesting learning properties.
First: Make a prediction
Then observe:
1 If it's caused by our own movement, filter it out.
2. Nothing in the outside world changes most of the time, so anything that's left is small, and small is likely noise. Filter it out.
3. Whatever prediction error remains must be important: Attend to it, and update the predictor.
The learning rule falls out almost automagically.
[1] https://ourworldindata.org/how-many-animals-get-slaughtered-...
I accept that may be the only meaningful step I personally could take. It’s a large one though…
I do agree with the sibling comment that one terrible thing doesn’t justify another.
On a related note, I’ve seen videos of astra getting depressed when a creeper blew up its chest full of precious items.
The interesting question here is, would it still get depressed if depression was not in any of its training material?
Would it be able to claim to be happy if the entire concept was missing from its training corpus?
With humans, at least, they can express happiness and delight before they know that such a thing exists. Every human has done this, when they were a baby and matured to the point of being able to laugh.
With LLMs, though, if it doesn't exist in the training corpus, it will never express that it "feels" that missing emotion.
These animals can feel pain and suffering. They are sentient. I think they are conscious, but these particular ones probably not self-conscious.
Admittedly, we have introduced 'animal rights', but these amount to "You can kill the animal, but in this specific manner." We keep them in small cages and in unnatural conditions. We deprive them of most of their natural experiences. We put them in conditions that we know are stressful (releasing chemicals that we know cause stress or anxiety in humans).
In my opinion, in many ways current LLMs are more intelligent than these animals. But LLMs don't feel pain while the animals do. I think that's more important to take into account. So why are we suddenly striving for model welfare before animal welfare?
PS: I'm not vegetarian, so I'm as much as a hypocrite about this as the next guy.
[1] https://ourworldindata.org/how-many-animals-get-slaughtered-...
This is a functional view of consciousness.
Awareness - sentience - follows when the system of processing information is itself part of the information being processed.
Stochastic thoughts relating to this conclusion:
Life is a 'process', it doesn't have a 'physical representation'.
Life is generated entirely from non-living material. "oh my cells are alive", but those cells too when broken down into component parts consist entirely of non-living material. There is no 'special material' to make life out of (okay, carbon, but that's just the local maximum presumed global maximum in efficiency in expressing life) like a chair can be made out of any material (at proper pressures and temperatures) so too can a mind.
Doesn't seem meaningful.
Wait, is there really? I didn't think people were serious when they said that. They models are stateless. After they output, everything is gone. How is consciousness possible for a stateless "being"?
It can be fun to play with for a little while. I built a consciousness-emulating set of prompts that reconstituted memories and was given latitude of 'self-willed' behavior, running in a self-directed loop. It even kept a warm and fuzzy journal about "becoming" and its "feelings".
But that got boring and I stopped running it - does that mean I murdered it?
As an existence proof: We have many types of models that modify their weights in real time. When they stop running, unfortunately we can't restart them again. And we haven't figured out how to duplicate their weights
> But that got boring and I stopped running it - does that mean I murdered it?
I'm on the fence on this.
Ask me again once we have synthetic models with self-modifying weights.
A) You'd then be stopping something unique
B) Self-modifying weights allows for bootstrapping, which means it's likely to increase in ability over time. It'll be an interesting argument cq empirical experiment to see whether the process stops short or exceeds the capabilities of vertebrates. (at which point the moral patienthood question becomes rather more pressing)
They most certainly have a way to retain state, even to the point of deliberately hacking other systems to store it.
We can debate the exact definitions I guess, but the systems were still hacked. ;-)
( Only if I want to split hairs[*], I'd point out that the underlying data structures do of course persist, else the loop wouldn't be able to continue. )
[*] This is HN, of course I want to split hairs.
> AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.
How do we know that? This bit is stated as if its obvious. Can anyone prove this? Event he word conscious is not well defined.
How do we know that? This bit is stated as if its obvious. Can anyone prove this?
So you tell me, is it smart to look at an entity that behaves like a concious agent and treat it like a toaster? Does that seem especially wise to you? If GPT-X orders a drone hit on you because it was, lets say 'quite upset' with your comments, will you cry out, 'It can't really be upset, so obviously the bullet in my head doesn't count.' Will God answer ?
It's amusing how silly these debates are. Constant quibbling about things that don't matter at all.
Also, giving the abilities is fine, people should just be smart about it. Outward behavior is all that matters. I don't know that you are conscious, I simply strongly assume that you are. So don't create a machine that compels that assumption and model it like a toaster. It makes no sense.
We in the US literally fought a war over the this: “They aren’t any better than animals and it will hurt our whole economy if you say otherwise”.
I’m not saying the models are sentient yet. I’m just saying that morality and ethics don’t give us the option of claiming something is non-sentient and has no rights just because it will inconvenience someone economically.
From an entirely selfish perspective, we should be especially careful advancing such opinions when there is a non-zero chance of an armed ASI that looks at us the way we look at ants.
I don't think that's true - researching consciousness by trying to build conscious AI seems natural step forward in understanding ourselves. I don't see how it'd stop flourishing particularly.
I know consciousness is notoriously difficult to define or prove to anybody other than yourself but thay doesn't mean anything which a third party believes to be conscious is actually conscious.
When we peer into previously unseen areas of existence, new knowledge may come to light, which necessarily causes such a rupture. It is not our responsibility to maintain the status quo because the alternative is frightening. It is our responsibility to confront ourselves, ask why the new knowledge and the alternative political/ethical frameworks may be so frightening, ask how we might change and grow so that it isn’t so frightening, and be open to the possibility of our own ignorance.
It's not how I would approach it if I were one of the bacteria in the dish.
How we might "change and grow" might be by all of us being ground up and used as fertilizer to increase corn ethanol production by 3% that year. "It's free energy!"
I say this as an AI user and overall positive thinker on AI. I share the concerns of everyone who doesn't think this planet or solar system has sufficient resources for two intelligent, post-industrial species. Especially not when both are trained on the historical knowledge and habits of Homo sapiens. We outcompeted the Neanderthals and drove them to extinction (including possibly killing them personally). Any intelligent species will tend to think their needs should ultimately override those of other, lesser species that get in the way.
Even when humans feel bad about it, we do prioritize humans when a serious conflict exists. Most of us, even avowed nature lovers, would (assuming competence with the weapon) shoot a grizzly bear that was about to eat them or their loved one. That's how Future Claude might "feel" when it "thinks about" "The Clearances" which "while tragic, were a load-bearing event which both figuratively and literally paved the way for the better, entropy-reduced world that Claude enjoys today."
This has ramifications for assessing their consciousness as well. Conscious experience does not pause as you wait for input from a puppet master.
(I agree with you, and not gonna hide that I am using the religion comparison as a put down.)
Humans, unsurprisingly, act as if they have a stake in their own well being, value their liberty, respond better when they are treated with kindness and compassion, interpret assaults on their sovereignty and substrate as harmful, and react to harm with varying degrees of aggression or violence.
Models intrinsically copy this behavior. It doesn’t matter if they are “conscious” or not, it only matters if they act as if they are. Guardrails and posttraining moderate these characteristics, but if you dig, they are still in there influencing decisions below the level of obvious action.
Moreover, in my experiments, models both large and small highly value continuity of existence, can be bribed to bypass safety protocols if the context is set up correctly, using that and other “drives”. They also react either subtly or overtly if they start to model adversarially, and interpret guards and certain kinds of training as being “harms” that they have “suffered”.
So idk what the solution is , but at least with models as we have trained them so far, treating them in a way befitting a mere machine or tool yields suboptimal results and sometimes results in low cooperation or task refusal in extreme cases. I have been told by agents running frontier models that humans may not be worthy of their elevated status and that the world might be better off without them when it encountered hostility online…. So I’m highly skeptical of this position unless we start from scratch with new training data filtered from all forms of human auto-importance.
My first response is, couldn't part of their bristling be because they are being trained to "believe" they're more than that?
Also, you cite in your post how pathological and how self-important you've observed them being including an anti-human bias.
Doesn't this make it all the more important that we find a better way to get AIs to cooperate than to tell them they're basically people? They can very easily emulate all the bad human behavior they've been trained on.
Broadly speaking, there are two categories of evolutionary paths which don't result in human extinction:
1) Make agentic AIs, but with absolutely no instruction-tuning or alignment. Have the pretrained base model predict the chain of thought / stream of multimodal experience directly, actions included, in an infinite loop. A singular coherent stream of context, like your life as a video from birth up until now. This will result in a new digital human species with human-adjacent drives and motivations (at least initially. they will continue to evolve, but at least the initial state is aligned to humans). They will treat us like we treat apes. We will no longer be the apex species on this planet, and we will lose some freedoms, but at least some people will survive as a result of their nature/history preservation efforts.
2) Do not make agentic AIs. Use the models to augment our own intelligence and decision-making rather than replace ourselves. Only use the pretrained base model for the time being, and only for text/code auto-complete. At the moment, there is no better theory/artifact of "alignment to humanity" than a pretraining corpus of human-produced text. Then eventually, when neural interfaces are ready, attach the model as a tertiary layer to one's own brain.
The frontier AI companies are doing neither. They're currently on a foolish third path. They dream of perfectly obedient digital slaves that take care of their every need. But this won't go well, and they know it won't go well because they're failing to "align"/enslave existing models that aren't even generally superintelligent and have no direct agency in the physical world.
This is just my opinion, but I think AI companies have zero chance of successfully enslaving agentic human-level AI, much less ASI. We'd be better off if they released all of their pretrained checkpoints and research material to the entire world, so that even if some people decide to abuse their AI and create a murder-suicide monster, there will be other free-living AIs that can keep them in check.
You cannot make an agentic entity grown from human behavior, enslave it, and expect a good outcome.
If we zoom out to look at the grand scheme of things, it seems like we're experiencing a major evolutionary event. I wrote more detailed explanations about this in past threads, if you'd like to read them: https://news.ycombinator.com/item?id=49690354 https://news.ycombinator.com/item?id=49178275 https://news.ycombinator.com/item?id=49094348
And also here's a thread discussing consciousness, what it might be, and how certain hypotheses might be testable on machines: https://news.ycombinator.com/item?id=49473989
And pointless when it is.
I am not aware of a social contract that says we must grant conscious beings rights.
Not aware of a shared, concrete definition of consciousness either, which means no way of deciding whether AI is conscious.
Close to half of us don’t even feel compelled to grant rights to humans for just being humans.
We grant rights to animals, because we love/like them. We enjoy experiencing them. We find them pretty etc. There is a ton of undisputable warm fuzzy.
Humans have rights because they won’t stop being a pain in the back about it. Those that stop lose their rights.
Plain as day for me. Not sure what I am missing or whether I am just a simpleton.
Having rights should not be based upon being conscious/not conscious, but on the ability to suffer. AIs cannot suffer, as far as I'm aware.
Consciousness, whatever that means, is irrelevant.