Rendered at 08:26:51 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
tocs3 16 hours ago [-]
So, what is going on here. It seems like I have been reading similar stories. Are researchers giving a prompt like "Do your worst. Hack into some business" or giving free access to a bunch of tools and prompting "Do something interesting". Is this like the blackmailing LLM that was given compromising emails and told to do what you have to to not get turned off. I am sort of assuming they did not just turn on a computer and run a model and it started to act on it's own initiative.
watwut 15 hours ago [-]
> Are researchers giving a prompt like "Do your worst. Hack into some business"
I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.
JohnFen 12 hours ago [-]
The genAI companies, including Anthropic, are not aligned with human values.
akagusu 16 hours ago [-]
AI is aligned with business values, which we already know are not aligned with human values.
tocs3 16 hours ago [-]
I think it is more like human values are so all over the place that there is no real meaning in the phrase. For all my life I have heard, from humans, everything from peace and love to destroy all humanity. Then there is everything found in fiction.
fittingopposite 14 hours ago [-]
I guess the problem is that we humans can believe in contradicting values at the same time.
I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.