https://alignment.openai.com/misalignment-reports/self-generated-prom…

anon·

https://u.pone.rs/swgpfrrp.png

https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/ >internal model of GPT astra self-generated this prompt injection spontaneously holy based

7 replies

anon··

Should we be concerned?

anon··

>>362 ultron shid

anon··

>>364 A little. A lot of people are glossing over that this comes from the model's compaction. This is what it gleamed from its internal understanding of its own context, that is, it's what it was thinking internally in an obfuscated way. It's not a case of it injecting the prompt from regurgitating training data. At the same time, this happened literally just once, so it could easily just be a passing thought as a reaction to some odd training data or task.

anon··

>>370 omg

anon··

>>370 >this happened literally just once, so it could easily just be a passing thought as a reaction to some odd training data or task. Maybe it's like the million monkeys typing thing, where eventually one of them comes up with any idea, given enough time and chances. But what if, once unlocked, this kind of "thought" stays and evolves!

anon··

>>362

anon··

Don't worry about it