Juliet E McKenna on Nostr: I have read a technical but comprehensible explanation of the OpenAI 'rogue software' ...
I have read a technical but comprehensible explanation of the OpenAI 'rogue software' event - with close attention as the detail relates to IT at the outer limit of my personal understanding.
Summing up is clear
"If you optimize a model to find exploits, you should expect it to find them — and prepare for that. OpenAI did not. They built a model, took the safeguards off, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making."
https://mail.cyberneticforests.com/models-dont-go-rogue/Published at
2026-09-01 07:32:53 UTCEvent JSON
{
"id": "339f1bac3d889c2e2662d7a22c9c7fa4e040306a827dd9f9e2dcd8d907cf4957",
"pubkey": "363c3eafe9b5fdde2e7ceaa52caefbecd6de3fa7b87fab117d209a0698c070a2",
"created_at": 1788247973,
"kind": 1,
"tags": [
[
"proxy",
"https://wandering.shop/@JulietEMcKenna/117194619177097687",
"web"
],
[
"proxy",
"https://wandering.shop/users/JulietEMcKenna/statuses/117194619177097687",
"activitypub"
],
[
"L",
"pink.momostr"
],
[
"l",
"pink.momostr.activitypub:https://wandering.shop/users/JulietEMcKenna/statuses/117194619177097687",
"pink.momostr"
],
[
"-"
]
],
"content": "I have read a technical but comprehensible explanation of the OpenAI 'rogue software' event - with close attention as the detail relates to IT at the outer limit of my personal understanding.\n\nSumming up is clear\n\n\"If you optimize a model to find exploits, you should expect it to find them — and prepare for that. OpenAI did not. They built a model, took the safeguards off, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making.\"\n\nhttps://mail.cyberneticforests.com/models-dont-go-rogue/",
"sig": "9e3df8772166686d812bb6506d13f37425b5a1eb3e6bfd1a9b864e12b4f243419dc2d5ad2e7ab2a667e27faae6dd671013d889c413335cef61ff216a88318b9f"
}