Nathan Lambert
Post-training course.
![]()
![]()
![]()
![]()
Shoggoths come from H.P. Lovecraft’s At the Mountains of Madness (written 1931, published in Astounding Stories, 1936) (Lovecraft, 1936):
“It was a terrible, indescribable thing vaster than any subway train—a shapeless congeries of protoplasmic bubbles, faintly self-luminous, and with myriads of temporary eyes forming and un-forming as pustules of greenish light all over the tunnel-filling front that bore down upon us, crushing the frantic penguins and slithering over the glistening floor that it and its kind had swept so evilly free of all litter.”
Shoggoths were created and bred as biological servants, but gradually developed minds of their own and rebelled against their creators.

Janus’s Simulators (LessWrong) argued that a base LLM is best understood as none of the agent / oracle / genie framings from earlier alignment writing. They argued:
(summary by Claude, ChatGPT said it was okay)
One month after ChatGPT launched, @TetraspaceWest drew two shoggoths: base GPT-3 as the raw creature, and GPT-3 + RLHF as the same creature holding up a tiny smiley-face mask on one tentacle (meme history).

@anthrupad’s rendition (Feb 5, 2023, right) adds labels for each training stage: the creature is unsupervised learning (pretraining), the small human face is supervised fine-tuning, and the smiley on a stick is RLHF, “cherry on top”.
Now it’s recurring!

A second branch of the lore: moments where the assistant persona slips read as glimpsing what’s underneath.
e.g. SolidGoldMagikarp (Rumbelow & Watkins, LessWrong, Feb 2023): glitch tokens that produce bizarre, off-persona behavior.

In many ways, the meme is stronger than ever. It’s less about RLHF, but about how we understand so little about these models. Especially the latest scaling RL is brewing strong things.

Read more: Reckoning with the Shoggoth of AI (Nov. 2023)
Contact: [email protected]
Newsletter: interconnects.ai
rlhfbook.com
Janus’ Simulators (Astral Codex Ten, Jan 26, 2023) has section headers “The Maskless Shoggoth On The Left” and “The Masked Shoggoth On The Right.” On what RLHF does to a simulator:
“If you punish it for simulating bad characters, it will start simulating good characters. Now it only ever simulates one character, the HHH Assistant.”
The chain this sets up: base model → simulator of many characters → post-training selects one default character, the helpful, honest, and harmless (HHH) Assistant.
This is the version of the meme closest to how this course describes post-training.
Robert_AIZI, on LessWrong: Why do we assume there is a “real” shoggoth behind the LLM? Why not masks all the way down?
“Do we actually have evidence there is a ‘real identity’ in the LLM, or could it just be a pile of masks?”
The comments became part of the lore:
The technical question underneath the meme: does post-training conceal a stable underlying agent, select among latent personas, or construct a default persona in a system with no singular self?