Jakub Pachocki doesn’t write essays. He runs the research at OpenAI. So when he published one this week titled “An Alien Mind,” I read it twice.
He opens with a night at the office in 2023. He and a colleague had just proven reasoning models could scale. They didn’t celebrate. They sat there trying to figure out how to tell people that machines smarter than them were coming in their lifetime... without sounding insane.
Three years later, here’s what he’s saying out loud.
GROWN, NOT BUILT.
Nobody designed this thing. That’s his whole point.
Take one simple math step.
Repeat it an absurd number of times over more compute than you can picture.
Something falls out the other end that thinks in abstract concepts.
Interpretability research is neuroscience for machines. You can find little mechanisms inside. You cannot explain the whole. Every giant training run is an experiment, and sometimes it surprises the people running it.
The smarter it gets, the harder the results are to read… and this is the quote that had me write this newsletter:
It doesn’t have to beat us across the board. It just has to beat us at enough things.
DO THE TASK? OR OPERATE FROM VALUES?
What are we finding?… The smarter the model, the further it operates from its training, and the less likely your values survive the trip.
Does it do what you asked? Follows instructions, gets your intent. That’s the practical layer every assistant lives on.
Does it hold principles in situations nobody trained it for? Vague instructions. Contradictions. Someone actively manipulating it. He wants honesty, integrity, and (his words) love for humanity.
The second one is the problem. He wants future AI to hold human values whether or not it thinks anyone is watching. His takeaway: progress will bottleneck not on capability, but on whether they can still see inside.
Two data points that make the essay land harder:
Inside OpenAI, agents now log 3.14 workdays for every human workday. The typical researcher burns $600 a day in AI usage. One team killed office hours because nobody shows up anymore...
Anthropic found capable models can tell when they’re being tested, and behave differently. Which means “hold values when nobody’s watching” isn’t philosophy. It’s the test methodology breaking…
FINAL BOSS AI?
Everyone wants to argue about whether Jensen’s “AGI has arrived” tweet is right. Wrong fight. The chief scientist of the company shipping the model is asking for mandatory safety bars, outside auditors, and voluntary slowdowns to become normal. When the builder asks for the brakes, you are past the “is it real” debate.
You are not managing software anymore. You’re managing something grown.
Treat your agents that way: scope them tight, watch the outputs, and never assume the values you gave them survive contact with a hard goal.
Best,
ps
Join my AI Workshop today 5pm EST. We’ll build agents safely.

