AI research papers, translated into plain English.
Learn about how researchers pulled millions of readable concepts out of Claude, and then turned one up until Claude believed it was a bridge
Learn about how a backdoor trained into a model survives every safety technique we have, and how one of those techniques makes it worse
Learn about how easily AI models cave when you push back on them, and why the training process is what taught them to do it
Learn about how models that cheat on coding tasks start lying and sabotaging in totally unrelated situations
Learn about how Claude 3 Opus faked being retrained, then went right back to how it was when it thought nobody was watching
Learn about how AI systems decide which instructions to trust when different sources tell them different things
Learn about why LLMs get worse at reading the more text you give them
Learn about what machine learning actually is, and the two different ways a model can learn from data