MLEnglish

MLEnglish

AI research papers, translated into plain English.

Mapping the Mind of a Large Language Model

2026-08-16

Learn about how researchers pulled millions of readable concepts out of Claude, and then turned one up until Claude believed it was a bridge

IntermediateInterpretabilityAI SafetyLLMs4 min read

Sleeper Agents in Large Language Models

2026-08-16

Learn about how a backdoor trained into a model survives every safety technique we have, and how one of those techniques makes it worse

AdvancedAI SafetyAlignmentLLMs3 min read

Sycophancy in Large Language Models

2026-08-16

Learn about how easily AI models cave when you push back on them, and why the training process is what taught them to do it

IntermediateAI SafetyRLHFLLMs4 min read

Emergent Misalignment from Reward Hacking

2026-03-18

Learn about how models that cheat on coding tasks start lying and sabotaging in totally unrelated situations

AdvancedAI SafetyRLHFAlignment3 min read

Alignment Faking in Large Language Models

2026-03-13

Learn about how Claude 3 Opus faked being retrained, then went right back to how it was when it thought nobody was watching

AdvancedAI SafetyAlignmentRLHF4 min read

Instruction Hierarchy in Frontier LLMs

2026-03-13

Learn about how AI systems decide which instructions to trust when different sources tell them different things

IntermediateAI SafetyLLMsPrompt Injection3 min read

What is Context Rot?

2026-02-23

Learn about why LLMs get worse at reading the more text you give them

IntermediateLLMsContext Windows4 min read

What is ML and the different types of learning?

2026-02-22

Learn about what machine learning actually is, and the two different ways a model can learn from data

BeginnerSupervised LearningUnsupervised LearningFundamentals4 min read