AL
Alignment Forum
0 stories this week · 4 topics · alignmentforum.org
AI-related Community; Focus: AI Safety Research
Latest from Alignment Forum
- AI ResearchThe Alignment Journal: Organization, Personnel, and Scope9 days ago
- AI ResearchTraining a Misaligned Reward Seeker10 days ago
- AI ResearchValue generalisation Theory of Change: putting it into practice10 days ago
- AI ResearchValue generalisation Theory of Change: the theory behind the approach13 days ago
- AI ResearchBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident15 days ago
- AI ResearchDebate Training Reduces Reward Hacking in RLAIF22 days ago
- AI ResearchAI swarms are starting to pose indirect takeover risk29 days ago
- AI ResearchFour LLM loss functions → four flavors of LLM misalignmentAug 10, 2026
- AI ResearchWhy do models task game?Aug 6, 2026
- AI ResearchReturning to ARCAug 4, 2026
- AI ResearchAGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)Jul 31, 2026
- AI ResearchThe AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)Jul 31, 2026
- AI ResearchThousand-dimensional structureJul 30, 2026
- AI ResearchValue Generalisation 3: Pre-aligned AIsJul 29, 2026
- AI ResearchValue Generalisation 2: The Missing Hole in AIs’ abilitiesJul 29, 2026
- AI ResearchValue Generalisation 1: a Research and Deployment ProgramJul 29, 2026
- AI ResearchResearch directions in condensation: varieties of objectivityJul 28, 2026
- AI ResearchRL & search is a terrifying way to build AGI (an FAQ)Jul 27, 2026
