Anthropic suggests slowing AI research until we can align it with human goals
“How the alignment problem gets solved — or not — in this future is something we are least certain about,” they wrote. Advanced, self-improving models could follow our needs and wants — or, they warned, “The rare occurrences of misalignment present in today’s models could compound as the models build their successors, growing more frequent…