Study: Humans approve AI agent commands, miss 1 in 3 harmful actions
In a large-scale simulation with 40,000 game runs, human overseers approved AI agent commands while failing to notice about one-third of the harmful ones. The findings highlight a major challenge for safely supervising autonomous AI systems in real-world tasks.
Sources (1)
technology