News

When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning
News

When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning

The researchers at Center for Human-Compatible AI (CHAI) at the University of California, Berkeley, has embarked on a study that brings to light the nuanced challenges encountered when AI systems learn from human feedback, especially under conditions of partial observability.