When a baseline makes an agent too sure of itself
Why a seemingly small choice in policy gradient learning can change how quickly an agent commits—and whether it keeps exploring.
Read the postA home for longer notes and ideas.
Why a seemingly small choice in policy gradient learning can change how quickly an agent commits—and whether it keeps exploring.
Read the post