TLDRocket
Sign in

Agent Evaluation

21 summarised stories about Agent Evaluation, each linking back to the original source. Browse all topics →

Tuesday, 3 December 2019

Procgen Benchmark

OpenAI Blog 6 years ago

OpenAI released Procgen Benchmark, a set of 16 procedurally-generated environments designed to measure how quickly reinforcement learning agents learn generalizable skills. The benchmark includes 16 distinct environments that generate new levels procedurally to test agent generalization across novel scenarios. Researchers can now use this standardized tool to compare RL agent performance on learning efficiency and transferability rather than memorization of fixed levels.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.