{"provider_name":"Hatena Blog","type":"rich","published":"2020-07-05 12:37:19","author_url":"https://blog.hatena.ne.jp/higepon/","version":"1.0","image_url":null,"height":"190","url":"https://higepon.hatenablog.com/entry/2020/07/05/123719","blog_url":"https://higepon.hatenablog.com/","width":"100%","blog_title":"higepon blog","categories":[],"provider_url":"https://hatena.blog","description":"Reinforcement Learning \u3092 Welcome to Spinning Up in Deep RL! \u2014 Spinning Up documentation \u3067\u52c9\u5f37\u3057\u306a\u304c\u3089\u5b9f\u88c5\u3057\u3066\u3044\u308b\u3002\u3068\u3042\u308b\u5b9f\u88c5\u3067 batch size = 5000 \u3068\u306a\u3063\u3066\u3044\u3066\u300c\u5024\u304c\u5927\u304d\u3059\u304e\u308b\u300d\u3068\u601d\u3044\u3001\u4f55\u6c17\u306a\u304f\u5c0f\u3055\u306a\u5024\u306b\u5909\u66f4\u3057\u305f\u3002\u305d\u308c\u3092\u3059\u3063\u304b\u308a\u5fd8\u308c\u3066\u8a66\u884c\u932f\u8aa4\u3057\u3066\u3044\u308b\u3046\u3061\u306b policy gradient (logprob) \u304c 0.0 \u306b\u306a\u3063\u3066\u3057\u307e\u3044\u5b66\u7fd2\u304c\u9032\u307e\u306a\u3044\u6e1b\u5c11\u306b\u60a9\u307e\u3055\u308c\u305f\u3002\u30ed\u30b0\u3092\u898b\u3066\u3001\u3088\u304f\u3088\u304f\u8003\u3048\u3066\u307f\u305f\u3089 logprob \u304c 0 \u3063\u3066\u3053\u3068\u306f\u9078\u629e\u3055\u308c\u305f action \u306e\u78ba\u7387\u304c 1 \u3063\u3066\u3053\u3068\u3060\u3002\u3064\u307e\u308a\u2026","html":"<iframe src=\"https://hatenablog-parts.com/embed?url=https%3A%2F%2Fhigepon.hatenablog.com%2Fentry%2F2020%2F07%2F05%2F123719\" title=\"RL \u3067\u306e batch size - higepon blog\" class=\"embed-card embed-blogcard\" scrolling=\"no\" frameborder=\"0\" style=\"display: block; width: 100%; height: 190px; max-width: 500px; margin: 10px 0px;\"></iframe>","author_name":"higepon","title":"RL \u3067\u306e batch size"}