Algorithms - optimal control
State-Action Value function:
Note
Qp(s,a) = E [ R(s,a) + g Vp(s’)]
p is deterministic.
השקופית הקודמת
השקופית הבאה
חזור אל השקופית הראשונה
הצג גירסה גרפית