Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
ppo
10 items with this tag.
Sep 02, 2026
Actor-Critic
reinforcement-learning
actor-critic
critic
variance-reduction
ppo
rl-for-llms
Sep 02, 2026
Generalized Advantage Estimation (GAE)
reinforcement-learning
gae
advantage-estimation
td-learning
variance-reduction
ppo
rl-for-llms
Sep 02, 2026
KL Regularization in RLHF
rlhf
kl-divergence
reward-hacking
reference-policy
ppo
rl-for-llms
Sep 02, 2026
PPO (Proximal Policy Optimization)
rl
ppo
trust-region
policy-gradient
rlhf
importance-sampling
rl-for-llms
Sep 02, 2026
RLHF (Reinforcement Learning from Human Feedback)
rlhf
alignment
reward-model
ppo
instructgpt
llm-training
rl-for-llms
Sep 02, 2026
Trust Region
rl
trust-region
trpo
ppo
natural-gradient
kl-divergence
rl-for-llms
Sep 02, 2026
03 · Actor-Critic & GAE
reinforcement-learning
actor-critic
gae
advantage-estimation
variance-reduction
ppo
rlhf
Sep 02, 2026
04 · PPO
rl
ppo
trpo
trust-region
policy-gradient
rlhf
importance-sampling
Sep 02, 2026
05 · RL on Token Sequences
rl
rlhf
llm
token-level-mdp
kl-regularization
ppo
credit-assignment
Sep 02, 2026
06 · The RLHF Pipeline
rlhf
reward-model
bradley-terry
instructgpt
ppo
preference-learning
alignment
reward-overoptimization