Gunjan Dhanuka — Learning Notes

ppo

10 items with this tag.

  • Sep 02, 2026

    Actor-Critic

    • reinforcement-learning
    • actor-critic
    • critic
    • variance-reduction
    • ppo
    • rl-for-llms
  • Sep 02, 2026

    Generalized Advantage Estimation (GAE)

    • reinforcement-learning
    • gae
    • advantage-estimation
    • td-learning
    • variance-reduction
    • ppo
    • rl-for-llms
  • Sep 02, 2026

    KL Regularization in RLHF

    • rlhf
    • kl-divergence
    • reward-hacking
    • reference-policy
    • ppo
    • rl-for-llms
  • Sep 02, 2026

    PPO (Proximal Policy Optimization)

    • rl
    • ppo
    • trust-region
    • policy-gradient
    • rlhf
    • importance-sampling
    • rl-for-llms
  • Sep 02, 2026

    RLHF (Reinforcement Learning from Human Feedback)

    • rlhf
    • alignment
    • reward-model
    • ppo
    • instructgpt
    • llm-training
    • rl-for-llms
  • Sep 02, 2026

    Trust Region

    • rl
    • trust-region
    • trpo
    • ppo
    • natural-gradient
    • kl-divergence
    • rl-for-llms
  • Sep 02, 2026

    03 · Actor-Critic & GAE

    • reinforcement-learning
    • actor-critic
    • gae
    • advantage-estimation
    • variance-reduction
    • ppo
    • rlhf
  • Sep 02, 2026

    04 · PPO

    • rl
    • ppo
    • trpo
    • trust-region
    • policy-gradient
    • rlhf
    • importance-sampling
  • Sep 02, 2026

    05 · RL on Token Sequences

    • rl
    • rlhf
    • llm
    • token-level-mdp
    • kl-regularization
    • ppo
    • credit-assignment
  • Sep 02, 2026

    06 · The RLHF Pipeline

    • rlhf
    • reward-model
    • bradley-terry
    • instructgpt
    • ppo
    • preference-learning
    • alignment
    • reward-overoptimization

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source