Gunjan Dhanuka — Learning Notes

Home

❯

learn

❯

rl

rl

10 items under this folder.

  • Sep 02, 2026

    01 · MDPs & the RL Objective

    • reinforcement-learning
    • mdp
    • value-functions
    • bellman
    • rl-for-llms
  • Sep 02, 2026

    02 · Policy Gradients

    • reinforcement-learning
    • policy-gradient
    • reinforce
    • baselines
    • variance-reduction
    • rlhf
  • Sep 02, 2026

    03 · Actor-Critic & GAE

    • reinforcement-learning
    • actor-critic
    • gae
    • advantage-estimation
    • variance-reduction
    • ppo
    • rlhf
  • Sep 02, 2026

    04 · PPO

    • rl
    • ppo
    • trpo
    • trust-region
    • policy-gradient
    • rlhf
    • importance-sampling
  • Sep 02, 2026

    05 · RL on Token Sequences

    • rl
    • rlhf
    • llm
    • token-level-mdp
    • kl-regularization
    • ppo
    • credit-assignment
  • Sep 02, 2026

    06 · The RLHF Pipeline

    • rlhf
    • reward-model
    • bradley-terry
    • instructgpt
    • ppo
    • preference-learning
    • alignment
    • reward-overoptimization
  • Sep 02, 2026

    07 · DPO & RL-Free Preference Optimization

    • dpo
    • preference-optimization
    • rlhf
    • bradley-terry
    • kl-regularization
    • ipo
    • kto
    • orpo
    • simpo
    • cpo
    • alignment
    • likelihood-displacement
  • Sep 02, 2026

    08 · Scaling RLHF & Alternatives

    • rlhf
    • rlaif
    • constitutional-ai
    • reward-model
    • reward-hacking
    • reward-overoptimization
    • llm-as-a-judge
    • rejection-sampling
    • best-of-n
    • rest
    • warm
    • alignment
  • Sep 02, 2026

    09 · RL for Reasoning (RLVR, GRPO, R1)

    • rl
    • rlvr
    • grpo
    • deepseek-r1
    • reasoning
    • verifiable-rewards
    • chain-of-thought
    • process-reward-model
    • dapo
    • dr-grpo
    • o1
    • test-time-compute
  • Sep 02, 2026

    10 · Frontier RL — Infra & Open Problems

    • rl-infra
    • async-rl
    • off-policy
    • rlvr
    • agentic-rl
    • entropy-collapse
    • reward-overoptimization
    • post-training
    • scaling-laws
    • frontier

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source