Gunjan Dhanuka — Learning Notes

kl-regularization

4 items with this tag.

  • Sep 02, 2026

    DPO (Direct Preference Optimization)

    • dpo
    • preference-optimization
    • rlhf
    • bradley-terry
    • kl-regularization
    • alignment
    • rl-for-llms
  • Sep 02, 2026

    Reward Overoptimization (Goodhart)

    • reward-overoptimization
    • reward-hacking
    • goodhart
    • scaling-laws
    • warm
    • kl-regularization
    • rl-for-llms
  • Sep 02, 2026

    05 · RL on Token Sequences

    • rl
    • rlhf
    • llm
    • token-level-mdp
    • kl-regularization
    • ppo
    • credit-assignment
  • Sep 02, 2026

    07 · DPO & RL-Free Preference Optimization

    • dpo
    • preference-optimization
    • rlhf
    • bradley-terry
    • kl-regularization
    • ipo
    • kto
    • orpo
    • simpo
    • cpo
    • alignment
    • likelihood-displacement

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source