Gunjan Dhanuka — Learning Notes

alignment

8 items with this tag.

  • Sep 02, 2026

    Constitutional AI (CAI)

    • constitutional-ai
    • rlaif
    • alignment
    • harmlessness
    • anthropic
    • rlhf
    • rl-for-llms
  • Sep 02, 2026

    DPO (Direct Preference Optimization)

    • dpo
    • preference-optimization
    • rlhf
    • bradley-terry
    • kl-regularization
    • alignment
    • rl-for-llms
  • Sep 02, 2026

    Preference Optimization (RL-free family)

    • preference-optimization
    • dpo
    • ipo
    • kto
    • orpo
    • simpo
    • cpo
    • alignment
    • rl-for-llms
  • Sep 02, 2026

    RLAIF (RL from AI Feedback)

    • rlaif
    • llm-as-a-judge
    • constitutional-ai
    • rlhf
    • reward-model
    • alignment
    • rl-for-llms
  • Sep 02, 2026

    RLHF (Reinforcement Learning from Human Feedback)

    • rlhf
    • alignment
    • reward-model
    • ppo
    • instructgpt
    • llm-training
    • rl-for-llms
  • Sep 02, 2026

    06 · The RLHF Pipeline

    • rlhf
    • reward-model
    • bradley-terry
    • instructgpt
    • ppo
    • preference-learning
    • alignment
    • reward-overoptimization
  • Sep 02, 2026

    07 · DPO & RL-Free Preference Optimization

    • dpo
    • preference-optimization
    • rlhf
    • bradley-terry
    • kl-regularization
    • ipo
    • kto
    • orpo
    • simpo
    • cpo
    • alignment
    • likelihood-displacement
  • Sep 02, 2026

    08 · Scaling RLHF & Alternatives

    • rlhf
    • rlaif
    • constitutional-ai
    • reward-model
    • reward-hacking
    • reward-overoptimization
    • llm-as-a-judge
    • rejection-sampling
    • best-of-n
    • rest
    • warm
    • alignment

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source