Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
rlhf
18 items with this tag.
Sep 02, 2026
Bradley–Terry Model
bradley-terry
preference-learning
reward-model
dpo
rlhf
rl-for-llms
Sep 02, 2026
Constitutional AI (CAI)
constitutional-ai
rlaif
alignment
harmlessness
anthropic
rlhf
rl-for-llms
Sep 02, 2026
DPO (Direct Preference Optimization)
dpo
preference-optimization
rlhf
bradley-terry
kl-regularization
alignment
rl-for-llms
Sep 02, 2026
KL Regularization in RLHF
rlhf
kl-divergence
reward-hacking
reference-policy
ppo
rl-for-llms
Sep 02, 2026
Policy Gradient Theorem
reinforcement-learning
policy-gradient
score-function
rlhf
rl-for-llms
Sep 02, 2026
PPO (Proximal Policy Optimization)
rl
ppo
trust-region
policy-gradient
rlhf
importance-sampling
rl-for-llms
Sep 02, 2026
REINFORCE
reinforcement-learning
policy-gradient
reinforce
monte-carlo
rlhf
rl-for-llms
Sep 02, 2026
Reward Model
rlhf
reward-model
preference-learning
bradley-terry
reward-hacking
rl-for-llms
Sep 02, 2026
RLAIF (RL from AI Feedback)
rlaif
llm-as-a-judge
constitutional-ai
rlhf
reward-model
alignment
rl-for-llms
Sep 02, 2026
RLHF (Reinforcement Learning from Human Feedback)
rlhf
alignment
reward-model
ppo
instructgpt
llm-training
rl-for-llms
Sep 02, 2026
02 · Policy Gradients
reinforcement-learning
policy-gradient
reinforce
baselines
variance-reduction
rlhf
Sep 02, 2026
03 · Actor-Critic & GAE
reinforcement-learning
actor-critic
gae
advantage-estimation
variance-reduction
ppo
rlhf
Sep 02, 2026
04 · PPO
rl
ppo
trpo
trust-region
policy-gradient
rlhf
importance-sampling
Sep 02, 2026
05 · RL on Token Sequences
rl
rlhf
llm
token-level-mdp
kl-regularization
ppo
credit-assignment
Sep 02, 2026
06 · The RLHF Pipeline
rlhf
reward-model
bradley-terry
instructgpt
ppo
preference-learning
alignment
reward-overoptimization
Sep 02, 2026
07 · DPO & RL-Free Preference Optimization
dpo
preference-optimization
rlhf
bradley-terry
kl-regularization
ipo
kto
orpo
simpo
cpo
alignment
likelihood-displacement
Sep 02, 2026
08 · Scaling RLHF & Alternatives
rlhf
rlaif
constitutional-ai
reward-model
reward-hacking
reward-overoptimization
llm-as-a-judge
rejection-sampling
best-of-n
rest
warm
alignment
Sep 02, 2026
RL for Frontier LLMs — Curriculum
reinforcement-learning
rlhf
llm-training
curriculum