Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
rl-for-llms
26 items with this tag.
Sep 02, 2026
Actor-Critic
reinforcement-learning
actor-critic
critic
variance-reduction
ppo
rl-for-llms
Sep 02, 2026
Advantage Function
reinforcement-learning
advantage
variance-reduction
policy-gradient
rl-for-llms
Sep 02, 2026
Agentic RL (multi-turn, tool-use)
agentic-rl
multi-turn
tool-use
credit-assignment
environments
rlvr
echo-trap
rl-for-llms
Sep 02, 2026
Best-of-N Sampling (and RFT / ReST)
best-of-n
rejection-sampling
rft
rest
expert-iteration
test-time-compute
rl-for-llms
Sep 02, 2026
Bradley–Terry Model
bradley-terry
preference-learning
reward-model
dpo
rlhf
rl-for-llms
Sep 02, 2026
Chain-of-Thought RL (emergent reasoning)
chain-of-thought
reasoning
rlvr
grpo
o1
deepseek-r1
test-time-compute
rl-for-llms
Sep 02, 2026
Constitutional AI (CAI)
constitutional-ai
rlaif
alignment
harmlessness
anthropic
rlhf
rl-for-llms
Sep 02, 2026
DPO (Direct Preference Optimization)
dpo
preference-optimization
rlhf
bradley-terry
kl-regularization
alignment
rl-for-llms
Sep 02, 2026
Generalized Advantage Estimation (GAE)
reinforcement-learning
gae
advantage-estimation
td-learning
variance-reduction
ppo
rl-for-llms
Sep 02, 2026
GRPO (Group Relative Policy Optimization)
grpo
rl
rlvr
deepseek
critic-free
dr-grpo
dapo
reasoning
rl-for-llms
Sep 02, 2026
KL Regularization in RLHF
rlhf
kl-divergence
reward-hacking
reference-policy
ppo
rl-for-llms
Sep 02, 2026
Markov Decision Process
reinforcement-learning
mdp
bellman
rl-for-llms
Sep 02, 2026
Policy Gradient Theorem
reinforcement-learning
policy-gradient
score-function
rlhf
rl-for-llms
Sep 02, 2026
PPO (Proximal Policy Optimization)
rl
ppo
trust-region
policy-gradient
rlhf
importance-sampling
rl-for-llms
Sep 02, 2026
Preference Optimization (RL-free family)
preference-optimization
dpo
ipo
kto
orpo
simpo
cpo
alignment
rl-for-llms
Sep 02, 2026
Process Reward Model (PRM)
prm
orm
process-supervision
reasoning
reward-model
rlvr
math-shepherd
rl-for-llms
Sep 02, 2026
REINFORCE
reinforcement-learning
policy-gradient
reinforce
monte-carlo
rlhf
rl-for-llms
Sep 02, 2026
Reward Model
rlhf
reward-model
preference-learning
bradley-terry
reward-hacking
rl-for-llms
Sep 02, 2026
Reward Overoptimization (Goodhart)
reward-overoptimization
reward-hacking
goodhart
scaling-laws
warm
kl-regularization
rl-for-llms
Sep 02, 2026
RL Post-Training Infrastructure
rl-infra
async-rl
off-policy
actor-learner
importance-sampling
verifier
vllm
rl-for-llms
Sep 02, 2026
RLAIF (RL from AI Feedback)
rlaif
llm-as-a-judge
constitutional-ai
rlhf
reward-model
alignment
rl-for-llms
Sep 02, 2026
RLHF (Reinforcement Learning from Human Feedback)
rlhf
alignment
reward-model
ppo
instructgpt
llm-training
rl-for-llms
Sep 02, 2026
RLVR (RL with Verifiable Rewards)
rlvr
rl
verifiable-rewards
reasoning
grpo
reward-hacking
deepseek-r1
rl-for-llms
Sep 02, 2026
Trust Region
rl
trust-region
trpo
ppo
natural-gradient
kl-divergence
rl-for-llms
Sep 02, 2026
Value Function
reinforcement-learning
value-function
bellman
critic
rl-for-llms
Sep 02, 2026
01 · MDPs & the RL Objective
reinforcement-learning
mdp
value-functions
bellman
rl-for-llms