Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
rl
7 items with this tag.
Sep 02, 2026
GRPO (Group Relative Policy Optimization)
grpo
rl
rlvr
deepseek
critic-free
dr-grpo
dapo
reasoning
rl-for-llms
Sep 02, 2026
PPO (Proximal Policy Optimization)
rl
ppo
trust-region
policy-gradient
rlhf
importance-sampling
rl-for-llms
Sep 02, 2026
RLVR (RL with Verifiable Rewards)
rlvr
rl
verifiable-rewards
reasoning
grpo
reward-hacking
deepseek-r1
rl-for-llms
Sep 02, 2026
Trust Region
rl
trust-region
trpo
ppo
natural-gradient
kl-divergence
rl-for-llms
Sep 02, 2026
04 · PPO
rl
ppo
trpo
trust-region
policy-gradient
rlhf
importance-sampling
Sep 02, 2026
05 · RL on Token Sequences
rl
rlhf
llm
token-level-mdp
kl-regularization
ppo
credit-assignment
Sep 02, 2026
09 · RL for Reasoning (RLVR, GRPO, R1)
rl
rlvr
grpo
deepseek-r1
reasoning
verifiable-rewards
chain-of-thought
process-reward-model
dapo
dr-grpo
o1
test-time-compute