Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
rlvr
7 items with this tag.
Sep 02, 2026
Agentic RL (multi-turn, tool-use)
agentic-rl
multi-turn
tool-use
credit-assignment
environments
rlvr
echo-trap
rl-for-llms
Sep 02, 2026
Chain-of-Thought RL (emergent reasoning)
chain-of-thought
reasoning
rlvr
grpo
o1
deepseek-r1
test-time-compute
rl-for-llms
Sep 02, 2026
GRPO (Group Relative Policy Optimization)
grpo
rl
rlvr
deepseek
critic-free
dr-grpo
dapo
reasoning
rl-for-llms
Sep 02, 2026
Process Reward Model (PRM)
prm
orm
process-supervision
reasoning
reward-model
rlvr
math-shepherd
rl-for-llms
Sep 02, 2026
RLVR (RL with Verifiable Rewards)
rlvr
rl
verifiable-rewards
reasoning
grpo
reward-hacking
deepseek-r1
rl-for-llms
Sep 02, 2026
09 · RL for Reasoning (RLVR, GRPO, R1)
rl
rlvr
grpo
deepseek-r1
reasoning
verifiable-rewards
chain-of-thought
process-reward-model
dapo
dr-grpo
o1
test-time-compute
Sep 02, 2026
10 · Frontier RL — Infra & Open Problems
rl-infra
async-rl
off-policy
rlvr
agentic-rl
entropy-collapse
reward-overoptimization
post-training
scaling-laws
frontier