Gunjan Dhanuka — Learning Notes

rlvr

7 items with this tag.

  • Sep 02, 2026

    Agentic RL (multi-turn, tool-use)

    • agentic-rl
    • multi-turn
    • tool-use
    • credit-assignment
    • environments
    • rlvr
    • echo-trap
    • rl-for-llms
  • Sep 02, 2026

    Chain-of-Thought RL (emergent reasoning)

    • chain-of-thought
    • reasoning
    • rlvr
    • grpo
    • o1
    • deepseek-r1
    • test-time-compute
    • rl-for-llms
  • Sep 02, 2026

    GRPO (Group Relative Policy Optimization)

    • grpo
    • rl
    • rlvr
    • deepseek
    • critic-free
    • dr-grpo
    • dapo
    • reasoning
    • rl-for-llms
  • Sep 02, 2026

    Process Reward Model (PRM)

    • prm
    • orm
    • process-supervision
    • reasoning
    • reward-model
    • rlvr
    • math-shepherd
    • rl-for-llms
  • Sep 02, 2026

    RLVR (RL with Verifiable Rewards)

    • rlvr
    • rl
    • verifiable-rewards
    • reasoning
    • grpo
    • reward-hacking
    • deepseek-r1
    • rl-for-llms
  • Sep 02, 2026

    09 · RL for Reasoning (RLVR, GRPO, R1)

    • rl
    • rlvr
    • grpo
    • deepseek-r1
    • reasoning
    • verifiable-rewards
    • chain-of-thought
    • process-reward-model
    • dapo
    • dr-grpo
    • o1
    • test-time-compute
  • Sep 02, 2026

    10 · Frontier RL — Infra & Open Problems

    • rl-infra
    • async-rl
    • off-policy
    • rlvr
    • agentic-rl
    • entropy-collapse
    • reward-overoptimization
    • post-training
    • scaling-laws
    • frontier

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source