Gunjan Dhanuka — Learning Notes

reinforcement-learning

11 items with this tag.

  • Sep 02, 2026

    Actor-Critic

    • reinforcement-learning
    • actor-critic
    • critic
    • variance-reduction
    • ppo
    • rl-for-llms
  • Sep 02, 2026

    Advantage Function

    • reinforcement-learning
    • advantage
    • variance-reduction
    • policy-gradient
    • rl-for-llms
  • Sep 02, 2026

    Generalized Advantage Estimation (GAE)

    • reinforcement-learning
    • gae
    • advantage-estimation
    • td-learning
    • variance-reduction
    • ppo
    • rl-for-llms
  • Sep 02, 2026

    Markov Decision Process

    • reinforcement-learning
    • mdp
    • bellman
    • rl-for-llms
  • Sep 02, 2026

    Policy Gradient Theorem

    • reinforcement-learning
    • policy-gradient
    • score-function
    • rlhf
    • rl-for-llms
  • Sep 02, 2026

    REINFORCE

    • reinforcement-learning
    • policy-gradient
    • reinforce
    • monte-carlo
    • rlhf
    • rl-for-llms
  • Sep 02, 2026

    Value Function

    • reinforcement-learning
    • value-function
    • bellman
    • critic
    • rl-for-llms
  • Sep 02, 2026

    01 · MDPs & the RL Objective

    • reinforcement-learning
    • mdp
    • value-functions
    • bellman
    • rl-for-llms
  • Sep 02, 2026

    02 · Policy Gradients

    • reinforcement-learning
    • policy-gradient
    • reinforce
    • baselines
    • variance-reduction
    • rlhf
  • Sep 02, 2026

    03 · Actor-Critic & GAE

    • reinforcement-learning
    • actor-critic
    • gae
    • advantage-estimation
    • variance-reduction
    • ppo
    • rlhf
  • Sep 02, 2026

    RL for Frontier LLMs — Curriculum

    • reinforcement-learning
    • rlhf
    • llm-training
    • curriculum

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source