Gunjan Dhanuka — Learning Notes

importance-sampling

3 items with this tag.

  • Sep 02, 2026

    PPO (Proximal Policy Optimization)

    • rl
    • ppo
    • trust-region
    • policy-gradient
    • rlhf
    • importance-sampling
    • rl-for-llms
  • Sep 02, 2026

    RL Post-Training Infrastructure

    • rl-infra
    • async-rl
    • off-policy
    • actor-learner
    • importance-sampling
    • verifier
    • vllm
    • rl-for-llms
  • Sep 02, 2026

    04 · PPO

    • rl
    • ppo
    • trpo
    • trust-region
    • policy-gradient
    • rlhf
    • importance-sampling

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source