Gunjan Dhanuka — Learning Notes

scaling-laws

2 items with this tag.

  • Sep 02, 2026

    Reward Overoptimization (Goodhart)

    • reward-overoptimization
    • reward-hacking
    • goodhart
    • scaling-laws
    • warm
    • kl-regularization
    • rl-for-llms
  • Sep 02, 2026

    10 · Frontier RL — Infra & Open Problems

    • rl-infra
    • async-rl
    • off-policy
    • rlvr
    • agentic-rl
    • entropy-collapse
    • reward-overoptimization
    • post-training
    • scaling-laws
    • frontier

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source