Gunjan Dhanuka — Learning Notes

cuda

10 items with this tag.

  • Sep 02, 2026

    CUDA Execution Model

    • gpu
    • cuda
    • simt
    • execution-model
    • threads
    • blocks
    • shared-memory
    • llm-systems
  • Sep 02, 2026

    Memory Coalescing

    • gpu
    • coalescing
    • hbm
    • global-memory
    • bandwidth
    • sectors
    • cache-line
    • cuda
    • llm-systems
  • Sep 02, 2026

    Occupancy

    • gpu
    • occupancy
    • latency-hiding
    • registers
    • shared-memory
    • littles-law
    • cuda
    • llm-systems
  • Sep 02, 2026

    Shared Memory Bank Conflicts

    • gpu
    • shared-memory
    • bank-conflicts
    • smem
    • padding
    • transpose
    • cuda
    • llm-systems
  • Sep 02, 2026

    Warp & SIMT

    • gpu
    • simt
    • warp
    • divergence
    • predication
    • lockstep
    • cuda
    • llm-systems
  • Sep 02, 2026

    Learn: GPU Architecture & SIMT

    • gpu
    • cuda
    • simt
    • tensor-cores
    • occupancy
    • hardware
    • llm-systems
  • Sep 02, 2026

    Learn: Your First CUDA Kernels

    • gpu
    • cuda
    • kernels
    • vector-add
    • matmul
    • grid-stride
    • cuda-events
    • arithmetic-intensity
    • llm-systems
  • Sep 02, 2026

    Learn: Optimizing GEMM

    • gpu
    • cuda
    • gemm
    • matmul
    • tiling
    • coalescing
    • shared-memory
    • register-blocking
    • bank-conflicts
    • warptiling
    • cublas
    • cutlass
    • performance
  • Sep 02, 2026

    Learn: Triton & Modern Kernel Authoring

    • gpu
    • triton
    • torch-compile
    • kernels
    • cuda
    • autotuning
    • llm-systems
  • Sep 02, 2026

    GPU Systems & ML Systems for LLMs — Curriculum

    • gpu
    • ml-systems
    • cuda
    • kernels
    • distributed-training
    • curriculum

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source