Gunjan Dhanuka — Learning Notes
Search
Search
Dark mode
Light mode
Reader mode
Explorer
cuda
10 items with this tag.
Sep 02, 2026
CUDA Execution Model
gpu
cuda
simt
execution-model
threads
blocks
shared-memory
llm-systems
Sep 02, 2026
Memory Coalescing
gpu
coalescing
hbm
global-memory
bandwidth
sectors
cache-line
cuda
llm-systems
Sep 02, 2026
Occupancy
gpu
occupancy
latency-hiding
registers
shared-memory
littles-law
cuda
llm-systems
Sep 02, 2026
Shared Memory Bank Conflicts
gpu
shared-memory
bank-conflicts
smem
padding
transpose
cuda
llm-systems
Sep 02, 2026
Warp & SIMT
gpu
simt
warp
divergence
predication
lockstep
cuda
llm-systems
Sep 02, 2026
Learn: GPU Architecture & SIMT
gpu
cuda
simt
tensor-cores
occupancy
hardware
llm-systems
Sep 02, 2026
Learn: Your First CUDA Kernels
gpu
cuda
kernels
vector-add
matmul
grid-stride
cuda-events
arithmetic-intensity
llm-systems
Sep 02, 2026
Learn: Optimizing GEMM
gpu
cuda
gemm
matmul
tiling
coalescing
shared-memory
register-blocking
bank-conflicts
warptiling
cublas
cutlass
performance
Sep 02, 2026
Learn: Triton & Modern Kernel Authoring
gpu
triton
torch-compile
kernels
cuda
autotuning
llm-systems
Sep 02, 2026
GPU Systems & ML Systems for LLMs — Curriculum
gpu
ml-systems
cuda
kernels
distributed-training
curriculum