Gunjan Dhanuka — Learning Notes

matmul

4 items with this tag.

  • Sep 02, 2026

    GEMM (General Matrix Multiply)

    • gpu
    • gemm
    • matmul
    • tensor-core
    • arithmetic-intensity
    • tiling
    • transformer
    • llm-systems
  • Sep 02, 2026

    Tensor Core

    • gpu
    • tensor-core
    • mma
    • precision
    • fp8
    • bf16
    • tf32
    • matmul
    • llm-systems
  • Sep 02, 2026

    Learn: Your First CUDA Kernels

    • gpu
    • cuda
    • kernels
    • vector-add
    • matmul
    • grid-stride
    • cuda-events
    • arithmetic-intensity
    • llm-systems
  • Sep 02, 2026

    Learn: Optimizing GEMM

    • gpu
    • cuda
    • gemm
    • matmul
    • tiling
    • coalescing
    • shared-memory
    • register-blocking
    • bank-conflicts
    • warptiling
    • cublas
    • cutlass
    • performance

Created with Quartz v5.0.0 © 2026

  • Personal site
  • Research
  • Field notes
  • Source