-
Notifications
You must be signed in to change notification settings - Fork 251
Pull requests: NVIDIA/cudnn-frontend
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Further reduce FROST LA CPU overhead + fully support all L2norm, beta, gate fusion
#616
opened Aug 16, 2026 by
jhjpark
Collaborator
Loading…
3 tasks done
frost: fix oversized-SMEM query crashing all frost GEMM on CUDA<13.4 cuda-python
#615
opened Aug 16, 2026 by
YangXu1990uiuc
Collaborator
Loading…
frost(sdpa): order every execute-path tensor prep on the launch stream (Rule 5)
cat-bugfix
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#614
opened Aug 16, 2026 by
vedaanta
Collaborator
Loading…
3 tasks done
First-class cudnn.Handle (create_handle returns an object owning {backend handle, device, stream})
#612
opened Aug 16, 2026 by
YangXu1990uiuc
Collaborator
Loading…
FE execute-path host overhead: idempotent set_stream + skip the discarded per-call context
#611
opened Aug 15, 2026 by
YangXu1990uiuc
Collaborator
Loading…
Prototype: fused dense SwiGLU-MLP autograd op + cuDNN framework-integration perf guide
#609
opened Aug 15, 2026 by
YangXu1990uiuc
Collaborator
Loading…
frost(sdpa): SM120 THD zero-host-read execute — device-built metadata, declared-S_q envelope grid, CUDA-graph capturable (issue #552)
cat-enhancements
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#608
opened Aug 15, 2026 by
vedaanta
Collaborator
Loading…
NSA compression: suppress aliased O1 store for static short queries
#599
opened Aug 14, 2026 by
mgoldfarb-nvidia
•
Draft
2 of 3 tasks
Add cudnn.fla: a cuDNN drop-in for flash-linear-attention (GDN + KDA)
#596
opened Aug 14, 2026 by
YangXu1990uiuc
Collaborator
Loading…
Add more features to sm120 frost fp8 sdpa_fwd kernel
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add SM100 FP8 SDPA support for d192/d128
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#594
opened Aug 14, 2026 by
adshen
Collaborator
Loading…
3 tasks done
Avoid zero-initializing JAX grouped GEMM outputs
#592
opened Aug 14, 2026 by
mgoldfarb-nvidia
•
Draft
2 of 3 tasks
Drive the grouped GEMM host overhead to the dispatch floor (wrapper 84->~22us; execute() trusts by default)
#591
opened Aug 13, 2026 by
YangXu1990uiuc
Collaborator
Loading…
3 of 4 tasks
Memoize the CuTeDSL hot-path dtype/stride conversions (~20% off per-call host)
#589
opened Aug 13, 2026 by
YangXu1990uiuc
Collaborator
Loading…
2 of 3 tasks
sdpa fp8: THD/varlen execute + d<=128 envelope (Rubin) for per-tensor FP8
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#587
opened Aug 13, 2026 by
vedaanta
Collaborator
Loading…
sdpa fp8 sm100: fused LDTM row-max on cc10.3 (per-tensor FP8)
cat-enhancements
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#586
opened Aug 13, 2026 by
vedaanta
Collaborator
Loading…
Add LPT_L2 scheduler
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#585
opened Aug 13, 2026 by
yanzhuo607
Collaborator
Loading…
2 of 3 tasks
Improve large tensor fuzzer plan fallback and failure output
#583
opened Aug 13, 2026 by
msalasooNV
Collaborator
Loading…
2 of 3 tasks
Emit the FROST gemm parameter tables, and a demo that launches from them
#582
opened Aug 13, 2026 by
YangXu1990uiuc
Collaborator
•
Draft
2 of 3 tasks
frost(sdpa): dense-padded cu_seq_len support (kernel CU read mode)
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#578
opened Aug 13, 2026 by
vedaanta
Collaborator
Loading…
3 tasks done
frost(sdpa): drop the redundant zero-fill of the never-read THD sinks dummy (SM100)
cat-cleanup
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
mod-frost
orig-nv-eng
Reported or requested by NVIDIA engineering.
#573
opened Aug 13, 2026 by
vedaanta
Collaborator
Loading…
dsa(indexer_backward): validate output & plan signature before the score-grad precompute (default SM100/SM90)
#572
opened Aug 13, 2026 by
zkyue
Contributor
Loading…
Add SM100 DSA sparse attention forward kernels
#569
opened Aug 13, 2026 by
jiayus-nvidia
Contributor
Loading…
2 of 3 tasks
Previous Next
ProTip!
no:milestone will show everything without a milestone.