Skip to content

Try enabling fwd broadcating Enzyme test - #3214

Open
kshyatt wants to merge 1 commit into
mainfrom
ksh/enz_cumemcpy
Open

Try enabling fwd broadcating Enzyme test#3214
kshyatt wants to merge 1 commit into
mainfrom
ksh/enz_cumemcpy

Conversation

@kshyatt

@kshyatt kshyatt commented Jul 24, 2026

Copy link
Copy Markdown
Member

No description provided.

@kshyatt
kshyatt requested a review from wsmoses July 24, 2026 06:43
@codecov

codecov Bot commented Jul 24, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 92.69%. Comparing base (cac0ac9) to head (1006860).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #3214   +/-   ##
=======================================
  Coverage   92.68%   92.69%           
=======================================
  Files         167      167           
  Lines       13901    13901           
=======================================
+ Hits        12884    12885    +1     
+ Misses       1017     1016    -1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown
Contributor

CUDA.jl Benchmarks

Details
Benchmark suite Current: 1006860 Previous: 069cdff Ratio
array/accumulate/Float32/1d 97155 ns 98074 ns 0.99
array/accumulate/Float32/dims=1 71139 ns 74919 ns 0.95
array/accumulate/Float32/dims=1L 1598600 ns 1600002 ns 1.00
array/accumulate/Float32/dims=2 136323 ns 140367 ns 0.97
array/accumulate/Float32/dims=2L 659812 ns 659745 ns 1.00
array/accumulate/Int64/1d 117189 ns 117818 ns 0.99
array/accumulate/Int64/dims=1 75434 ns 79251 ns 0.95
array/accumulate/Int64/dims=1L 1715039 ns 1717606 ns 1.00
array/accumulate/Int64/dims=2 148320 ns 153507 ns 0.97
array/accumulate/Int64/dims=2L 986302 ns 986800 ns 1.00
array/broadcast 17442 ns 18060 ns 0.97
array/broadcast launch 8142.333333333333 ns
array/construct 900.8840579710145 ns 869.1052631578947 ns 1.04
array/copy 16120 ns 15967 ns 1.01
array/copyto!/cpu_to_gpu 208371 ns 208599 ns 1.00
array/copyto!/gpu_to_cpu 241018 ns 241280 ns 1.00
array/copyto!/gpu_to_gpu 8782.333333333334 ns 9209.333333333334 ns 0.95
array/iteration/findall/bool 130642 ns 132021 ns 0.99
array/iteration/findall/int 144511 ns 145286 ns 0.99
array/iteration/findfirst/bool 67218 ns 67274 ns 1.00
array/iteration/findfirst/int 68875 ns 68529 ns 1.01
array/iteration/findmin/1d 61659 ns 64309 ns 0.96
array/iteration/findmin/2d 98627 ns 99940 ns 0.99
array/iteration/logical 181793 ns 185815 ns 0.98
array/iteration/scalar 61876 ns 63136 ns 0.98
array/permutedims/2d 47540 ns 48431 ns 0.98
array/permutedims/3d 48910 ns 50026 ns 0.98
array/permutedims/4d 49196 ns 49849 ns 0.99
array/random/rand/Float32 11660 ns 11669 ns 1.00
array/random/rand/Int64 22239 ns 22883 ns 0.97
array/random/rand!/Float32 7653.75 ns 7731.75 ns 0.99
array/random/rand!/Int64 19867 ns 20210 ns 0.98
array/random/randn/Float32 32444 ns 32574 ns 1.00
array/random/randn!/Float32 23053 ns 23189 ns 0.99
array/reductions/mapreduce/Float32/1d 31100 ns 31371 ns 0.99
array/reductions/mapreduce/Float32/dims=1 36772 ns 37491 ns 0.98
array/reductions/mapreduce/Float32/dims=1L 49816 ns 49977 ns 1.00
array/reductions/mapreduce/Float32/dims=2 54135 ns 54508 ns 0.99
array/reductions/mapreduce/Float32/dims=2L 65938 ns 66329 ns 0.99
array/reductions/mapreduce/Int64/1d 37892 ns 37643 ns 1.01
array/reductions/mapreduce/Int64/dims=1 40039 ns 39961 ns 1.00
array/reductions/mapreduce/Int64/dims=1L 87861 ns 87917 ns 1.00
array/reductions/mapreduce/Int64/dims=2 56739 ns 57047 ns 0.99
array/reductions/mapreduce/Int64/dims=2L 81825 ns 82751 ns 0.99
array/reductions/reduce/Float32/1d 31155 ns 31482 ns 0.99
array/reductions/reduce/Float32/dims=1 36991 ns 37630 ns 0.98
array/reductions/reduce/Float32/dims=1L 49513 ns 50073 ns 0.99
array/reductions/reduce/Float32/dims=2 54316 ns 54518 ns 1.00
array/reductions/reduce/Float32/dims=2L 67515 ns 67963 ns 0.99
array/reductions/reduce/Int64/1d 38471 ns 38054 ns 1.01
array/reductions/reduce/Int64/dims=1 39921 ns 39983 ns 1.00
array/reductions/reduce/Int64/dims=1L 87622 ns 87988 ns 1.00
array/reductions/reduce/Int64/dims=2 56699 ns 56860 ns 1.00
array/reductions/reduce/Int64/dims=2L 82090 ns 82191 ns 1.00
array/reverse/1d 16416 ns 16098 ns 1.02
array/reverse/1dL 69054 ns 68844 ns 1.00
array/reverse/1dL_inplace 67058 ns 66918 ns 1.00
array/reverse/1d_inplace 8172.333333333333 ns 9757 ns 0.84
array/reverse/2d 19106 ns 19108 ns 1.00
array/reverse/2dL 72632 ns 72561 ns 1.00
array/reverse/2dL_inplace 66890 ns 66785 ns 1.00
array/reverse/2d_inplace 9557 ns 9397 ns 1.02
array/sorting/1d 2638930 ns 2656939 ns 0.99
array/sorting/2d 1027515 ns 1038211 ns 0.99
array/sorting/by 3181446 ns 3192363 ns 1.00
cuda/synchronization/context/auto 1014.5 ns 1015.3 ns 1.00
cuda/synchronization/context/blocking 787.9514563106796 ns 784.6960784313726 ns 1.00
cuda/synchronization/context/nonblocking 5791.2 ns 5655.166666666667 ns 1.02
cuda/synchronization/stream/auto 874.5555555555555 ns 857.3174603174604 ns 1.02
cuda/synchronization/stream/blocking 661.2784810126582 ns 654.7730061349694 ns 1.01
cuda/synchronization/stream/nonblocking 5682.666666666667 ns 5491 ns 1.03
integration/byval/reference 147357 ns 147245 ns 1.00
integration/byval/slices=1 151394 ns 149267 ns 1.01
integration/byval/slices=2 294544 ns 292191 ns 1.01
integration/byval/slices=3 437519 ns 434868 ns 1.01
integration/cudadevrt 104490 ns 104357 ns 1.00
integration/volumerhs 9144772 ns 9315717 ns 0.98
kernel/indexing 12473 ns 12183 ns 1.02
kernel/indexing_checked 13171 ns 13078 ns 1.01
kernel/launch 2015.111111111111 ns 2032.3333333333333 ns 0.99
kernel/occupancy 653.3475609756098 ns 667.8466257668712 ns 0.98
kernel/rand 13690 ns 14903 ns 0.92
latency/import 4128007713 ns 4116528672 ns 1.00
latency/precompile 4814647233 ns 4814427609 ns 1.00
latency/ttfp 5205590998 ns 5198634459 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@wsmoses wsmoses left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you'll need to backport this to an earlier cuda to test per the gpucompiler incompat

@kshyatt

kshyatt commented Jul 27, 2026

Copy link
Copy Markdown
Member Author

I think I'll just wait until you merge Valentin's PR :D

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants