A small, header-only C++ layer that adds to NVIDIA CCCL the pieces it structurally cannot provide — CUDA graphs, async pipeline containers, and graphics (OpenGL/Vulkan) interop — and nothing else.
cue used to reimplement the whole CUDA-runtime RAII layer (streams, events,
devices, buffers, launch config, reductions). ~65% of it was a thinner
reimplementation of APIs NVIDIA ships free inside the CUDA Toolkit, versioned
to the toolkit, tested across every arch. cue now delegates all of that to the
cuda:: namespace (libcudacxx), Thrust, and CUB, and spends its effort on the
layer CCCL's mission explicitly excludes.
| Component | What it is |
|---|---|
cue::Error / cue::check |
Typed CUDA errors carrying the cudaError_t and the call-site std::source_location. |
cue::Graph + Recorder |
CUDA graphs — stream capture or explicit DAG node builders (add_kernel / add_memcpy_* / add_memset). |
cue::RingBuffer<T,N> |
N-buffer ring for async pipelines; templated on your owning buffer. |
cue::PingPong<T> / cue::TripleBuffer<T> |
Double / triple buffering over the same owning type. |
cue::gl::InteropTexture |
CUDA↔OpenGL shared texture (register-only; no GL loader). |
cue::vk::InteropImage / InteropMemory / InteropSemaphore |
CUDA↔Vulkan sharing (external memory + semaphores; no Vulkan loader). |
| You want | Use this (ships in the CUDA Toolkit) |
|---|---|
| stream / event / timing | cuda::stream, cuda::event, cuda::timed_event (s.record_timed_event(), end - start → ns), cuda::stream_ref |
| device query | cuda::device_ref, cuda::devices, cuda::device_attributes::compute_capability / total_global_memory |
| launch config | cuda::launch, cuda::kernel_config, or plain <<<grid, block, smem, stream>>> with s.get() |
| device / pinned / managed memory | thrust::device_vector<T>, raw cudaMallocHost, cuda::mr::legacy_pinned_memory_resource |
| 2D views | cuda::mdspan |
| reductions / scans / sorts | thrust::reduce / scan / sort, cub::Device* |
| profiling / tensor cores | <nvtx3/nvtx3.hpp>, nvcuda::wmma |
The async containers take your owning buffer — typically thrust::device_vector<T>:
cue::RingBuffer<thrust::device_vector<float>, 3> slots{CHUNK};Kept pieces take a raw cudaStream_t (plus cuda::stream_ref overloads).
#include <cue/cue.hpp>
#include <cuda/stream>
#include <thrust/device_vector.h>
cuda::device_ref dev{0};
cuda::stream s{dev};
cue::Graph g;
{
auto rec = g.record_scope(s); // begin capture
my_kernel<<<grid, block, 0, s.get()>>>(...);
} // capture ends, graph instantiated
g.launch(s); // replay
s.sync();Typed errors:
try { cue::check(cudaMalloc(&p, huge)); }
catch (const cue::Error& e) {
if (e.code() == cudaErrorMemoryAllocation) { /* ... */ }
}Build with -DCUE_BUILD_EXAMPLES=ON. Each showcases one cue feature against
the CCCL resource layer:
| Example | cue feature | Notes |
|---|---|---|
saxpy |
cue::check |
The canonical first kernel: cuda::stream + thrust::device_vector + cuda::timed_event. |
graph_pipeline |
cue::Graph + Recorder |
Capture a 12-kernel chain, replay it, measure the CPU-overhead savings vs naked launches. |
ring_buffer_overlap |
cue::TripleBuffer |
Hide H2D/compute/D2H latency with a 3-stream staged pipeline; cuda::event joins the stages. |
heat_diffusion |
cue::PingPong |
Time-stepped 2D heat equation (Jacobi sweep); writes a PPM flipbook. |
cmake -B build -DCUE_BUILD_EXAMPLES=ON -DCMAKE_CUDA_ARCHITECTURES=89
cmake --build build
./build/examples/saxpyadd_subdirectory(/path/to/cue cue_build)
target_link_libraries(my_app PRIVATE cue::cue)include(FetchContent)
FetchContent_Declare(cue
GIT_REPOSITORY https://github.com/momokrono/cue.git
GIT_TAG main
)
FetchContent_MakeAvailable(cue)
target_link_libraries(my_app PRIVATE cue::cue)cmake -B build -DCMAKE_INSTALL_PREFIX=~/.local
cmake --install buildfind_package(cue REQUIRED)
target_link_libraries(my_app PRIVATE cue::cue)cue::cue is the only CMake target — a header-only interface library linking
the CUDA runtime. There are no separate cue::gl or cue::vk targets: the
interop adapters are ordinary headers (<cue/gl/interop.hpp>,
<cue/vk/interop.hpp>) you include alongside cue::cue. They make zero
graphics API calls — no GL function loader and no Vulkan loader are linked —
and are always usable in-tree. CUE_ENABLE_GL / CUE_ENABLE_VK only gate
whether the adapter headers get installed.
Headers (flat layout under include/cue/):
| Header | Contents |
|---|---|
<cue/cue.hpp> |
aggregator: error + graph + ring_buffer |
<cue/error.hpp> |
cue::Error / cue::check |
<cue/graph.hpp> |
cue::Graph / cue::GraphNode + Recorder + DAG node builders |
<cue/ring_buffer.hpp> |
cue::RingBuffer<T,N> / PingPong<T> / TripleBuffer<T> |
<cue/gl/interop.hpp> |
cue::gl::InteropTexture (always usable in-tree; install with CUE_ENABLE_GL) |
<cue/vk/interop.hpp> |
cue::vk::InteropImage / InteropMemory / InteropSemaphore (always usable in-tree; install with CUE_ENABLE_VK) |
cue makes zero GL calls — no function loader (glad/GLEW) is needed. The header
is always usable in-tree; pass -DCUE_ENABLE_GL=ON when installing to ship it.
#include <cue/cue.hpp>
#include <cue/gl/interop.hpp>
GLuint my_tex = ...; // user-created, user-owned
cue::gl::InteropTexture interop{my_tex};
{
auto scope = interop.map_to_cuda(stream); // map for surface writes
render_kernel<<<...>>>(scope.get(), ...);
} // unmaps on scope exit
// draw my_tex with OpenGL as usualLike the GL adapter, cue makes zero Vulkan calls — no find_package(Vulkan),
no loader link. The caller creates the VkImage / VkDeviceMemory /
VkSemaphore with external-memory extensions, exports the native handle (an
fd on Linux via vkGetMemoryFdKHR, a Win32 handle on Windows), and hands it
to cue, which imports it into CUDA.
Unlike GL there is no single register call: CUDA↔Vulkan sharing goes through
external memory (cudaImportExternalMemory) plus external semaphores
for queue ordering. Synchronization is explicit — wait() before CUDA touches
shared memory, signal() to hand it back.
#include <cue/cue.hpp>
#include <cue/vk/interop.hpp>
cue::vk::InteropImage image{tex_fd, {.alloc_size = alloc_sz,
.width = W, .height = H,
.format = cudaCreateChannelDesc<uchar4>()}};
cue::vk::InteropSemaphore vk_done{sem_fd}; // Vulkan signals when its writes land
cue::vk::InteropSemaphore cu_done{sem2_fd}; // CUDA signals when its writes land
vk_done.wait(stream); // honour Vulkan's ownership
{
auto surf = image.map_to_cuda(stream);
render_kernel<<<grid, block, 0, stream.get()>>>(surf.get(), W, H);
}
cu_done.signal(stream); // hand ownership back to Vulkancue::vk::InteropMemory::map_buffer(offset, size) shares linear buffers
(SSBOs / vertex data) as a plain device pointer.
- CUDA Toolkit with CCCL (ships in-toolkit; nvcc auto-includes it).
Developed and tested on CUDA 13.3 / CCCL 3.4.0. Older toolkits work
if their bundled CCCL exposes
cuda::stream,cuda::device_ref,cuda::timed_event, andcuda::event(roughly CCCL 2.7+ / CUDA 12.6+). Note CCCL is never forward-compatible with a newer CUDA Toolkit — always same-version-or-newer. - C++20 minimum for cue's own headers (
std::span,std::source_location)..cutranslation units compile at nvcc's max (-std=c++23). - CMake 3.26+.
| Option | Default | Effect |
|---|---|---|
CUE_BUILD_EXAMPLES |
OFF |
Build the example programs |
CUE_BUILD_TESTS |
OFF |
Build the unit tests |
CUE_ENABLE_GL |
OFF |
Install the OpenGL interop adapter header |
CUE_ENABLE_VK |
OFF |
Install the Vulkan interop adapter header (external memory + semaphores; no Vulkan link) |
CUE_CXX_STANDARD |
26 |
C++ standard required by the cue::cue target (20/23/26) |
MIT License