Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CUE: CUDA Unified Engine

A small, header-only C++ layer that adds to NVIDIA CCCL the pieces it structurally cannot provide — CUDA graphs, async pipeline containers, and graphics (OpenGL/Vulkan) interop — and nothing else.

cue used to reimplement the whole CUDA-runtime RAII layer (streams, events, devices, buffers, launch config, reductions). ~65% of it was a thinner reimplementation of APIs NVIDIA ships free inside the CUDA Toolkit, versioned to the toolkit, tested across every arch. cue now delegates all of that to the cuda:: namespace (libcudacxx), Thrust, and CUB, and spends its effort on the layer CCCL's mission explicitly excludes.

What cue provides

Component What it is
cue::Error / cue::check Typed CUDA errors carrying the cudaError_t and the call-site std::source_location.
cue::Graph + Recorder CUDA graphs — stream capture or explicit DAG node builders (add_kernel / add_memcpy_* / add_memset).
cue::RingBuffer<T,N> N-buffer ring for async pipelines; templated on your owning buffer.
cue::PingPong<T> / cue::TripleBuffer<T> Double / triple buffering over the same owning type.
cue::gl::InteropTexture CUDA↔OpenGL shared texture (register-only; no GL loader).
cue::vk::InteropImage / InteropMemory / InteropSemaphore CUDA↔Vulkan sharing (external memory + semaphores; no Vulkan loader).

What cue does not provide (use CCCL directly)

You want Use this (ships in the CUDA Toolkit)
stream / event / timing cuda::stream, cuda::event, cuda::timed_event (s.record_timed_event(), end - start → ns), cuda::stream_ref
device query cuda::device_ref, cuda::devices, cuda::device_attributes::compute_capability / total_global_memory
launch config cuda::launch, cuda::kernel_config, or plain <<<grid, block, smem, stream>>> with s.get()
device / pinned / managed memory thrust::device_vector<T>, raw cudaMallocHost, cuda::mr::legacy_pinned_memory_resource
2D views cuda::mdspan
reductions / scans / sorts thrust::reduce / scan / sort, cub::Device*
profiling / tensor cores <nvtx3/nvtx3.hpp>, nvcuda::wmma

The async containers take your owning buffer — typically thrust::device_vector<T>:

cue::RingBuffer<thrust::device_vector<float>, 3> slots{CHUNK};

Kept pieces take a raw cudaStream_t (plus cuda::stream_ref overloads).

Quick start

#include <cue/cue.hpp>
#include <cuda/stream>
#include <thrust/device_vector.h>

cuda::device_ref dev{0};
cuda::stream s{dev};

cue::Graph g;
{
    auto rec = g.record_scope(s);          // begin capture
    my_kernel<<<grid, block, 0, s.get()>>>(...);
}                                         // capture ends, graph instantiated
g.launch(s);                              // replay
s.sync();

Typed errors:

try { cue::check(cudaMalloc(&p, huge)); }
catch (const cue::Error& e) {
    if (e.code() == cudaErrorMemoryAllocation) { /* ... */ }
}

Examples

Build with -DCUE_BUILD_EXAMPLES=ON. Each showcases one cue feature against the CCCL resource layer:

Example cue feature Notes
saxpy cue::check The canonical first kernel: cuda::stream + thrust::device_vector + cuda::timed_event.
graph_pipeline cue::Graph + Recorder Capture a 12-kernel chain, replay it, measure the CPU-overhead savings vs naked launches.
ring_buffer_overlap cue::TripleBuffer Hide H2D/compute/D2H latency with a 3-stream staged pipeline; cuda::event joins the stages.
heat_diffusion cue::PingPong Time-stepped 2D heat equation (Jacobi sweep); writes a PPM flipbook.
cmake -B build -DCUE_BUILD_EXAMPLES=ON -DCMAKE_CUDA_ARCHITECTURES=89
cmake --build build
./build/examples/saxpy

Installation

Option 1: add_subdirectory (recommended)

add_subdirectory(/path/to/cue cue_build)
target_link_libraries(my_app PRIVATE cue::cue)

Option 2: FetchContent

include(FetchContent)
FetchContent_Declare(cue
    GIT_REPOSITORY https://github.com/momokrono/cue.git
    GIT_TAG main
)
FetchContent_MakeAvailable(cue)
target_link_libraries(my_app PRIVATE cue::cue)

Option 3: install + find_package

cmake -B build -DCMAKE_INSTALL_PREFIX=~/.local
cmake --install build
find_package(cue REQUIRED)
target_link_libraries(my_app PRIVATE cue::cue)

Components

cue::cue is the only CMake target — a header-only interface library linking the CUDA runtime. There are no separate cue::gl or cue::vk targets: the interop adapters are ordinary headers (<cue/gl/interop.hpp>, <cue/vk/interop.hpp>) you include alongside cue::cue. They make zero graphics API calls — no GL function loader and no Vulkan loader are linked — and are always usable in-tree. CUE_ENABLE_GL / CUE_ENABLE_VK only gate whether the adapter headers get installed.

Headers (flat layout under include/cue/):

Header Contents
<cue/cue.hpp> aggregator: error + graph + ring_buffer
<cue/error.hpp> cue::Error / cue::check
<cue/graph.hpp> cue::Graph / cue::GraphNode + Recorder + DAG node builders
<cue/ring_buffer.hpp> cue::RingBuffer<T,N> / PingPong<T> / TripleBuffer<T>
<cue/gl/interop.hpp> cue::gl::InteropTexture (always usable in-tree; install with CUE_ENABLE_GL)
<cue/vk/interop.hpp> cue::vk::InteropImage / InteropMemory / InteropSemaphore (always usable in-tree; install with CUE_ENABLE_VK)

OpenGL interop

cue makes zero GL calls — no function loader (glad/GLEW) is needed. The header is always usable in-tree; pass -DCUE_ENABLE_GL=ON when installing to ship it.

#include <cue/cue.hpp>
#include <cue/gl/interop.hpp>

GLuint my_tex = ...;                          // user-created, user-owned
cue::gl::InteropTexture interop{my_tex};
{
    auto scope = interop.map_to_cuda(stream); // map for surface writes
    render_kernel<<<...>>>(scope.get(), ...);
}                                            // unmaps on scope exit
// draw my_tex with OpenGL as usual

Vulkan interop

Like the GL adapter, cue makes zero Vulkan calls — no find_package(Vulkan), no loader link. The caller creates the VkImage / VkDeviceMemory / VkSemaphore with external-memory extensions, exports the native handle (an fd on Linux via vkGetMemoryFdKHR, a Win32 handle on Windows), and hands it to cue, which imports it into CUDA.

Unlike GL there is no single register call: CUDA↔Vulkan sharing goes through external memory (cudaImportExternalMemory) plus external semaphores for queue ordering. Synchronization is explicit — wait() before CUDA touches shared memory, signal() to hand it back.

#include <cue/cue.hpp>
#include <cue/vk/interop.hpp>

cue::vk::InteropImage  image{tex_fd, {.alloc_size = alloc_sz,
                                      .width = W, .height = H,
                                      .format = cudaCreateChannelDesc<uchar4>()}};
cue::vk::InteropSemaphore vk_done{sem_fd};   // Vulkan signals when its writes land
cue::vk::InteropSemaphore cu_done{sem2_fd};  // CUDA signals when its writes land

vk_done.wait(stream);                        // honour Vulkan's ownership
{
    auto surf = image.map_to_cuda(stream);
    render_kernel<<<grid, block, 0, stream.get()>>>(surf.get(), W, H);
}
cu_done.signal(stream);                      // hand ownership back to Vulkan

cue::vk::InteropMemory::map_buffer(offset, size) shares linear buffers (SSBOs / vertex data) as a plain device pointer.

Requirements

  • CUDA Toolkit with CCCL (ships in-toolkit; nvcc auto-includes it). Developed and tested on CUDA 13.3 / CCCL 3.4.0. Older toolkits work if their bundled CCCL exposes cuda::stream, cuda::device_ref, cuda::timed_event, and cuda::event (roughly CCCL 2.7+ / CUDA 12.6+). Note CCCL is never forward-compatible with a newer CUDA Toolkit — always same-version-or-newer.
  • C++20 minimum for cue's own headers (std::span, std::source_location). .cu translation units compile at nvcc's max (-std=c++23).
  • CMake 3.26+.

Build options

Option Default Effect
CUE_BUILD_EXAMPLES OFF Build the example programs
CUE_BUILD_TESTS OFF Build the unit tests
CUE_ENABLE_GL OFF Install the OpenGL interop adapter header
CUE_ENABLE_VK OFF Install the Vulkan interop adapter header (external memory + semaphores; no Vulkan link)
CUE_CXX_STANDARD 26 C++ standard required by the cue::cue target (20/23/26)

License

MIT License

About

cuda engine

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages