Experimental native AMD runtime · v1.4.0

CUDA-oriented workflows.
Native AMD execution.

CUDAtoAMD is a clean-room developer runtime for assessing and incrementally adapting supported CUDA-oriented source work to HIP/ROCm. It reports real AMD devices, compiles a bounded source/PTX subset to native AMD code objects, and keeps unsupported paths explicit.

Build locally

No driver spoofing. No NVIDIA DLL redistribution. No fabricated capability reports.

Truthful hardwareReports HIP/ROCm devices and native architecture as detected.

Clean-room surfaceImplements documented compatibility behavior without bundling NVIDIA software.

Fail-closed boundariesUnsupported calls and CUDA-native wheels are identified rather than silently misrun.

Get started

Verify the AMD path in three steps.

  1. 01

    Build the HIP runtime

    cmake -S . -B build-hip -G "NMake Makefiles" -DCOMPATCUDA_ENABLE_HIP=ON -DCOMPATCUDA_BUILD_TESTS=ON
    cmake --build build-hip
  2. 02

    Discover the toolchain

    python -m compat toolchain --format json
  3. 03

    Initialize a real device

    python -m compat init --library build-hip/compatcuda.dll --self-test --format json

Run from a 64-bit Visual Studio developer prompt on Windows. HIP and hipBLAS are required for the native build; hipFFT is optional.

Implemented surface

Useful where the boundary is known.

Each capability is deliberately scoped and documented. The project does not represent this list as universal CUDA compatibility.

01

Native runtime

HIP-backed device discovery, allocations, streams, events, memory pools, HSACO module loading and kernel dispatch.

compat init

02

CUDA-facing subset

Clean-room Runtime and Driver API facades with deterministic error behavior and explicit native AMD interop.

include/cuda_*.h

03

Source to code object

CUDA-syntax lowering plus a bounded PTX compiler path targeting AMDGPU code objects through HIP tools.

compat cc

04

Numerical building blocks

FP32 and mixed-precision GEMM subset, C2C FFT when hipFFT is present, and selected neural operators.

hipBLAS · hipFFT

05

Python workflow

Native binding lifecycle, explicit NumPy/CPU-PyTorch staging, environment diagnostics and guarded wheel admission.

compat wheel-doctor

06

Release checks

Static version alignment, native header coverage, install support, capability documentation and wheel packaging checks.

compat release-check

Compatibility boundary

Honest limits make the supported path dependable.

CUDAtoAMD does not run arbitrary CUDA binaries, transparently execute CUDA-locked wheels, impersonate NVIDIA hardware, or provide full CUDA, PTX/NVVM or framework parity.

Use the included analyzer and wheel doctor before committing to a porting path. Unsupported constructs should return a clear result rather than an unreliable approximation.

Repository documentation

Everything needed to evaluate the project is in the source tree.

docs/local-developer-guide.mdBuild, test and use the native runtime.

docs/initialization.mdReal-device initialization and troubleshooting.

docs/cuda-source-compatibility.mdSupported CUDA-facing source workflow and limitations.

docs/cuda-wheel-execution.mdWheel admission and refusal behavior.