CUDA / GPU Performance Engineer (Kernel Optimization)
About Us
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
Role Overview
We are looking for experienced CUDA and GPU performance engineers to analyze, profile, and optimize high-performance kernels and supporting C++ code. The role combines CUDA optimization, GPU profiling, C++, shader development, and performance analysis across different GPU architectures. No prior AI experience is required; strong systems and GPU engineering expertise is the key requirement.
CONTRACT: Freelance contractor, paid per completed task
COMMITMENT: Flexible, based on available tasks and project demand
LOCATIONS: Fully remote - GLOBAL
PROCESS: Application review, technical assessment, and onboarding
HOURLY RATE: $60-$100/h
Responsibilities
- Analyze and optimize CUDA kernels for throughput, latency, and hardware utilization.
- Profile GPU workloads to identify compute, memory, synchronization, and execution bottlenecks.
- Develop and implement targeted kernel optimization strategies.
- Refactor C++ and CUDA codebases for performance, maintainability, and portability.
- Evaluate kernel behavior across different GPU architectures and hardware generations.
- Develop or adapt shader and compute workflows using GLSL and WebGPU.
- Use GPU profiling tools to validate improvements and compare performance.
- Document optimization approaches, benchmarks, findings, and performance gains.
- Contribute technical input to GPU architecture and performance-design discussions.
- Evaluate emerging GPU programming techniques and apply relevant improvements.
Requirements
- Strong professional experience with CUDA programming and GPU kernel optimization.
- Advanced proficiency in C++, ideally in high-performance or systems programming environments.
- Proven experience profiling and tuning GPU workloads for performance.
- Hands-on experience with GPU profiling tools such as NVIDIA Nsight or comparable tools.
- Strong understanding of GPU architecture, memory hierarchy, parallel execution, and synchronization.
- Experience analyzing performance across different GPU hardware generations.
- Hands-on experience with GLSL and/or WebGPU for shader or compute development.
- Ability to document performance findings and technical decisions clearly in English.