Senior HPC / CUDA Optimization Engineer
Remote, PolskaKey offer highlights
Min. 5 years of experience
Backend: Java / .NET / Node / Python
Full-time
Remote work - no commuting
Description
We are seeking an experienced HPC / CUDA Optimization Engineer with strong expertise in modern C++, CUDA, GPU computing, and performance optimization for large-scale scientific and engineering applications. Responsibilities Design, develop, and optimize CUDA kernels for large-scale scientific and engineering applications Analyze and improve GPU performance using profiling tools to identify and resolve bottlenecks Implement and optimize distributed and multi-GPU computing solutions using MPI, NCCL, and NVSHMEM Develop and maintain high-performance C++ and Python code for performance-critical, high-load systems Collaborate with cross-functional teams to architect scalable and efficient software solutions Debug complex performance and memory issues across GPU architectures and distributed systems Research and apply best practices in HPC, numerical methods, and GPU computing to improve existing codebases Requirements 5+ years of professional experience with C++ and Python development Deep knowledge of CUDA, GPU architecture, PTX, memory hierarchy, and kernel optimization Experience with GPU profiling, performance analysis, and bottleneck identification Hands-on experience with distributed and multi-GPU computing (MPI, NCCL, NVSHMEM) Experience developing and optimizing performance-critical, high-load software systems Strong debugging and software architecture skills Nice to have Knowledge of Oil & Gas domain is a huge bonus Experience with RTM/FWI, seismic imaging, or geoscience applications Development of HDF5-based data viewing and analysis tools Familiarity with scientific computing and numerical methods
Requirements
5+ years of professional experience with C++ and Python development
Deep knowledge of CUDA, GPU architecture, PTX, memory hierarchy, and kernel optimization
Experience with GPU profiling, performance analysis, and bottleneck identification
Hands-on experience with distributed and multi-GPU computing (MPI, NCCL, NVSHMEM)
Experience developing and optimizing performance-critical, high-load software systems
Strong debugging and software architecture skills
Responsibilities
Design, develop, and optimize CUDA kernels for large-scale scientific and engineering applications
Analyze and improve GPU performance using profiling tools to identify and resolve bottlenecks
Implement and optimize distributed and multi-GPU computing solutions using MPI, NCCL, and NVSHMEM
Develop and maintain high-performance C++ and Python code for performance-critical, high-load systems
Collaborate with cross-functional teams to architect scalable and efficient software solutions
Debug complex performance and memory issues across GPU architectures and distributed systems
Research and apply best practices in HPC, numerical methods, and GPU computing to improve existing codebases
Seniority
Senior
Nice to have
Knowledge of Oil & Gas domain is a huge bonus
Experience with RTM/FWI, seismic imaging, or geoscience applications
Development of HDF5-based data viewing and analysis tools
Familiarity with scientific computing and numerical methods
Keywords / Skills