Senior HPC DevOps Engineer
Remote, PolskaKey offer highlights
Min. 5 years of experience
DevOps / Cloud: AWS, Azure, Docker, Kubernetes
Backend: Java / .NET / Node / Python
Full-time
Remote work - no commuting
Description
We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform. Responsibilities Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads Maintain automated pipelines for infrastructure provisioning and platform service deployments Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas Collaborate with developer experience teams to improve documentation Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines Ensure compute availability through capacity planning and reservation management Deploy containerized environments tuned for HPC and GPU pass-through Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions Requirements 5+ years of experience in HPC or DevOps engineering roles Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre) Proficiency in English at a B2+ level
Requirements
5+ years of experience in HPC or DevOps engineering roles
Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot
Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
Proficiency in English at a B2+ level
Responsibilities
Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction
Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads
Maintain automated pipelines for infrastructure provisioning and platform service deployments
Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
Collaborate with developer experience teams to improve documentation
Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus
Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
Ensure compute availability through capacity planning and reservation management
Deploy containerized environments tuned for HPC and GPU pass-through
Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions
Seniority
Senior
Keywords / Skills