Lead HPC Kubernetes Engineer
Remote, PolskaKey offer highlights
Min. 5 years of experience
Backend: Java / .NET / Node / Python
DevOps / Cloud: AWS, Azure, Docker, Kubernetes
Full-time
Remote work - no commuting
Description
We are seeking a Lead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, GCP, and other providers, spanning several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours, at a scale where novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI, with additional providers to be added in the near future Manage job scheduling to allocate GPU compute across training and inference workloads Define and maintain SLIs/SLOs Build monitoring and alerting systems Participate in severity escalation response and author post-incident reviews Coordinate daily with Networking, Storage, Security, and AI/ML platform teams Requirements 5+ years of experience in infrastructure engineering, cloud platforms, or HPC Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Proficiency in Terraform for writing and reviewing infrastructure-as-code daily Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) Skills in Python for tooling and automation English proficiency at B2 level or higher Nice to have Familiarity with Google Kubernetes Engine Familiarity with Amazon Elastic Kubernetes Service Knowledge of Google Cloud Platform
Requirements
5+ years of experience in infrastructure engineering, cloud platforms, or HPC
Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
Proficiency in Terraform for writing and reviewing infrastructure-as-code daily
Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
Skills in Python for tooling and automation
English proficiency at B2 level or higher
Responsibilities
Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers
Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth
Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI, with additional providers to be added in the near future
Manage job scheduling to allocate GPU compute across training and inference workloads
Define and maintain SLIs/SLOs
Build monitoring and alerting systems
Participate in severity escalation response and author post-incident reviews
Coordinate daily with Networking, Storage, Security, and AI/ML platform teams
Seniority
Lead
Nice to have
Familiarity with Google Kubernetes Engine
Familiarity with Amazon Elastic Kubernetes Service
Knowledge of Google Cloud Platform
Keywords / Skills