pracaon.plpracaon.pl

Lead HPC Kubernetes Engineer

Remote, Polska
EPAM
Partner
5d
Wynagrodzenie do ustalenia
Pełny etat • Zdalna • IT, Data i AI

Najważniejsze cechy oferty

  • Min. 5 lat doświadczenia

  • Backend: Java / .NET / Node / Python

  • DevOps / Cloud: AWS, Azure, Docker, Kubernetes

  • Pełny etat

  • Praca zdalna - bez dojazdów

Description

We are seeking a Lead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, GCP, and other providers, spanning several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours, at a scale where novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI, with additional providers to be added in the near future Manage job scheduling to allocate GPU compute across training and inference workloads Define and maintain SLIs/SLOs Build monitoring and alerting systems Participate in severity escalation response and author post-incident reviews Coordinate daily with Networking, Storage, Security, and AI/ML platform teams Requirements 5+ years of experience in infrastructure engineering, cloud platforms, or HPC Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Proficiency in Terraform for writing and reviewing infrastructure-as-code daily Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) Skills in Python for tooling and automation English proficiency at B2 level or higher Nice to have Familiarity with Google Kubernetes Engine Familiarity with Amazon Elastic Kubernetes Service Knowledge of Google Cloud Platform

Requirements

  • 5+ years of experience in infrastructure engineering, cloud platforms, or HPC

  • Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets

  • Proficiency in Terraform for writing and reviewing infrastructure-as-code daily

  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)

  • Skills in Python for tooling and automation

  • English proficiency at B2 level or higher

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers

  • Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth

  • Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, GCP, and OCI, with additional providers to be added in the near future

  • Manage job scheduling to allocate GPU compute across training and inference workloads

  • Define and maintain SLIs/SLOs

  • Build monitoring and alerting systems

  • Participate in severity escalation response and author post-incident reviews

  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams

Seniority

  • Lead

Nice to have

  • Familiarity with Google Kubernetes Engine

  • Familiarity with Amazon Elastic Kubernetes Service

  • Knowledge of Google Cloud Platform

Słowa kluczowe / Umiejętności

Platform Engineering
DevOps
Kubernetes
Python
Site Reliability Engineering
Terraform
Amazon Elastic Kubernetes Service
Google Cloud Platform
Google Kubernetes Engine
Oferta została zaimportowana ze źródła zewnętrznego. Sprawdź link, aby przejść do oryginalnego ogłoszenia.Źródło ogłoszenia

Więcej podobnych ofert