Senior Engineer - SRE & Infrastructure Services
Remote, PolskaОсновні характеристики вакансії
Мін. 5 років досвіду
Сервер: Java / .NET / Node / Python
DevOps / Хмара: AWS, Azure, Docker, Kubernetes
Повний робочий день
Віддалена робота - без поїздок
Description
We are seeking a Senior Engineer to join our SRE/DevOps organization. The successful candidate will design and implement automated provisioning, deployment, management, and monitoring solutions for a large-scale, rapidly evolving portfolio of SaaS services. Working closely with architecture and development teams, the Senior Engineer will contribute to engineering standards and best practices, drive CI/CD and observability improvements, and support the team in delivering reliable, scalable infrastructure. Responsibilities Design and implement CI/CD pipelines leveraging Kubernetes, Flux, and related cloud-native technologies Implement and maintain monitoring and management solutions for cloud-based products using a combination of commercial off-the-shelf (COTS) and in-house tooling Collaborate with architects and development teams on standardized, scalable approaches for log management, service components, and infrastructure elements Evaluate and integrate AI-enabled tooling across observability, pipeline efficiency, and SRE troubleshooting workflows Develop DevOps tooling that reduces manual toil, strengthens security posture, and minimizes human error Build and maintain resilient, self-scaling systems that minimize customer impact while supporting a sustainable operational environment Participate in incident response, root cause analysis, and post-incident review processes Requirements Minimum 7 years of professional experience in software development, DevOps, and/or Site Reliability Engineering Minimum 3 years of experience building and maintaining CI/CD pipelines and SRE automation within cloud environments at scale Experience with monitoring and alerting platforms (e.g., PagerDuty, Prometheus, Grafana) Hands-on experience deploying and managing cloud infrastructure Experience with Amazon Web Services (e.g., EC2, Elasticsearch, Lambda, CloudFormation) Working experience with at least one additional cloud provider (GCP, Azure, or OCI) Experience with CI/CD toolchains (Jenkins, Kubernetes, Flux) Proficiency in one or more of the following languages: Python, Go, Java, or C Minimum English language level of B1+ Nice to have Experience applying Generative AI or ML-based tooling within an operations context Experience with DevSecOps practices and security automation Background in architecting monitoring and management systems for enterprise SaaS products
Requirements
Minimum 7 years of professional experience in software development, DevOps, and/or Site Reliability Engineering
Minimum 3 years of experience building and maintaining CI/CD pipelines and SRE automation within cloud environments at scale
Experience with monitoring and alerting platforms (e.g., PagerDuty, Prometheus, Grafana)
Hands-on experience deploying and managing cloud infrastructure
Experience with Amazon Web Services (e.g., EC2, Elasticsearch, Lambda, CloudFormation)
Working experience with at least one additional cloud provider (GCP, Azure, or OCI)
Experience with CI/CD toolchains (Jenkins, Kubernetes, Flux)
Proficiency in one or more of the following languages: Python, Go, Java, or C
Minimum English language level of B1+
Responsibilities
Design and implement CI/CD pipelines leveraging Kubernetes, Flux, and related cloud-native technologies
Implement and maintain monitoring and management solutions for cloud-based products using a combination of commercial off-the-shelf (COTS) and in-house tooling
Collaborate with architects and development teams on standardized, scalable approaches for log management, service components, and infrastructure elements
Evaluate and integrate AI-enabled tooling across observability, pipeline efficiency, and SRE troubleshooting workflows
Develop DevOps tooling that reduces manual toil, strengthens security posture, and minimizes human error
Build and maintain resilient, self-scaling systems that minimize customer impact while supporting a sustainable operational environment
Participate in incident response, root cause analysis, and post-incident review processes
Seniority
Senior
Nice to have
Experience applying Generative AI or ML-based tooling within an operations context
Experience with DevSecOps practices and security automation
Background in architecting monitoring and management systems for enterprise SaaS products
Ключові слова / Навички