pracaon.plpracaon.pl

Senior Site Reliability Engineer

Hybrid, Polska
EPAM
Partner
13д
Зарплата за домовленістю
Повна зайнятість • Гібридна • IT, дані та AI

Основні характеристики вакансії

  • Мін. 5 років досвіду

  • DevOps / Хмара: AWS, Azure, Docker, Kubernetes

  • Повний робочий день

  • Гібридний формат - частково віддалено

Description

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle. Responsibilities Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect Requirements 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation Knowledge of FinOps basics for AI workloads Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code

Requirements

  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away

  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling

  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation

  • Knowledge of FinOps basics for AI workloads

  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code

Responsibilities

  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source

  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards

  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets

  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward

  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively

  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle

  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated

  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect

Seniority

  • Senior

Ключові слова / Навички

Site Reliability Engineering
Azure DevOps
Azure Kubernetes Service
Microsoft Azure
Terraform
AI Observability
Цю пропозицію імпортовано із зовнішнього порталу.Джерело оголошення

Більше схожих вакансій