Senior DevOps / SRE Engineer
Remote Poland, PolskaKey offer highlights
Backend: Java / .NET / Node / Python
DevOps / Cloud: AWS, Azure, Docker, Kubernetes
Looking for experts - senior/expert
English B2/C1
Full-time
Description
Join a healthcare data platform team building client-scoped SLO monitoring and alerting solutions. The project includes automated dashboard and alert generation across 110-140 services, as well as an internal application providing customer-specific SLA visibility. This role focuses on application instrumentation, telemetry quality, and observability adoption across distributed systems. AI-assisted development is expected and encouraged. In Luxoft, our culture is one that strives on solving difficult problems focusing on product engineering based on hypothesis testing to empower people to come up with ideas. Great Place to Work Institute certifies us as one of the Top 10 companies to work for in Mexico, and we do it with a truly flexible environment, high impact projects in Agile environments, a culture focused on results, training and strong support to grow your career.
What we offer
Global Relocation - (Relocation options; Experience in an international environment; Cross-cultural experience)
Recognition and Evaluation - (Feedback culture; Regular appraisals)
Time Off - (Annual holiday - 20 or 26 days. The duration of the leave depends on the overall seniority; Occasional leave - 1 or 2 days/ depending on the circumstances; Child care leave - 2 days or 16 hours per year; Absence due to force majeure - 2 days or 16 hours per year; Maternity Leave - 20 weeks; Parental Leave - 41 weeks; Paternity Leave - 14 days)
Luxoft Training Center - (Expert-led tech courses covering basic to advanced topics; Internal instructor-led soft skills courses; Comprehensive in-house self-learning resources for both soft and hard skills; Access to external self-learning libraries like ProQuest eBook and Udemy for Business; Cloud Programs: MS Cloud Academy, AWS Partner Academy, Google Cloud Academy; Custom Learning Programs: upskilling, reskilling, technical mentorship; Leadership Programs for Managers)
Well-being and Work-life Balance - (Multisport card; Possibility to order Multisport card at the corporate rate for family members; LuxGood Program: wellbeing seminars, contests, relaxation sessions, yoga sessions, etc.; One Team Program: Buddy for each New Joiner; seminars, meeting and workplace space to support integration with local community and culture; “Hire me” workshops for partners; Preferential banking offer; Preferential car leasing offer; Cafeteria program discounts for shops, cinema tickets, holiday offers; Luxoft Social Benefit Fund: sport and recreation benefits, the possibility to receive financial support)
Health Care - (Private Healthcare Insurance with unlimited access to specialists; Full dental support; Travel Insurance; Possibility to add private healthcare coverage for family members at the corporate rate; Life insurance at the corporate rate for employees and family members, including payment of the basic package for the employee by the employer; Reimbursement for corrective glasses)
Company Events and Friendly Environment - (Many fun social activities organized by the Luxoft team offline in your city; Online entertainment events for whole company and local team events; A workplace where you’re treated with respect within a multicultural team)
Internal Mobility - (Rotation between projects and accounts; New career opportunities)
Self-Learning Library
CSR Projects
Other
Languages: English: B2 Upper Intermediate
Seniority: Senior
Requirements
Strong experience in Observability, SRE, or Platform Engineering.
Advanced PromQL (or equivalent) expertise, including the ability to identify misleading or incorrect query results.
Hands-on experience designing and implementing SLOs and multi-window burn-rate alerting.
Experience with Grafana provisioning and configuration as code.
Ability to build automation tools and code generators using Python or TypeScript.
Hands-on experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or equivalent. AI-assisted engineering is an expected part of the development workflow.
Strong analytical mindset with a healthy skepticism toward telemetry data.
Experience validating signals through multiple independent sources before operationalizing alerts.
Experience with cloud-native or distributed systems environments.
Groundcover
VictoriaMetrics
OpenTelemetry Collector
AWS EKS
Healthcare or regulated-industry experience
Experience building observability tooling and service health reporting solutions
Domain: US Healthcare (PHI-adjacent)
Okta and observability platform access
AI-assisted engineering is a standard part of the development workflow.
Responsibilities
Design and maintain SLO-based monitoring and alerting solutions.
Create and optimize PromQL queries and multi-window burn-rate alerts.
Build and manage Grafana dashboards using configuration-as-code practices.
Develop automation and monitoring configuration generators in Python or TypeScript.
Contribute to Terraform-based observability infrastructure.
Validate monitoring signals and improve alert quality.
Collaborate with engineering and platform teams to enhance reliability and operational visibility.
Keywords / Skills