pracaon.plpracaon.pl

Senior Platform SRE

Kraków - Poland, Polska
1d
Wynagrodzenie do ustalenia
Pełny etat • Hybrydowa • IT, Data i AI

Najważniejsze cechy oferty

  • DevOps / Cloud: AWS, Azure, Docker, Kubernetes

  • Backend: Java / .NET / Node / Python

  • Model hybrydowy - część pracy zdalnie

  • Szukamy ekspertów - senior/ekspert

  • Pełny etat

Description

IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions – built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated. What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses.

Core Competencies

  • Systems thinking approach to problem-solving

  • Excellent communication skills for cross-functional collaboration and technical enablement

  • Ability to balance hands-on development work with operational responsibilities

  • Strong bias toward automation and eliminating manual toil

How we work

  • Lead and Inspire: Drives trust, alignment, and enthusiasm

  • Think Big: Focus on the problems that most impact commercial outcomes

  • Champion the client: Understand and prioritise client's needs

  • Deliver at pace: Push for fast, sustainable growth;

  • Raise the bar: Take ownership, be accountable and share feedback

Set and uphold standards

  • Author and evolve the SRE standards that underpin the Guild: SLO methodology, error budget policy, observability instrumentation guide, and Production Readiness Review (PRR) checklist

  • Mentor developers on reliability patterns including circuit breakers, retry logic, and fault tolerance

  • Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per-service.

  • Assist and guide teams in system design, capacity planning, architectural reviews and closing observability gaps.

Essential Technical Skills

  • Observability and instrumentation: hands-on OpenTelemetry experience (spans, metrics, traces, context propagation) and production use of Honeycomb, Datadog, Dynatrace, or Grafana; able to instrument Java or Python services directly.

  • SLOs and error budgets: proven track record designing customer-meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and working with development teams on reliability measurement

  • CI/CD and release engineering: experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration

  • Container orchestration: Kubernetes (EKS, AKS, or GKE) required; HashiCorp Nomad is a strong advantage on IG’s hybrid estate; solid understanding of cloud networking and IaC (Terraform preferred)

  • Software engineering: production-quality coding in Java and/or Python; comfortable contributing to application codebases to implement reliability patterns, not just configuring infrastructure around them

  • Distributed systems: strong understanding of how large-scale systems fail and how to make them fail safely; circuit breakers, bulkheads, idempotency, graceful degradation, and load-shedding; high-throughput, low-latency environments preferred

  • Incident management: on-call experience on production systems, blameless PIR facilitation, contributing-factor analysis, and driving action items to closure; PagerDuty and ServiceNow familiarity helpful

  • Chaos engineering: experience designing and executing hypothesis-driven experiments with blast-radius controls and gap-to-impact-tolerance analysis; AWS FIS, Gremlin, or equivalent

  • Community and standards: at ease in a guild or community-of-practice model; comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt

Experience Requirements

  • Track record in high-throughput, production environments (financial services, trading platforms, or similar mission-critical systems preferred)

  • Demonstrated ability to improve system reliability and performance at scale

  • Experience working collaboratively with development teams to implement observability and reliability improvements

  • Strong troubleshooting skills in distributed systems environments

Own incident response and learning

  • Facilitate blameless post-incident reviews (PIRs) within five working days using contributing-factor methodology

  • Maintain the Lessons Register, track remediation actions to closure, and surface patterns across incidents quarterly

Build and own the reliability platform

  • Implement comprehensive monitoring and observability using OpenTelemetry and distributed tracing. Maintain SLO, error budgets and burn-rate tracking

  • Establish and maintain 24/7 operational readiness including automated deployments, blue/green releases, and zero-downtime patching strategies

  • Engineer self-healing capabilities: auto-remediation, error-budget-gated rollback, and automated traffic rerouting

  • Design and run chaos experiments across the AWS estate, turning severe-but-plausible failure scenarios into engineering improvements

  • Build automation tools and CI/CD pipelines that embed reliability practices, while applying software engineering discipline including version control, code reviews, and testing.

  • Contribute to the SRE AI agent, IG’s agentic tooling for incident investigation and reliability review, built on AWS frontier models

  • Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline

Oferta została zaimportowana ze źródła zewnętrznego. Sprawdź link, aby przejść do oryginalnego ogłoszenia.Źródło ogłoszenia

Więcej podobnych ofert