pracaon.plpracaon.pl

Senior Site Reliability Engineer

Warszawa, Dolnośląskie, Polska, 00-841
Allegro
Partner
8d
Salary to be agreed
Full-time • On-site • IT, Data & AI

Key offer highlights

  • DevOps / Cloud: AWS, Azure, Docker, Kubernetes

  • Backend: Java / .NET / Node / Python

  • Looking for experts - senior/expert

  • Full-time

#Goodtobehere means that:

  • You will join a team you can count on - we work with top-class specialists who have knowledge- and experience-sharing in their DNA.

  • You will love our level of autonomy in team organization, the space for continuous development, and the opportunity to try new things.

  • You get to choose which technology solves the problem and you are responsible for what you create.

  • You will value our Developer Experience and the full platform of tools and technologies that make creating software easier. We rely on an internal ecosystem based on self-service and widely used tools such as Kubernetes, Docker, Consul, GitHub, and GitHub Actions. Thanks to this, you can contribute to Allegro from your very first days on the job.

  • You will be equipped with modern AI tools to automate repetitive tasks, allowing you to focus on developing new services and refining existing ones (also leveraging AI support).

  • You will create solutions that will be used (and loved!) by your friends, family and millions of our customers.

  • You will meet the Allegro Scale, which starts with over 1000 microservices, an open-source data bus (Hermes) with 300K+ rps, a Service Mesh with 1M+ rps, tens of petabytes of data, and production-used machine learning.

  • You will become part of Allegro Tech - We speak at industry conferences, cooperate with tech communities, run our own blog (it's been over 10 years!), record podcasts, lead guilds, and we organize our own internal conference - the Allegro Tech Meeting. We create solutions we love (and can) to talk about!

  • Send us your CV and... see you at Allegro!

We are looking for people with:

  • Solid Engineering Background with a deep understanding of distributed systems and microservices architecture.

  • Proficiency in at least one programming language used for infrastructure automation, tooling, and backend services (Kotlin, Java, Python or Go).

  • Hands-on, advanced experience managing containerized environments and orchestration tools, specifically Kubernetes and Docker, alongside networking and service mesh concepts.

  • Hands-on experience with Chaos Engineering practices (e.g., fault injection, failure scenarios) and performance testing frameworks (practical knowledge of Gatling or similar tooling is a plus).

  • Deep experience configuring monitoring and observability stacks (e.g., Prometheus, Grafana, ELK), paired with a proven ability to identify single points of failure in complex systems and act as an Incident Commander during major outages.

  • Strong practical knowledge of building and maintaining CI/CD pipelines and managing infrastructure using IaC tools.

  • A strong sense of ownership and a data-driven mindset, with a track record of mentoring less experienced engineers, setting code quality standards, and shaping team architecture decisions.

In your daily work, you will handle the following tasks:

  • Leading and coordinating team-level projects to improve system reliability across our massive distributed architecture.

  • Driving the design and implementation of complex reliability solutions and internal platforms, heavily utilizing our core ecosystem.

  • Designing and building an AI agent to revolutionize our incident response workflow, alongside acting as a senior Incident Commander during major outages and leading blameless post-mortems.

  • Developing and scaling our self-service performance testing platform, empowering engineering teams to run large-scale, distributed load tests with real-time observability.

  • Building an automated resiliency platform for continuous fault injection and driving high-stakes, quarterly full Data Center failover experiments.

  • Owning the reliability of the team's systems, managing technical debt, and automating infrastructure.

  • Mentoring less experienced engineers and setting the standard for team code quality.

  • Influencing how the team communicates and makes decisions on difficult system architecture problems at the Allegro scale.

This offer was imported from an external portal.Listing source

More similar listings