Site Reliability Engineer — SRE
PolskaKey offer highlights
Employment: contract
Remote work - no commuting
DevOps / Cloud: AWS, Azure, Docker, Kubernetes
Full-time
Description
Your new company You will join a fast-growing technology company building an AI-powered real-time customer interaction platform. The productoperates in a live environment where uptime, latency and reliability are central to the customer experience. Your new role Your responsibilities will include: Defining and managing SLOs, error budgets and production reliability standards. Improving observability across metrics, logs, traces and latency indicators. Leading incident response, postmortems and reliability improvements. Automating operational tasks using Go or Python. Supporting capacity planning, load testing and production scaling. Partnering with engineering teams on safe deployments, rollback strategies and platform resilience. What you'll need to succeed Strong experience with SLOs, error budgets and incident management. Strong Go or Python skills for automation and reliability engineering. Experience with observability tools, monitoring, alerting, logs and traces. Understanding of event-driven or real-time systems such as NATS, Kafka, WebSockets or streaming pipelines. Good networking knowledge, including TCP/IP, TLS, DNS and load balancing. Experience with cloud platforms, ideally GCP. Calm incident communication style, strong ownership and a continuous improvement mindset. What you'll get in return Contract of Employment 100% remote work Chance to work in an environment that supports autonomy and ownership What you need to do now If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV, or call us now. Hays Poland sp. z o.o. is an employment agency registered in a registry kept by Marshal of the Mazowieckie Voivodeship under the number 361.
Płaca
33000-38000 PLN gross
Responsibilities
As a Site Reliability Engineer, you will take ownership of production reliability for a real-time platform. You will work across SLOs, observability, incident response, capacity planning and automation to ensure the platform remains scalable, resilient and performant.
Informacje dodatkowe
Nr ref.: 1200945
Lokalizacja: Polska
Rodzaj pracy: Stała