Senior Cloud Native Developer
Remote, PolskaWichtige Merkmale des Angebots
Mind. 3 Jahre Erfahrung
Backend: Java / .NET / Node / Python
Hybridmodell - teilweise remote
DevOps / Cloud: AWS, Azure, Docker, Kubernetes
Vollzeit
Description
We are looking for a Senior Cloud Native Developer to join our team in building an Enterprise Agent Development Platform — a production-grade, cloud-native ecosystem that enables engineering teams to define, orchestrate, deploy, and observe AI agents at scale. The platform standardizes agent development across the organization using LangGraph and Strands Agents on AWS AgentCore Runtime, providing a single-file developer experience where the platform handles runtime, tooling, observability, and deployment automatically. The initiative spans agent framework design, runtime architecture, marketplace integration, CI/CD automation, and enterprise-grade observability — reducing agent development from months to days while enforcing consistent security, quality, and governance standards. Responsibilities Define and implement enterprise-wide telemetry standards, monitoring architecture, and observability frameworks for AI agents, MCP servers, and runtime platforms Design and maintain OpenTelemetry-based instrumentation across agent frameworks and runtime environments Architect collector pipelines, processor configurations, and multi-exporter setups using AWS Distro for OpenTelemetry (ADOT) Build and optimize monitoring, alerting, and distributed tracing solutions using AWS CloudWatch and X-Ray at scale Implement enterprise-grade APM, custom dashboards, and NRQL-based alerting within New Relic Establish GenAI SIG semantic conventions for AI/LLM workloads and drive adoption across engineering teams Develop cross-runtime observability architectures to consolidate multi-vendor telemetry Instrument agent and LLM workflows using tools such as LangSmith or MLflow tracing Translate enterprise requirements into actionable engineering plans and provide technical leadership across observability initiatives Collaborate with platform, agent framework, and CI/CD teams to embed observability into the developer experience Requirements 3+ years of experience in observability or platform engineering with hands-on OpenTelemetry implementation Expertise in OpenTelemetry architecture, including SDK, Collector, and semantic conventions (including GenAI SIG conventions for AI workloads) Proficiency in AWS Distro for OpenTelemetry (ADOT), including collector pipeline design, processor configuration, and multi-exporter setup Skills in AWS CloudWatch (Logs Insights, EMF, composite alarms, cross-account aggregation) and X-Ray at scale (distributed tracing, service maps, trace groups, sampling rules) Background in New Relic in enterprise environments, including APM, custom dashboards, alerting, distributed tracing, and NRQL Knowledge of OpenTelemetry GenAI SIG semantic conventions or AI/LLM observability implementation Familiarity with agent/LLM observability tools such as LangSmith, MLflow tracing, or equivalents Understanding of cross-runtime observability architecture and multi-vendor telemetry consolidation patterns Proficiency in Python for ADOT instrumentation, SDK development, and integration with agent frameworks Capability to provide technical leadership and translate enterprise requirements into actionable engineering plans Excellent command of written and spoken English (B2+ level) Nice to have Experience in cross-team coordination on observability standards across multiple engineering teams Familiarity with AWS AgentCore Observability Background in Datadog or Dynatrace in hybrid environments Skills in Strands or LangGraph agent framework instrumentation
Requirements
3+ years of experience in observability or platform engineering with hands-on OpenTelemetry implementation
Expertise in OpenTelemetry architecture, including SDK, Collector, and semantic conventions (including GenAI SIG conventions for AI workloads)
Proficiency in AWS Distro for OpenTelemetry (ADOT), including collector pipeline design, processor configuration, and multi-exporter setup
Skills in AWS CloudWatch (Logs Insights, EMF, composite alarms, cross-account aggregation) and X-Ray at scale (distributed tracing, service maps, trace groups, sampling rules)
Background in New Relic in enterprise environments, including APM, custom dashboards, alerting, distributed tracing, and NRQL
Knowledge of OpenTelemetry GenAI SIG semantic conventions or AI/LLM observability implementation
Familiarity with agent/LLM observability tools such as LangSmith, MLflow tracing, or equivalents
Understanding of cross-runtime observability architecture and multi-vendor telemetry consolidation patterns
Proficiency in Python for ADOT instrumentation, SDK development, and integration with agent frameworks
Capability to provide technical leadership and translate enterprise requirements into actionable engineering plans
Excellent command of written and spoken English (B2+ level)
Responsibilities
Define and implement enterprise-wide telemetry standards, monitoring architecture, and observability frameworks for AI agents, MCP servers, and runtime platforms
Design and maintain OpenTelemetry-based instrumentation across agent frameworks and runtime environments
Architect collector pipelines, processor configurations, and multi-exporter setups using AWS Distro for OpenTelemetry (ADOT)
Build and optimize monitoring, alerting, and distributed tracing solutions using AWS CloudWatch and X-Ray at scale
Implement enterprise-grade APM, custom dashboards, and NRQL-based alerting within New Relic
Establish GenAI SIG semantic conventions for AI/LLM workloads and drive adoption across engineering teams
Develop cross-runtime observability architectures to consolidate multi-vendor telemetry
Instrument agent and LLM workflows using tools such as LangSmith or MLflow tracing
Translate enterprise requirements into actionable engineering plans and provide technical leadership across observability initiatives
Collaborate with platform, agent framework, and CI/CD teams to embed observability into the developer experience
Seniority
Senior
Nice to have
Experience in cross-team coordination on observability standards across multiple engineering teams
Familiarity with AWS AgentCore Observability
Background in Datadog or Dynatrace in hybrid environments
Skills in Strands or LangGraph agent framework instrumentation
Stichwörter / Fähigkeiten