Senior QA / ML Tester
Remote, PolskaWichtige Merkmale des Angebots
Mind. 5 Jahre Erfahrung
Backend: Java / .NET / Node / Python
Qualitätssicherung: manuelle / automatisierte Tests
Vollzeit
Remote-Arbeit - kein Pendeln
Description
We are looking for a Senior QA / ML Tester to join the AI Platform team and take ownership of quality assurance for the Agent Evaluation Framework built on AWS AgentCore. This role involves designing, implementing, and maintaining a functional test suite that validates the correctness of AI agent evaluation pipelines — covering on-demand integration testing, online sampling accuracy, and multi-evaluator execution — and delivering a feasibility assessment for non-AgentCore runtime evaluation scenarios. This is a hands-on, production-focused role at the intersection of software quality engineering and AI/ML system testing, operating within an Agile delivery team and contributing to the reliability of enterprise-grade agentic AI infrastructure. Responsibilities Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good / known-bad session pair validation, multi-evaluator execution correctness, and edge case handling Develop and maintain integration tests for on-demand evaluation mode, integrated into the CI/CD pipeline with automated execution on each build Validate online mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidence Conduct and document a feasibility assessment for non-AgentCore runtime evaluation: analyze alternative runtimes, define evaluation methodology, and deliver a structured findings report Test OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrity Collaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gates Maintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence/Jira Participate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectives Contribute to EngX practices such as code review of test scripts, CI/CD pipeline integration, and test coverage reporting Requirements 5+ years of production experience in QA automation or ML/AI system testing Proven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validation Proficiency in Python test automation, including pytest, fixtures, parametrize, and mocking (unittest.mock, moto) Knowledge of AWS AgentCore Evaluation API, covering on-demand and online evaluation modes Familiarity with OpenTelemetry / ADOT for trace-based evaluation input testing and trace structure validation Skills in REST API testing, including request/response validation and authentication (SigV4, bearer tokens) Experience with CI/CD integration using GitHub Actions, Jenkins, or equivalent, including test pipeline configuration Background in test data management, including known-good / known-bad session pair design and synthetic trace generation Expertise in functional and integration test design for AI/ML evaluation pipelines Competency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysis Understanding of LLM/agent evaluation concepts, such as correctness scoring, sampling strategies, and evaluator chaining Ability to work independently after onboarding, manage own tasks, report status, and escalate blockers proactively Strong analytical skills to define test scenarios from ambiguous or evolving specifications English B2+ level, written and verbal, for daily collaboration with distributed teams Nice to have Hands-on experience with AWS AgentCore Evaluation API or AWS Bedrock testing Experience testing OpenTelemetry / distributed tracing pipelines Familiarity with multi-evaluator execution patterns and correctness validation strategies Experience writing feasibility assessments or technical reports for stakeholders Knowledge of agentic AI frameworks (LangGraph, Strands Agents) sufficient to understand evaluation contracts ISTQB CT-AI certification or equivalent AI testing qualification Experience with AI Ready / AI Practitioner practices at EPAM (prompt engineering, AI-assisted test design)
Requirements
5+ years of production experience in QA automation or ML/AI system testing
Proven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validation
Proficiency in Python test automation, including pytest, fixtures, parametrize, and mocking (unittest.mock, moto)
Knowledge of AWS AgentCore Evaluation API, covering on-demand and online evaluation modes
Familiarity with OpenTelemetry / ADOT for trace-based evaluation input testing and trace structure validation
Skills in REST API testing, including request/response validation and authentication (SigV4, bearer tokens)
Experience with CI/CD integration using GitHub Actions, Jenkins, or equivalent, including test pipeline configuration
Background in test data management, including known-good / known-bad session pair design and synthetic trace generation
Expertise in functional and integration test design for AI/ML evaluation pipelines
Competency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysis
Understanding of LLM/agent evaluation concepts, such as correctness scoring, sampling strategies, and evaluator chaining
Ability to work independently after onboarding, manage own tasks, report status, and escalate blockers proactively
Strong analytical skills to define test scenarios from ambiguous or evolving specifications
English B2+ level, written and verbal, for daily collaboration with distributed teams
Responsibilities
Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good / known-bad session pair validation, multi-evaluator execution correctness, and edge case handling
Develop and maintain integration tests for on-demand evaluation mode, integrated into the CI/CD pipeline with automated execution on each build
Validate online mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidence
Conduct and document a feasibility assessment for non-AgentCore runtime evaluation: analyze alternative runtimes, define evaluation methodology, and deliver a structured findings report
Test OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrity
Collaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gates
Maintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence/Jira
Participate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectives
Contribute to EngX practices such as code review of test scripts, CI/CD pipeline integration, and test coverage reporting
Seniority
Senior
Nice to have
Hands-on experience with AWS AgentCore Evaluation API or AWS Bedrock testing
Experience testing OpenTelemetry / distributed tracing pipelines
Familiarity with multi-evaluator execution patterns and correctness validation strategies
Experience writing feasibility assessments or technical reports for stakeholders
Knowledge of agentic AI frameworks (LangGraph, Strands Agents) sufficient to understand evaluation contracts
ISTQB CT-AI certification or equivalent AI testing qualification
Experience with AI Ready / AI Practitioner practices at EPAM (prompt engineering, AI-assisted test design)
Stichwörter / Fähigkeiten