site reliability engineer

4 dni temu

Warszawa, Województwo mazowieckie, Polska Enfint Pełny etat 180 000 zł - 280 000 zł/rok

Описание

EPAM develops enterprise software products, open source solutions, and accelerators, including solutions for production AI systems.

Задачи

  • Own end-to-end deployment, including infrastructure as code, CI/CD pipelines, and environment management on Azure
  • Build LLM-aware observability with model-call traces, production quality signals, drift detection, guardrail-trigger rates, and cost and latency dashboards
  • Define and defend SLOs for availability, latency, and quality objectives, balancing delivery speed against stability with data‑driven error budgets
  • Run incident management, including on‑call models, pre‑incident runbooks, and blameless postmortems
  • Manage intelligence costs by monitoring token economics, configuring budgets and alerts, and conducting proactive capacity planning
  • Maintain the security posture through patching, secret rotation, access reviews, and audit readiness across the solution lifecycle
  • Shape operability requirements before handover and ensure their inclusion in the pod's definition of done
  • Run hypercare jointly and sign off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect

Требования

  • 5+ Years of experience operating cloud production systems, including scaling, defining SLOs, managing on-call rotations, and automating manual work
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code with Terraform/Bicep, and CI/CD tooling
  • Proficiency with observability stacks, LLM tracing, container orchestration, and Python/Bash automation
  • Knowledge of FinOps basics for AI workloads
  • Familiarity with daily AI use in operations, including incident triage, runbook drafting, log analysis, and automation code

Условия

No conditions specified