site reliability engineer
4 dni temu
Warszawa, Województwo mazowieckie, Polska
Enfint
Pełny etat
180 000 zł - 280 000 zł/rok
Bezpłatnie za pośrednictwem poczty elektronicznej lub Google
Zapisz tę ofertę pracy i uporządkuj wyszukiwanie
Utwórz bezpłatne konto, aby zapisywać oferty pracy, tworzyć alerty i wracać do tego ogłoszenia ze swojego pulpitu nawigacyjnego.
Bezpłatnie za pośrednictwem poczty elektronicznej lub Google
Kontynuując, akceptujesz nasze Warunki & Politykę prywatności.
Описание
EPAM develops enterprise software products, open source solutions, and accelerators, including solutions for production AI systems.
Задачи
- Own end-to-end deployment, including infrastructure as code, CI/CD pipelines, and environment management on Azure
- Build LLM-aware observability with model-call traces, production quality signals, drift detection, guardrail-trigger rates, and cost and latency dashboards
- Define and defend SLOs for availability, latency, and quality objectives, balancing delivery speed against stability with data‑driven error budgets
- Run incident management, including on‑call models, pre‑incident runbooks, and blameless postmortems
- Manage intelligence costs by monitoring token economics, configuring budgets and alerts, and conducting proactive capacity planning
- Maintain the security posture through patching, secret rotation, access reviews, and audit readiness across the solution lifecycle
- Shape operability requirements before handover and ensure their inclusion in the pod's definition of done
- Run hypercare jointly and sign off on what will be operated
- Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect
Требования
- 5+ Years of experience operating cloud production systems, including scaling, defining SLOs, managing on-call rotations, and automating manual work
- Expertise in Azure IaaS/PaaS operations, infrastructure as code with Terraform/Bicep, and CI/CD tooling
- Proficiency with observability stacks, LLM tracing, container orchestration, and Python/Bash automation
- Knowledge of FinOps basics for AI workloads
- Familiarity with daily AI use in operations, including incident triage, runbook drafting, log analysis, and automation code
Условия
No conditions specified