T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
Zapisz tę ofertę pracy i uporządkuj wyszukiwanie
Utwórz bezpłatne konto, aby zapisywać oferty pracy, tworzyć alerty i wracać do tego ogłoszenia ze swojego pulpitu nawigacyjnego.
- Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
- Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
- Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
- Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
- Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
- Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
- Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
- Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
- Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
- Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
- Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
- Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
- Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
- Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
WHAT SKILLS WILL BE APPRECIATED?
5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
Strong experience with Kubernetes and OpenShift administration in production environments.
Proven experience deploying and operating vLLM-based inference platforms in production.
Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
Strong Python programming skills with experience developing automation and operational tooling.
Experience with Bash scripting and Linux systems administration.
Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.
OUR OFFER FOR YOU
Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.
No dress code - you can just be yourself here
Medical, sport and life insurance packages at preferential terms
Access to our products and services at preferential terms
Employment contract-based cooperation
Know Talent - receive training or financial bonus for recommending new employees
WHAT WILL YOUR RECRUITMENT PROCESS BE LIKE?
A fair approach to all people who want to join T Hub means that:
- The recruitment process is transparent;
- Our recruitment decision is based solely on an assessment of your skills (your race, skin color, sexual orientation, gender identity, origin, disability, political view, appearance, or religion will not have any influence on he outcome of the process):
- Regardless of the outcome of the process, you will get detailed feedback.
Let’s meet to better understand each other's expectations.
Our screening meeting will last about 20-30 minutes. We will ask you about our areas