We are looking for a Senior Infrastructure Engineer to build and operate enterprise-scale AI infrastructure platforms. The role focuses on Kubernetes environments, Rancher administration, Site Reliability Engineering (SRE), monitoring systems, and AI platform deployment for large organizations and government entities.
The successful candidate will help establish highly available, secure, and scalable environments supporting advanced AI workloads.
Key Responsibilities
Design and manage large-scale Kubernetes environments.
Administer and optimize Rancher-managed clusters.
Implement monitoring, observability, and alerting systems.
Support CI/CD pipelines and DevOps automation initiatives.
Maintain reliability, scalability, and security of AI infrastructure.
Collaborate with engineering teams on AI platform deployment and operations.
Develop operational standards and SRE best practices.
Support on-premise AI environments for enterprise and government customers.
Ensure infrastructure compliance with security and governance requirements.
Requirements
5+ years of Infrastructure Engineering, DevOps, Platform Engineering, or SRE experience.
Strong experience with:
Kubernetes
Rancher
CI/CD pipelines
Monitoring & Observability tools
Infrastructure automation
Experience supporting AI platforms and machine learning environments.
Knowledge of enterprise or government-grade infrastructure environments.
Strong understanding of security, resilience, and operational excellence principles.
Preferred
Experience supporting GenAI or Agentic AI platforms.
Experience with on-premise GPU infrastructure.
Experience with cloud-native architectures.