Summary: We are seeking an experienced Senior DevOps Engineer with over 8 years in the field to lead the design, implementation, and management of infrastructure and deployment pipelines across hybrid environments, including on-premises data centers and the cloud. This role requires deep expertise in AWS, Azure, Kubernetes, Infrastructure as Code (IaC), CI/CD, and monitoring tools. The engineer will collaborate closely with development, QA, and operations teams to ensure the delivery of reliable, scalable, and secure enterprise-grade applications. Additionally, the role involves mentoring junior team members and leading automation initiatives.
Responsibilities:
- Manage and support hybrid infrastructure, including physical data centers, AWS, and Azure services.
- Design and implement Infrastructure-as-Code using tools like Terraform, Pulumi, and CloudFormation.
- Configure and maintain Kubernetes clusters, ensuring scalability and high availability.
- Build, maintain, and optimize CI/CD pipelines using Jenkins, GitLab CI, Bitbucket, and GitHub.
- Automate deployment, scaling, and monitoring processes.
- Set up and maintain observability using Grafana, Prometheus, and other monitoring tools.
- Perform root cause analysis of performance issues and participate in on-call support.
- Manage SSL/TLS and code signing certificates and implement infrastructure security best practices.
- Support compliance and audit requirements through documentation and monitoring.
- Collaborate with cross-functional teams to improve infrastructure reliability and delivery processes.
- Document processes in Confluence, track work in Jira, and contribute to knowledge sharing.
- Write root cause analysis and incident post-mortems to prevent future issues.
Required Skills & Experience:
- 8+ years of experience in DevOps, Site Reliability Engineering, or related roles.
- Strong expertise with AWS and Azure services.
- Experience with Infrastructure-as-Code tools and GitOps.
- Proven experience with Kubernetes/EKS setup and operations.
- Proficiency in CI/CD pipelines and GitOps tools.
- Strong knowledge of monitoring and observability tools and incident management.
- Scripting and automation skills in Python, TypeScript, and Bash/shell.
- Understanding of networking, Linux/Windows Server administration, and troubleshooting.
- Experience with data streaming, CDC pipelines, and administering relational and in-memory data stores.
- Configuration management and server automation experience.
- Knowledge of security best practices and familiarity with compliance/audit frameworks.
Soft Skills:
- Strong problem-solving and root cause analysis capabilities.
- Excellent communication and documentation skills.
- Ability to collaborate in team meetings and contribute to on-call rotations.
- Mentoring and knowledge-sharing mindset.