Site Reliability Engineer Lead (Immediate Joiner)
Marketing & Communications, Software Engineering
Pune, Maharashtra, India
Experience: 7 to 10 Years
Location: Pune, Kharadi
Work Mode: 5 Days Work from Office
Joining: Immediate Joiners Only
We are looking for a highly skilled Senior/Lead DevOps Engineer with strong Site Reliability Engineering (SRE) experience to join our growing engineering team.
The ideal candidate should have strong hands-on expertise in AWS, Kubernetes, Infrastructure as Code, CI/CD automation, and production reliability engineering.
This role requires an engineer who can design, automate, deploy, monitor, troubleshoot, and optimize cloud-native environments while ensuring high availability, scalability, security, and operational excellence.
Key Responsibilities
- Design, deploy, and manage highly available cloud infrastructure on AWS.
- Build, manage, and troubleshoot production-grade Kubernetes environments.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation.
- Develop, maintain, and optimize CI/CD pipelines using Jenkins.
- Drive automation across infrastructure provisioning, deployments, monitoring, and incident management.
- Work closely with engineering teams to improve system reliability, performance, and scalability.
- Handle production incidents, perform root cause analysis, and drive post-incident improvements.
- Implement and maintain monitoring, alerting, logging, and observability solutions.
- Optimize cloud infrastructure for cost, security, performance, and reliability.
- Support containerized applications and microservices architectures.
- Maintain and improve deployment workflows using Git and Bitbucket.
- Contribute to SRE best practices, including SLIs, SLOs, and Error Budgets.
Mandatory Skills
- 7 to 10 years of hands-on DevOps / SRE experience.
- Strong hands-on experience with AWS.
- Production-grade Kubernetes (K8s) experience.
- Strong expertise in Terraform.
- Hands-on experience with AWS CloudFormation.
- Strong experience creating and managing Jenkins pipelines.
- Strong understanding of CI/CD implementation and automation.
- Strong Linux administration and troubleshooting skills.
- Proficiency in Shell Scripting, Bash, or Python.
- Experience with monitoring and logging tools such as Prometheus, Grafana, ELK, CloudWatch, or similar tools.
- Experience managing production environments and handling critical incidents.
Preferred Skills
- Strong Site Reliability Engineering experience.
- Bitbucket administration and pipeline integration.
- Docker and containerization technologies.
- Exposure to Azure.
- Security, compliance, and infrastructure governance.
- Experience working in Agile environments.
Ideal Candidate Profile
- Strong hands-on engineer who enjoys solving complex infrastructure and reliability challenges.
- Individual contributor, not a people-management role.
- Experience supporting large-scale production environments.
- Strong understanding of reliability engineering and operational excellence.
- Startup or product-based company experience is preferred.
- Excellent troubleshooting and problem-solving skills.
- Strong communication and stakeholder management skills.
- Comfortable taking end-to-end ownership of infrastructure, deployments, and production reliability.
Nice to Have
- AWS certifications.
- Kubernetes certifications.
- Azure exposure.
- Experience in healthcare, fintech, SaaS, or other product-based organizations.