Site Reliability Engineer Lead (Immediate Joiner)

HiLabs
HiLabs

Marketing & Communications, Software Engineering

Pune, Maharashtra, India

Posted on Aug 25, 2026

Experience: 7 to 10 Years

Location: Pune, Kharadi

Work Mode: 5 Days Work from Office

Joining: Immediate Joiners Only

We are looking for a highly skilled Senior/Lead DevOps Engineer with strong Site Reliability Engineering (SRE) experience to join our growing engineering team.

The ideal candidate should have strong hands-on expertise in AWS, Kubernetes, Infrastructure as Code, CI/CD automation, and production reliability engineering.

This role requires an engineer who can design, automate, deploy, monitor, troubleshoot, and optimize cloud-native environments while ensuring high availability, scalability, security, and operational excellence.

Key Responsibilities

  • Design, deploy, and manage highly available cloud infrastructure on AWS.
  • Build, manage, and troubleshoot production-grade Kubernetes environments.
  • Implement Infrastructure as Code using Terraform and AWS CloudFormation.
  • Develop, maintain, and optimize CI/CD pipelines using Jenkins.
  • Drive automation across infrastructure provisioning, deployments, monitoring, and incident management.
  • Work closely with engineering teams to improve system reliability, performance, and scalability.
  • Handle production incidents, perform root cause analysis, and drive post-incident improvements.
  • Implement and maintain monitoring, alerting, logging, and observability solutions.
  • Optimize cloud infrastructure for cost, security, performance, and reliability.
  • Support containerized applications and microservices architectures.
  • Maintain and improve deployment workflows using Git and Bitbucket.
  • Contribute to SRE best practices, including SLIs, SLOs, and Error Budgets.

Mandatory Skills

  • 7 to 10 years of hands-on DevOps / SRE experience.
  • Strong hands-on experience with AWS.
  • Production-grade Kubernetes (K8s) experience.
  • Strong expertise in Terraform.
  • Hands-on experience with AWS CloudFormation.
  • Strong experience creating and managing Jenkins pipelines.
  • Strong understanding of CI/CD implementation and automation.
  • Strong Linux administration and troubleshooting skills.
  • Proficiency in Shell Scripting, Bash, or Python.
  • Experience with monitoring and logging tools such as Prometheus, Grafana, ELK, CloudWatch, or similar tools.
  • Experience managing production environments and handling critical incidents.

Preferred Skills

  • Strong Site Reliability Engineering experience.
  • Bitbucket administration and pipeline integration.
  • Docker and containerization technologies.
  • Exposure to Azure.
  • Security, compliance, and infrastructure governance.
  • Experience working in Agile environments.

Ideal Candidate Profile

  • Strong hands-on engineer who enjoys solving complex infrastructure and reliability challenges.
  • Individual contributor, not a people-management role.
  • Experience supporting large-scale production environments.
  • Strong understanding of reliability engineering and operational excellence.
  • Startup or product-based company experience is preferred.
  • Excellent troubleshooting and problem-solving skills.
  • Strong communication and stakeholder management skills.
  • Comfortable taking end-to-end ownership of infrastructure, deployments, and production reliability.

Nice to Have

  • AWS certifications.
  • Kubernetes certifications.
  • Azure exposure.
  • Experience in healthcare, fintech, SaaS, or other product-based organizations.