Site Reliability Engineer (SRE)
Techdome · Remote
Experience: 3-6 Years
Site Reliability Engineer (SRE) | DevOps Engineer Location: Hyderabad / Indore Experience: 3–6 Years About the Role Techdome is hiring a Site Reliability Engineer (SRE) to build, operate, and continuously improve highly available, secure, and scalable cloud infrastructure across our healthcare, fintech, AI, and SaaS products. This role goes beyond traditional DevOps. You'll own production environments, improve system reliability, automate operations, build resilient deployment pipelines, manage incidents, and ensure seamless releases using strategies such as Blue-Green Deployments, Rolling Deployments, and Zero-Downtime Releases. If you're passionate about automation, cloud infrastructure, Kubernetes, observability, and AI-powered operations, we'd love to hear from you. Key Responsibilities Maintain the availability, reliability, scalability, and performance of production systems. Manage and optimize production environments across cloud platforms. Design and automate deployment pipelines using CI/CD best practices. Implement Blue-Green, Rolling, and Zero-Downtime deployment strategies. Build Infrastructure as Code using Terraform and Ansible. Implement observability using Prometheus, Grafana, ELK, Datadog, OpenTelemetry, and centralized logging. Define and maintain SLIs, SLOs, and Error Budgets. Lead production incident management, Root Cause Analysis (RCA), and post-incident reviews. Perform cloud cost optimization and capacity planning. Automate operational workflows using scripting and AI-powered tooling. Participate in on-call rotations and production support. Required Skills 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or Cloud Engineer. Hands-on experience with AWS, Azure, or GCP. Strong experience with Docker and Kubernetes. Expertise in Terraform, Ansible, or other Infrastructure as Code tools. Experience building CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar. Strong Linux administration, n