Share this job
Senior Software Engineer - SRE (JR1016)
Apply for this job

Site Reliability Engineer

Position Overview

Our client is seeking experienced Site Reliability Engineers (SREs) to build, operate, and continuously improve mission-critical, production-grade systems. This role is ideal for engineers who take ownership of infrastructure, prioritize reliability and security, and are passionate about preventing incidents rather than simply responding to them.

The Site Reliability Engineer will own highly available cloud infrastructure end to end, operate production Kubernetes environments, strengthen observability and reliability practices, and develop automated CI/CD systems that enable safe, fast, and repeatable deployments.

This is an ownership-heavy position for an engineer who thrives in complex environments, enjoys solving difficult infrastructure problems, and understands that reliability is an engineering discipline—not simply an operations function.

Key Responsibilities

  • Own Cloud Infrastructure: Design, deploy, operate, and continuously improve highly available and scalable AWS infrastructure across production environments.
  • Operate Kubernetes Platforms: Manage and maintain production EKS clusters, including upgrades, scaling, networking, security, workload reliability, and performance optimization.
  • Drive Reliability Engineering: Establish and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), reliability metrics, and operational standards.
  • Build Observability: Develop comprehensive monitoring, logging, tracing, alerting, and application performance monitoring capabilities to proactively identify and resolve reliability issues.
  • Develop Infrastructure as Code: Build and maintain scalable, reusable infrastructure using Terraform while applying strong engineering and security practices.
  • Automate Operations: Develop production-quality automation and tooling in Python, Go, or similar languages to eliminate manual processes and improve operational efficiency.
  • Build CI/CD Systems: Design, maintain, and optimize automated deployment pipelines that enable safe, fast, consistent, and repeatable software releases.
  • Implement GitOps Practices: Leverage tools such as GitHub Actions and ArgoCD to establish reliable, auditable, and automated deployment workflows.
  • Troubleshoot Complex Systems: Diagnose and resolve multi-layer infrastructure and application issues spanning Kubernetes, networking, compute, storage, security, and application layers.
  • Improve Incident Prevention: Analyze system behavior, operational data, and past incidents to identify systemic weaknesses and implement preventative improvements.
  • Support Security & Compliance: Build infrastructure and operational processes that meet rigorous security, availability, and compliance requirements, particularly in regulated or public-sector environments.
  • Collaborate Across Engineering: Partner with software engineers, security teams, platform teams, and other technical stakeholders to improve system reliability and deployment practices.
  • Participate in Incident Response: Provide technical leadership during critical incidents, restore service quickly, and drive thorough post-incident analysis and remediation.

Environment & Work Style

This is a highly technical, ownership-oriented engineering environment where reliability, automation, and operational excellence are core priorities. The successful candidate will have significant responsibility for production infrastructure and will be expected to make sound technical decisions while balancing availability, security, performance, and deployment velocity.

Engineers are encouraged to identify opportunities for automation and preventative improvements rather than relying on manual operational processes or repeatedly responding to the same incidents.

Essential Qualifications

  • 5+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps, or a closely related discipline.
  • Deep, hands-on expertise with Amazon Web Services (AWS), including networking, compute, IAM, scaling, and security.
  • Strong experience with Terraform and infrastructure as code at scale.
  • Very strong Kubernetes fundamentals with hands-on experience operating Amazon EKS in production.
  • Proven ability to troubleshoot complex, multi-layer infrastructure and application issues.
  • Production-quality programming or scripting experience in Python, Go, or a similar language.
  • Strong automation mindset with a track record of eliminating manual operational processes.
  • Experience designing and maintaining CI/CD pipelines.
  • Hands-on experience with GitHub Actions, ArgoCD, or comparable CI/CD and GitOps technologies.
  • Strong understanding of observability practices, including metrics, logs, traces, APM, and alerting.
  • Experience working with Datadog or a comparable observability platform.
  • Understanding and practical application of SLIs, SLOs, and reliability engineering principles.

Preferred Qualifications

  • Experience building and operating infrastructure for highly regulated or security-sensitive environments.
  • Experience supporting GovCloud, FedRAMP, or other public-sector compliance requirements.
  • Strong understanding of AWS networking, including VPCs, routing, load balancing, DNS, and network security.
  • Experience designing highly available and fault-tolerant distributed systems.
  • Experience implementing Kubernetes security, networking, scaling, and cluster lifecycle management.
  • Experience with GitOps-based infrastructure and application deployment models.
  • Familiarity with incident management, postmortems, capacity planning, and disaster recovery.
  • Experience establishing reliability standards and operational best practices across engineering organizations.


Apply for this job
Powered by