AgileEngine is seeking a Senior Site Reliability Engineer to strengthen the reliability of our enterprise on‑premise and SaaS platforms. You will focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry, and participate in on‑call rotations and incident RCA.
You will automate infrastructure tasks with Python, Bash, or Go, scale containerized environments, and ensure health across ESM and ECP environments.