Agile Engine is seeking a Senior Site Reliability Engineer to ensure stability across on-premise and SaaS-hosted systems, with a focus on Kubernetes cluster management, monitoring, and observability using Snowflake and Open Telemetry.
You will participate in on-call rotations, incident response, and root-cause analysis, while automating infrastructure tasks with Python, Bash, or Go to support ESM and ECP environments.