Agile Engine is seeking a Senior Site Reliability Engineer to ensure core system administration and operational stability for enterprise on-premise and SaaS-hosted systems. You will focus on Kubernetes cluster management, monitoring, and observability using Snowflake and Open Telemetry, while participating in on-call rotations and incident response.
You will automate infrastructure tasks with Bash, Python, or Go and support health across ESM and ECP environments.