30 sep
|
EPAM Systems
|
Argentina
30 sep
EPAM Systems
Argentina
We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.
Responsibilities
- Provide on-call support for Java backend identity services during business hours
- Troubleshoot complex production issues using logs and telemetry to identify root causes
- Prepare and deploy patches to address issues in cloud infrastructure
- Implement reliability improvements for key identity services through practical code and configuration changes
- Build and refine metrics and dashboards to enable rapid assessment of platform health
- Monitor SLOs across backend services and drive remediation when error rates increase
- Create and improve runbooks to standardize operational responses across services
Requirements
- 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
- Strong experience with Amazon Web Services in production environments
- Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
- Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
- Hands-on experience with Git-based workflows
- Hands-on experience with Gradle in Java service environments
- Leadership skills to guide reliability improvements and support operational decision-making
- Incident response skills to communicate operational issues clearly and concisely in writing
- Fast learning ability to absorb information quickly and apply it during on-call support
- SLO management skills to track, evaluate, and improve reliability through repeatable processes
- English proficiency: B2 (Upper-Intermediate)
Nice to have
- Kubernetes
- Terraform
- Grafana
- Apache Kafka
- New Relic
📌 Lead Site Reliability Engineer (Argentina)
🏢 EPAM Systems
📍 Argentina