30 jul
|
Teladoc Health
|
Argentina
30 jul
Teladoc Health
Argentina
Summary of Position
We are seeking a highly skilled Site Reliability Engineer (SRE) with deep experience in Azure environments, specializing in Observability, Monitoring, and Incident Response. This role is critical to ensuring the availability, reliability, and performance of our hybrid cloud infrastructure and services.
The idóneo candidate will design and implement observability frameworks, drive automation in monitoring and alerting, and lead effective incident management processes across multi-cloud environments.
This position requires strong technical acumen in cloud-native operations, a proactive mindset toward reliability engineering, and the ability to collaborate with engineering, operations, and security teams to maintain mission-critical healthcare and enterprise workloads.
Essential Duties and Responsibilities
Observability & Monitoring
- Design, implement, and maintain observability solutions across Azure (e.g., Azure Monitor, Datadog, Grafana-Prometheus, Dynatrace, Elastic).
- Define and standardize SLIs/SLOs/SLAs to measure service health and customer experience.
- Develop dashboards and automated alerting to proactively identify service degradations.
Incident Response
- Build and maintain on-call runbooks and playbooks to reduce time-to-resolution.
- Drive post-incident “blameless” retrospectives and continuous improvement initiatives.
Reliability Engineering
- Develop automation for self-healing systems, monitoring remediation, and incident mitigation.
- Contribute to disaster recovery and business continuity planning across multi-cloud platforms.
- Work with engineering teams to design for resiliency, scalability, and reliability from the ground up.
Collaboration & Governance
- Partner with security, network, and system engineering teams to ensure observability integrates with compliance and governance frameworks.
- Advocate for best practices in cloud-native reliability engineering.
- Mentor engineering staff in observability tools, monitoring strategies, and incident management.
The time spent on each responsibility reflects an estimate and is subject to change dependent on business needs.
Supervisory Responsibilities
No
Required Qualifications
- +3 years of experience, or equivalent demonstrated through a combination of work experience, training, military experience, or education for:
- Cloud Platforms: Expertise in Azure (VMs, AKS, Application Insights, Azure Monitor).
- Observability Tools: Hands-on experience with enterprise observability platform such as Datadog and Dynatrace, Elastic, Grafana, Prometheus orLogicMonitor.
- Monitoring & Alerting: Deep understanding of metrics, logs, traces, and distributed system monitoring.
Preferred Qualifications
- Automation & Infrastructure as Code (IaC): Proficiency with Terraform, Bicep, Ansible, or similar tools to automate monitoring and remediation workflows.
- Kubernetes Observability: Knowledge of AKS logging, tracing, and monitoring in containerized environments.
- AI & Observability: Exposure to AI-driven monitoring, anomaly detection, or predictive alerting tools.
- Programming/Scripting: Scripting skills in Python, PowerShell, or similar languages for automation and tool integration.
- Chaos Engineering: Experience with resiliency testing and tools such as Gremlin or Chaos Mesh.
- Healthcare & Compliance: Experience in healthcare IT environments with HIPAA, HITRUST, or other compliance frameworks.
📌 Site Reliability Engineer III (Argentina)
🏢 Teladoc Health
📍 Argentina