Senior Site Reliability Engineer (Buenos Aires)

Senior Site Reliability Engineer (Buenos Aires)

07 sep
|
teladoc
|
Buenos Aires

07 sep

teladoc

Buenos Aires

Join the team leading the next evolution of virtual care.
At Teladoc Health, you are empowered to bring your true self to work while helping millions of people live their healthiest lives.
Here you will be part of a high-performance culture where colleagues embrace challenges, drive transformative solutions, and create opportunities for growth. Together, we're transforming how better health happens.

Position Summary

We are seeking a Senior Site Reliability Engineer with strong Microsoft Azure experience to help operate, scale, and improve the reliability of our critical cloud-based services. This role will anchor an SRE Team that will focus on production reliability, observability, automation, incident response, cloud operations, and continuous improvement.

The Sr. SRE partners with software engineering, product, security, and operations leadership to embed reliability principles into the software delivery lifecycle, mentors the SRE team, and serves as the technical authority for observability and reliability across mission-critical healthcare workloads.

Role and Responsibilities

Reliability Engineering

Define, implement, and improve Service Level Indicators , Service Level Objectives , and error budgets for critical applications and platform services.
Partner with application teams to improve service reliability, fault tolerance, scalability, and operational readiness.
Identify and eliminate recurring reliability issues through root cause analysis, automation, and architectural improvements.
Help design systems that are resilient to Azure region, zone, network, dependency, and deployment failures.
Participate in production readiness reviews for new services, major releases, and infrastructure changes.

Observability and Monitoring

Build and improve observability across applications, infrastructure, networks,



and cloud services. Implement monitoring for the four golden signals of Latency, Traffic, Errors, and Saturation
Develop dashboards, alerts, logs, traces, and metrics using tools such as Azure Monitor, Log Analytics, Elastic/ELK, Grafana, Open Telemetry, Datadog, Dynatrace, New Relic, or similar APM platforms
Create service health dashboards for engineering, operations, and leadership audiences.

Performance, Capacity, and Resilience

Analyze system performance, bottlenecks, saturation trends, and capacity risks.
Improve backup, disaster recovery, failover, and business continuity practices.
Partner with engineering teams to implement resiliency patterns such as retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling.

Azure Cloud Operations

Support and improve production workloads running on Microsoft Azure .
Collaborate with cloud and network teams on secure, scalable Azure architecture.
Help enforce Azure operational standards, including tagging, monitoring, backup, recovery, identity, security, and cost awareness.

Incident Management and Response

Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention plans.
Work with Incident Management, NOC, Help Desk, and application teams to improve response processes and runbooks.

Security and Compliance Support





Work with Security Engineering to ensure production systems follow cloud security and compliance standards.
Support operational controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.

Required Qualifications

7+ years in site reliability, including hands-on ownership of mission-critical services, through a combination of applicable work experience, training, military experience, or education.
Deep Azure experience including Azure Monitor, Application Insights, AKS, and cloud-native operations across hybrid infrastructure.
Proven track record designing and rolling out an SLO program with SLIs, SLOs, and error budget policy in production environments.
Observability Tools: Hands-on experience with enterprise observability platform such as Datadog and Dynatrace, Elastic, Grafana, Prometheus or Logic Monitor.
Hands-on experience implementing and configuring Datadog for monitoring, observability, and alerting.

Preferred Qualifications

Prior experience standing up or anchoring an SRE practice, including operating model, rituals, and adoption across multiple engineering teams.
Strong incident command experience and a track record of running blameless postmortems that drive measurable improvement.
Healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent compliance frameworks.
Multi-cloud reliability experience (AWS) in addition to Azure.
Chaos engineering and resiliency testing experience (e.g., Gremlin, Chaos Mesh, Azure Chaos Studio).
Infrastructure as Code expertise (Terraform, Bicep, Ansible) for observability and remediation automation.
Scripting and programming proficiency (Python, Power Shell, Go) for automation, tooling, and integration work.

Experien

#J-18808-Ljbffr

📌 Senior Site Reliability Engineer (Buenos Aires)
🏢 teladoc
📍 Buenos Aires

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (buenos aires) / buenos aires

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: senior site reliability engineer (buenos aires) / buenos aires