07 oct
|
Flux It
|
Buenos Aires
07 oct
Flux It
Buenos Aires
AgileEngine is an Inc. **** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.
We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a
Senior Cloud/DevOps Engineer
to operate and improve the cloud infrastructure and reliability layers behind an enterprise data platform in a regulated healthcare environment.
The mandatory requirements are 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering, advanced hands-on experience with
AWS EKS
and Kubernetes, experience administering and troubleshooting
Argo Workflows
, and strong English communication skills.
MUST HAVES
-
5+ years
of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
- Strong hands-on experience operating
AWS infrastructure
in production environments.
- Advanced experience with
Kubernetes and Amazon EKS
, including workload operations, troubleshooting, access, observability, capacity, and reliability.
- Hands-on experience administering and troubleshooting
Argo Workflows
or comparable workflow orchestration platforms.
- Strong Infrastructure as Code experience with
Terraform
and source-controlled infrastructure practices.
- Experience building, hardening, and supporting CI/CD pipelines and production release processes.
- Strong experience with monitoring, logging, alerting, and incident-routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
- Demonstrated ability to lead complex incident resolution, perform root-cause analysis, and translate findings into preventive improvements.
- Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
- Ability to make well-reasoned technical decisions, identify tradeoffs, estimate work,
and drive improvements across a complex platform.
- Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
- Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
- Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on-call rotation.
NICE TO HAVES
- Experience supporting data-platform infrastructure involving Snowflake, dbt, Fivetran, HVR, Tableau Cloud, or custom ingestion pipelines.
- Familiarity with data-specific observability platforms such as SYNQ.
- Experience modernizing or migrating legacy orchestration and ingestion solutions such as Boomi or AWS Data Pipeline.
- Experience with service-management and change-control tools such as Freshservice and Jira.
- Experience operating in healthcare, life sciences, financial services, or another regulated environment.
- Familiarity with HIPAA, GDPR, FDA-related controls, least-privilege access, separation of duties, and audit-ready operational practices.
WHAT YOU WILL DO
- Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day-to-day operations, complex troubleshooting, and L2/L3 escalation.
- Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
- Administer Kubernetes-hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
- Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git-based delivery practices.
- Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
- Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data‐specific observability platforms.
- Lead or support major incident response, root‐cause analysis, post‐incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
- Design and implement reliability improvements such as selective auto‐remediation, dependency‐aware alert correlation, impact analysis, and automation of repetitive operational work.
- Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
- Apply disciplined change‐management, access‐control, secrets‐management, auditability, and documentation practices appropriate for a HIPAA‐, GDPR‐, and FDA‐regulated environment.
- Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge‐transfer materials.
- Mentor Middle‐level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
- Participate in the Cloud / DevOps on‐call rotation for critical incidents outside staffed service hours.
PERKS AND BENEFITS
-
Professional growth
: Mentorship, TechTalks, and personalized growth roadmaps.
-
Competitive compensation
: USD‐based pay with education, fitness, and team activity budgets.
-
Exciting projects
: Modern solutions with Fortune 500 and top product companies.
-
Flextime
: Adaptable schedule with remote and office options.
#J-*****-Ljbffr
📌 Senior Devops Engineer — Remote (Aws/Eks/Terraform) (Buenos Aires)
🏢 Flux It
📍 Buenos Aires