03 oct
|
Bonzzu
|
Argentina
? WE’RE HIRING
DevOps Engineer — AI & Streaming Platform
? 100% Remote — Argentina | Mexico | Peru | Uruguay
We are looking for an experienced DevOps Engineer to take real ownership of the production infrastructure behind a large-scale AI and real-time video streaming platform running on AWS .
This is not a build-only DevOps role.
You’ll own the full delivery and production lifecycle — from CI/CD, Infrastructure as Code and staged deployments to observability, incident response and automated remediation .
The platform already operates with a high degree of automation. Deployments are hands-off, and AI agents are used to support incident triage, root cause analysis and remediation workflows. Your mission will be to push that automation even further: build production systems that detect regressions, recover safely and continuously improve with less human intervention.
If you enjoy solving complex production problems, improving deployment reliability and turning repetitive operational work into automation, this role offers significant technical ownership and impact.
? WHAT YOU’LL OWN
As a DevOps Engineer, you will:
• Own infrastructure deployment for AI and streaming services across development, QA and production AWS environments.
• Build and maintain end-to-end CI/CD pipelines using Harness and GitHub Actions , covering build, testing, artifact promotion, staged rollout and rollback.
• Participate in the PagerDuty on-call rotation during LATAM working hours , triaging production incidents, mitigating impact and driving root causes toward permanent solutions.
• Automate operational and release processes including environment provisioning, configuration promotion and health validation .
• Manage production infrastructure through Terraform / CloudFormation and Helm , keeping environments reproducible and consistent.
• Work extensively with Docker and Kubernetes to support scalable production workloads.
• Own and improve observability and alerting using tools such as Splunk, Datadog and CloudWatch.
- Design actionable alerts, reduce operational noise and improve visibility into system health.
• Partner with ML and backend engineers to support deployment paths including SageMaker endpoints,
AWS Lambda and containerized inference services .
- Improve deployment safety, release velocity and recovery time by reducing failed releases and strengthening rollback strategies.
• Build auto-remediation and automated incident response , turning recurring incidents into automated detection and automated fixes.
• Use Python and Bash to build internal tooling and automation rather than relying exclusively on configuration and YAML.
• Use modern AI coding agents such as Claude Code, GitHub Copilot and Cursor as part of the development and operational workflow.
? WHAT WE’RE LOOKING FOR
• 4+ years of professional experience in DevOps, SRE or Production Infrastructure , with real ownership of production environments.
• Strong hands-on experience with AWS , including services and concepts such as IAM, VPC, EC2, ECS/EKS, Lambda, S3, CloudWatch and multi-account environments .
• Strong ownership of CI/CD pipelines end-to-end , covering both build and deployment workflows.
• Professional experience with GitHub Actions and exposure to or experience with Harness .
• Strong experience with Infrastructure as Code , particularly Terraform and/or CloudFormation .
• Experience with Helm, Docker and Kubernetes in production environments.
• Real-world experience handling production incidents , including on-call rotations, incident triage, mitigation and root cause analysis.
• Strong scripting and automation capabilities with Python and Bash / Shell .
• Solid understanding of observability , including dashboards, alert design, SLOs and production monitoring using Splunk, Datadog, CloudWatch or equivalent tools.
- Ability to diagnose complex operational issues across infrastructure, deployment pipelines and production services.
• Hands-on experience using AI coding agents as part of your engineering workflow.
- High level of autonomy — able to take an ambiguous operational problem and turn it into a reliable, documented solution.
- Strong written English and ability to collaborate effectively with distributed international teams.
⭐ NICE TO HAVE
• Experience with ML infrastructure , including SageMaker endpoints, GPU workloads, inference autoscaling or ML cost optimization.
• Experience operating streaming, real-time media or high-throughput infrastructure at scale .
• Experience with security and compliance , including secrets management, least-privilege IAM and audit requirements.
• Experience optimizing AWS infrastructure costs .
• Experience building auto-remediation, runbooks as code or automated incident-response systems .
? AUTOMATION-FIRST ENGINEERING
Automation is not an afterthought in this environment.
The platform already uses automated deployments and AI-assisted operational workflows. The goal is to continue moving toward infrastructure that can identify regressions, diagnose failures and trigger remediation with increasingly less human intervention .
You won’t simply be expected to respond to incidents.
You’ll be expected to ask:
“How do we prevent this from requiring a human next time?”
That mindset is central to this role.
? WHY THIS ROLE
You’ll have real ownership of the infrastructure supporting AI, computer vision and real-time video workloads used at significant production scale .
This role gives you the opportunity to work across both sides of modern DevOps:
Production Operations — reliability, observability, incident response and infrastructure ownership.
Platform Engineering — CI/CD, Infrastructure as Code, deployment automation and self-healing systems.
Success is measurable: fewer production incidents, safer deployments, faster recovery and more operational work handled automatically rather than manually.
? ABOUT BONZZU
Bonzzu connects talented software professionals with organizations worldwide. With a network of more than 40,000 professionals , we help build high-performance distributed engineering teams across a wide range of technologies.
📌 DevOps Engineer (Argentina)
🏢 Bonzzu
📍 Argentina