21 sep
|
OpenRelay (YC S26)
|
Argentina
21 sep
OpenRelay (YC S26)
Argentina
Why OpenRelay
OpenRelay is GPU compute and LLM inference for people who ship. Customers rent GPU machines and call inference APIs; providers plug their hardware into our network and get paid for it. We run a distributed fleet on real hardware, not a reseller skin over someone else's cloud. Small team, high bar, and room for A-players to own things that are live in production.
No degree required, show us what you have run
We care about what you can do, not diplomas. The profile we want most is simple: you have built infrastructure and kept it alive. Not a tutorial cluster, real machines with real workloads, where a mistake takes something down and you are the one who brings it back. Show us the systems you have run, what broke, and what you did about it.
Core tech
- Terraform, AWS (ECS Fargate, DynamoDB, S3, VPC networking)
- Nomad on bare-metal GPU hosts (QEMU/KVM, VFIO GPU passthrough)
- Go (control plane, gateways, node agents)
- Linux deep enough to argue with: kernel parameters, systemd, networking, drivers
- Overlay networking (Nebula, WireGuard, Tailscale), mTLS PKI
- VictoriaMetrics + Grafana, GitHub Actions
- Claude Code and agentic tooling, daily
What you will own
- The Terraform that defines both AWS environments, and the discipline of applying it: plan, review, beta, then prod
- The provider fleet: enrollment, node identity (mTLS certificates), health reconciliation, and the lifecycle of VMs scheduled onto bare-metal GPU hosts through Nomad
- The ugly, real parts of GPU hosts: VFIO passthrough, IOMMU groups, kernel flags, why this exact box will not release its GPU
- The network paths: customer traffic through our gateways, the control overlay between nodes and the control plane, and the debugging when a path silently degrades
- Monitoring and the alerts a human actually acts on:
dashboards, runbooks, and the discipline of making the next incident boring
- Ops tooling for a fleet you cannot always SSH into: safe remote execution, self-updating agents, automation that fails loudly
Skills and qualifications
- 2+ years running production infrastructure, professional or a serious homelab/fleet you can show
- Terraform or equivalent infra-as-code on a real cloud account
- Advanced Linux: you have debugged a boot problem, a driver problem, and a network problem, and can tell the stories
- Comfortable reading and writing Go, or strong in one systems language and willing
- Docker and at least one orchestrator (Nomad, Kubernetes, or equivalent)
- Able to own a system end to end: design, rollout, monitoring, incident response
- Working English, written and spoken
Nice to have: GPU or bare-metal experience, QEMU/KVM or VFIO, PKI/mTLS, Nebula/WireGuard/Tailscale, DynamoDB, on-call experience at a small company.
The video
Required, five minutes or less, in English. Two parts: your face on camera introducing yourself, and a screen recording walking through the best infrastructure you have built or run and the hardest problem it gave you. Casual is fine, content matters more than polish.
What to expect
- Fast-moving U.S. startup building GPU compute and inference infrastructure
- High autonomy, high expectations, fair and human culture
- Direct influence on product decisions
- Fully remote, anywhere in Argentina
- Competitive pay, strong mentorship and growth
How to apply
Apply on our site, not here: https://openrelay.inc/careers/infrastructure-engineer-argentina . The form takes a few minutes and asks for a short demo video, five minutes or less, in English. The video is required. There is an example on the page showing exactly what we are looking for.
📌 Infrastructure Engineer (Argentina)
🏢 OpenRelay (YC S26)
📍 Argentina