About Meitner Energy
Meitner Energy is developing advanced nuclear energy solutions intended to deliver reliable, scalable, and carbon-free electricity for industrial and grid applications.
Our international team brings together nuclear, engineering, commercial, regulatory, and project-development experience. We are building an organization focused on disciplined engineering, responsible execution, and the deployment of nuclear energy at meaningful scale.
The Opportunity
Meitner’s AI platform runs on infrastructure that has to be right, and this role owns it. You will operate the bare-metal GPU cluster, the private networking and storage beneath it, and the compliance boundary that keeps regulated nuclear material separated from general-purpose workloads. You will also own the deployment automation and staging environments that other platform engineers rely on to ship reliably, which makes this the layer everything else is built on.
The work rewards someone who treats uptime and security posture as fixed requirements, can implement a default-deny network architecture and explain every exception in it, and has the discipline to automate and document rather than carry the environment in their head. You will work directly with the AI Platform Lead and partner with corporate IT on shared networking and identity where that makes sense.
This role is idóneo for someone who:
- Takes personal ownership of uptime and treats incidents as failures of process rather than bad luck.
- Builds infrastructure through code and automation instead of one-off manual changes.
- Is comfortable in regulated, security-conscious environments where the network boundary is a hard requirement.
- Wants to be the person the rest of the engineering team relies on when the cluster is the constraint.
What You'll Do
Cluster Operations and Security Boundaries
Operate and maintain the GPU cluster serving both the regulated and the non-regulated tiers, enforcing default-deny egress, network segmentation, and management-plane protection through zero-trust access patterns, bastion hosts, and privileged access controls. You will own the network architecture that separates export-controlled and otherwise sensitive workloads from general-purpose compute, and ensure those controls satisfy the export-control and controlled-information requirements applicable to advanced nuclear technology, working alongside the AI Platform Lead and compliance stakeholders. You will also partner with corporate IT on shared networking and identity infrastructure while maintaining clean boundaries between platform and corporate systems.
Infrastructure as Code and Deployment Automation
Own the infrastructure-as-code for the platform environment, using a declarative and version-controlled toolchain, so that every infrastructure change is version-controlled, reviewed, and reproducible from source. You will build and maintain staging and test environments that mirror production closely enough to catch failures before they reach the live cluster, and implement deployment automation that supports reliable,
auditable releases in coordination with the platform engineers who depend on it.
Reliability, Observability, and Capacity
Define and enforce backup procedures, restore testing cadence, and recovery objectives for the platform, ensuring recovery procedures are documented and verified rather than assumed. You will own the platform observability stack covering metrics, logs, and alerting, building dashboards and alerts that surface problems before they become incidents, and maintain the capacity and cost telemetry for GPU and storage resources that informs infrastructure investment decisions.
You will carry on-call responsibility for the platform infrastructure, respond to incidents with urgency, and conduct blameless post-incident reviews that close the loop on root cause rather than stopping at service restoration.
What We're Looking For
Required Qualifications
- Bachelor’s degree in computer science, engineering, or a related technical discipline, or equivalent professional experience.
- 4+ years of professional experience in site reliability engineering, infrastructure engineering, or systems engineering.
- Strong Linux systems administration, networking fundamentals, and firewall and access control list management, with the ability to diagnose a network problem and trace a storage failure independently.
- Production container orchestration experience and demonstrated declarative infrastructure-as-code proficiency.
- Experience operating GPU clusters or comparable high-performance compute infrastructure.
- Hands-on experience with production metrics, logging, and alerting stacks, including building alerts that reflect real failure modes.
- Demonstrated on-call ownership with incident response and post-incident review practice, including verified restore testing rather than backup reports alone.
- Experience implementing network segmentation, default-deny policies, and privileged access controls in a production environment, and the ability to work on-site in Dallas.
Preferred Qualifications
- Working proficiency in Spanish. Spanish is preferred because the position involves infrastructure collaboration with colleagues in Argentina.
- Experience implementing zero-trust network access, bastion, or privileged access management patterns in production.
- Familiarity with regulated or export-controlled data environments, including controlled unclassified information handling.
- Experience with storage systems for high-throughput machine learning workloads, including high-speed local storage, distributed file systems, and object storage.
- Background in nuclear energy, defense, aerospace, or another security-sensitive industrial sector.
The Candidate We Are Seeking The strongest candidate will be an engineer who has carried a pager for infrastructure they built themselves and can describe what they learned the hard way. This may be an excellent next step for a site reliability engineer, infrastructure engineer, or systems engineer who wants full ownership of a cluster and its security boundary rather than a share of a large fleet. Candidates should be prepared to discuss the worst outage they owned, not the cleanest one.
Why Join Meitner Energy?
Consequential Work. Build and operate the infrastructure foundation supporting the delivery of reliable, scalable, carbon-free energy.
Direct Ownership. Own the cluster, the network boundary, and the recovery plan outright rather than filing requests against another team. Meitner is small enough that your decisions visibly shape the company.
International Scope. Work with professionals across the United States, Argentina, and the United Kingdom.
Strong Benefits. Meitner offers comprehensive health insurance, a 401(k) retirement plan, and participation in the company's employee stock option program, subject to plan terms and eligibility. The final offer will reflect the candidate's depth and relevant skills, including infrastructure depth, security design experience, Spanish proficiency, and regulated-industry background.
High-Quality Workplace. Work from a modern Dallas office designed to support collaboration, productivity, and employee well-being, including an on-site fitness facility.
To Apply
Email your resume and a written response of no more than one page to
[email protected]. Use this exact subject line: Site Reliability Engineer / your full name / hands-on. Applications without it will not be reviewed.
Write the response yourself, in your own words, for this posting specifically. Answer directly with specific examples; generic statements of engineering philosophy or best practices are not considered responsive, and responses that appear mass-produced or machine-generated are declined without review. Your one page should address the following:
- Quote the single sentence from this posting that best describes how this role differs from your current position, and explain why that difference appeals to you.
- Describe the most complex AI or LLM platform system you have designed, built, or operated in the past three years. What was the stack, what were the hardest trade-offs, what broke, and how did you fix it?
- Walk us through how you would fit a large open-weight model onto a fixed GPU memory budget. What trade-offs would you consider, and how would you verify the result?
- Describe how you have approached data-sensitivity routing or access control in a regulated or security-sensitive environment. What did you build, and how did you verify it held?
- Explain why this role is the right next step in your career, and identify one aspect of the position that may be more demanding than your current role.
Please do not include confidential, proprietary, export-controlled, or otherwise restricted information belonging to a current or former employer.
📌 Site Reliability / Infrastructure Engineer (Buenos Aires)
🏢 Meitner Energy
📍 Buenos Aires