13 ago
|
axiombio
|
Santa Fe
Agent Harness Engineer at axiombio.
About the role Axiom is constructing an agentic infrastructure layer designed to supplant animal testing and fundamentally reshape pharmaceutical drug discovery workflows. In this position, you will architect the foundational scaffolding that empowers frontier models to execute complex, long-horizon scientific reasoning across proprietary human biological datasets. The work sits at the intersection of advanced AI systems engineering and computational biology, requiring the translation of wet-lab scientific rigor into reliable, automated software processes. You will be responsible for the reliability and observability of the autonomous loops that drive scientific insight generation. This is a high-agency role where the engineering playbook is being written in real-time alongside the core technology.
Key facts
Location: SF General HQ Engagement: FullTime
What you'll do
- Architect and implement the core infrastructure and tooling necessary to evolve frontier large language models into fully autonomous scientific agents capable of independent hypothesis generation and validation.
- Construct sandboxed, deterministic execution environments that guarantee reproducibility for rigorous evaluation suites and reinforcement learning training runs.
- Develop and maintain high-throughput data pipelines for ingesting, storing, and versioning agent trajectories, runtime context windows, and synthetic training datasets.
- Design comprehensive evaluation frameworks encompassing offline static test suites, automated regression gates, and sophisticated LLM-as-judge pipelines for subjective scientific quality assessment.
- Implement deep observability and introspection tooling to trace every model invocation, tool interaction, and state transition for debugging and auditability.
- Engineer robust memory management, retrieval-augmented generation (RAG) systems, and checkpoint recovery mechanisms to maintain agent coherence over extended, multi-step analytical tasks.
- Partner directly with computational biologists and domain experts to codify scientific standards, experimental protocols, and regulatory requirements into quantifiable rubrics and golden evaluation sets.
- Orchestrate the central agent control loop, implementing safety guardrails, output verification protocols,
and dynamic resource budgeting (token/compute limits).
- Build reward modeling instrumentation and scalable rollout infrastructure to accelerate machine learning research iteration cycles.
- Optimize latency and cost efficiency of agent inference pipelines through caching strategies, model routing, and batch processing architectures.
- Establish CI/CD workflows specific to agent behavior, including automated red-teaming, capability benchmarking, and drift detection against production baselines.
- Document system architectures, failure modes, and operational runbooks to scale knowledge across a growing engineering team.
Requirements
- Demonstrated track record of designing, deploying, and maintaining production-grade agentic systems leveraging LLM APIs, including complex tool use schemas, multi-step planning, and recursive control loops.
- Deep software engineering expertise with a specialization in infrastructure engineering, platform development, data-intensive systems, or developer tooling.
- Prior experience authoring custom evaluation harnesses, real-time monitoring dashboards, or reinforcement learning environment observability stacks from the ground up.
- High agency mindset with a proven ability to drive ambiguous projects from vague problem statements to shipped solutions without requiring detailed specifications.
- Strong architectural philosophy favoring simplicity, maintainability, and operational robustness over premature optimization or unnecessary abstraction layers.
- Proficiency in post-mortem analysis of complex agent failure traces, including diagnosing tool misselection, context window overflow, reasoning hallucinations, and state corruption.
- Fluency in Python and comfort operating within a modern data stack involving columnar databases and containerized orchestration.
- Experience with infrastructure-as-code and cloud-native deployment patterns for GPU-intensive or high-I/O workloads.
Nice to have
- Direct experience applying AI/ML tooling to life sciences, bioinformatics, cheminformatics, or pharmaceutical R&D; workflows.
- Background in building simulation environments, digital twins, or scientific computing platforms requiring high numerical fidelity.
- Familiarity with reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), or online RL training loops for LLM agents.
- Contributions to open-source agent frameworks, evaluation libraries (e.g., LangSmith, Weights & Biases integrations), or LLM observability tools.
- Experience with frontend frameworks (SvelteKit, React) for building internal developer tooling, annotation UIs, or agent playgrounds.
- Knowledge of regulatory landscapes (FDA, EMA) regarding computational toxicology or non-clinical safety assessment validation.
Skills & tools
- Python (Advanced)
- Modal (Serverless compute)
- DuckDB (Analytical database)
- FastAPI (High-performance APIs)
- Docker (Containerization)
- Terraform (Infrastructure as Code)
- SvelteKit / Svelte 5 (Frontend framework)
- React (Frontend library)Git / GitHub Actions (Version control & CI/CD)
- Linux / Bash (Systems engineering)
Practical notes
- You will operate in a rapidly evolving technical domain where established best practices do not yet exist; you are expected to define the playbook through experimentation and rigorous engineering.
- The team prioritizes engineers who possess deep intellectual curiosity and a hands-on tinkering mindset, comfortable reading research papers and prototyping novel architectures on weekends.
- Success is defined by reliability engineering: preventing silent environment failures, data corruption, or evaluation drift that would poison training runs or invalidate scientific conclusions.
- The role is based at the San Francisco Global Headquarters with an expectation of high in-person collaboration cadence.
- Compensation package includes competitive base salary, significant equity upside in a seed‑stage venture‑backed company, and comprehensive benefits.
- Visa sponsorship (H-1B, O-1, TN, E-3) is available for exceptional candidates; relocation assistance provided for non‑local hires.
- META Company: axiombio Title: Agent Harness Engineer Listed location: SF Integral HQ Job type: full_time
📌 Agent Harness Engineer (Santa Fe)
🏢 axiombio
📍 Santa Fe