11 sep
|
simplyblock
|
Argentina
11 sep
simplyblock
Argentina
Back
Why this role exists
Simplyblock builds Kubernetes-native, software-defined NVMe-over-Fabrics block storage. Our customers run it underneath databases and stateful workloads where downtime is measured in money and latency is measured in microseconds.
When something behaves oddly in a large multi-tier distributed storage cluster, the customer talks to you. You are the person who turns a vague symptom into a reproducible root cause, a fix, and a change that keeps it from recurring.
This is not a ticket-queue job. You sit between our customers and our engineering team, you own the analysis end to end, and you have a direct hand in how the product evolves.
What You’ll Do
- Analyse complex operational and software issues in a large multi-tier distributed system: data path, control plane, cluster state, and the Kubernetes layer around it
- Own customer communication. Respond to incidents, guide customers through upgrades and migrations, and answer advanced product, architecture, and tuning questions
- Go deep on root cause. Logs, metrics, traces, core dumps, packet captures, reproductions: whatever it takes to explain the behaviour rather than describe it
- Coordinate resolution across engineering. Bring the pieces together from storage, networking, control plane, and platform teams, and keep momentum until the issue is genuinely closed
- Push proactively for improvements. Permanent fixes over workarounds, better observability, better defaults, and automation that reduces both the number of incidents and the manual operations effort
- Turn what you learn into assets: runbooks, diagnostic tooling, product feedback, and documentation the next engineer doesn’t have to rediscover
What You Bring
- Networking engineering depth. TCP/IP internals, RDMA/RoCE, congestion and flow control, multipathing, MTU and offload issues:
you can read a capture and tell a story with it
- Distributed systems experience in production.
A real feel for failure modes: partitions, split brain, quorum loss, tail latency, silent degradation, clock and ordering problems
- Kubernetes operations. CSI, operators, StatefulSets, scheduling and node lifecycle, and the day-2 realities of running storage under an orchestrator
- AI-assisted engineering as part of your workflow. Claude or comparable models and agents for log and code analysis, reproduction scripting, and customer-facing drafting. We expect you to use these tools to make yourself materially faster, and to have opinions about where they help and where they don’t
- Excellent written and spoken communication with technical customers. You stay precise and calm when the other side is not
- A proactive, high-responsiveness mindset. You chase the loose end nobody assigned to you, and you close the loop without being asked
Preferred
- DPDK, SPDK, NVMe-oF, or other kernel-bypass and userspace storage and network stacks
- Low-level Linux debugging: perf, eBPF, gdb, crash dumps, kernel and driver behaviour
- Comfortable writing Python, Go, or C for tooling, diagnostics, and reproductions
- Prior customer-facing experience at an infrastructure or storage vendor
- Fluent English required; German or other European languages a plus
What We Offer
- Direct impact on a technically hard product, with a short path from your findings to shipped changes
- Work alongside the engineers who wrote the code. No escalation wall between you and them
- The AI tooling budget to match a Claude-native company: Claude Max, Claude Code, and whatever else you can justify
- Competitive salary + equity
- A team that celebrates ambition and outstanding results. No participation trophies.
If you’d rather explain the behaviour than describe it: apply.
📌 Senior Support Engineer (Distributed Storage) (Argentina)
🏢 simplyblock
📍 Argentina