Founding Engineer, Core Runtime
Zibra Labs is building the post-training and inference infrastructure that currently only exists at the frontier labs. Our vision is to enable everyone to post-train and serve open weight models at a fraction of the cost with higher capability.
The core runtime team is building the scheduling, memory and process management, and high performance networking stack on specialized hardware.
What You Will Work On
- An intra-node scheduler that fair-shares memory, CPU, and GPU across multiple workloads.
- A high-performance lock-free shared-memory communication protocol for intra-process and intra-container communication.
- Low latency runtime container checkpointing and migration.
- Soft and hard failure detection in network and shared memory interfaces on specialized hardware.
About You
- You have a strong grasp of a systems programming language.
- You have taste in system design and can determine when to build, when to buy, and when to cut losses from an incorrect decision.
- You can reason about dynamic systems problems such as queue overflows, unreliable data transmission, and impedance mismatch between producers and consumers.
- You care about performance when it matters. You can tell when it doesn't matter.
- You're high agency: you can take a problem, figure out what to do, and ship it.
- You care about the quality of your work.
Jargon That Might Be Useful
None of this is a checklist. It is a map of the territory, so you can tell whether it is the territory you want to be in.
- Transport libraries. NCCL, UCX
- Orchestration. Ray, Kubernetes, Slurm
- Inference. SGLang, vLLM, TensorRT
- RL. veRL, SkyRL
- Training. PyTorch, JAX
- Execution. gVisor, Firecracker, Linux
Apply
Send a note about what you have built and what you want to work on, with anything that shows the work: a repo, a design doc, or the story of a system you took from broken to boring.
Our address: careers (at) zibralabs (dot) ai