Founding ML Researcher (Post Training)
Zibra Labs is building the post-training and inference infrastructure that currently only exists at the frontier labs. Our vision is to enable everyone to post-train and serve open weight models at a fraction of the cost with higher capability.
The research side of that vision owns the recipes: how an open weight model gets post-trained cheaply and served cheaply without giving up capability. You will sit next to the team building the runtime, close enough that a recipe needing a new primitive can get one.
What You Will Work On
- Post-training recipes for open weight models that hold up outside the benchmark they were tuned against.
- Reinforcement learning at the algorithm level: reward design, rollout efficiency, and keeping long-horizon runs stable.
- Making a served model cheaper without making it worse, through quantization, distillation, sparsity, and speculative decoding.
- Evaluation that survives contact with reality, catching the regressions a leaderboard number hides.
- Shaping the product direction: what we post-train, what we serve, and what ships next are research calls as much as engineering ones.
About You
- You have trained models at a scale where the infrastructure fought back, and you can tell which failures were the recipe and which were the machine.
- You read the literature closely enough to know which results reproduce.
- You have taste in experiment design: you can pick the smallest run that answers the question.
- You write code other people can run. Research is not an excuse.
- You're high agency: you can take a capability goal, turn it into experiments, and ship the result.
- You care about the quality of your work.
Jargon That Might Be Useful
None of this is a checklist. It is a map of the territory, so you can tell whether it is the territory you want to be in.
- Post-training. SFT, DPO, GRPO, PPO
- RL. veRL, SkyRL
- Training. PyTorch, JAX, FSDP, DeepSpeed
- Inference. SGLang, vLLM, TensorRT
- Efficiency. FP8 and INT4 quantization, distillation, speculative decoding, LoRA
- Evaluation. lm-evaluation-harness, task-specific harnesses
Apply
Send a note about what you have built and what you want to work on, with anything that shows the work: a paper, a repo, a training run you debugged to the bottom.
Our address: careers (at) zibralabs (dot) ai