Zorky CRMZorky CRM
EN|RU
@ekaterinovikova
All jobs

Member of Technical Staff - Foundations

principalremoteSan Francisco / Tel Aviv / Zurich, USScore undefined/1005d ago
Stack
fine-tuningpythonpytorchtriton
Apply
Upload your CV — we will connect you with the employer directly through our pool.
Send your CV →
Description
Tzafon is a foundation model lab building scalable compute systems and advancing machine intelligence, with offices in San Francisco, Zurich & Tel Aviv. We’ve raised over $12m in funding to advance our mission of expanding the frontiers of machine intelligence. We're a team of engineers and scientists with deep backgrounds in ML infrastructure & research. Founded by IOI and IMO medalists, PhDs, and alumni from leading tech companies, such as Google Deepmind, Character, and NVIDIA, we train models and build infrastructure for swarms of agents to automate work across real-world environments. You'll work between our product and post-training teams to ship Large Action Models that actually work. Build evals, benchmarks, and fine-tuning pipelines. Define what good model behavior means and make it happen at scale. What you'll do Design and execute large scale training runs on our clusters Build and optimize distributed training infrastructure across massive multi-node systems Implement post-training pipelines at scale Develop data pipelines that process and filter trillions of tokens for pre-training Research and implement architectural improvements, scaling laws, and training optimizations Debug training instabilities, loss spikes, and convergence issues in long-running jobs Build tooling for cluster utilization, fault tolerance, and checkpoint management Write custom CUDA/Triton kernels to optimize critical training operations (attention, normalization, activations) Collaborate on research that advances the state of the art in foundation model training We're looking for Deep experience pre-training or post-training foundation models on large clusters Expert-level at Python and ML frameworks (PyTorch, JAX, Torchtitan) Strong systems skills: distributed training, FSDP/ZeRO, tensor parallelism, pipeline parallelism Experience writing performant CUDA or Triton kernels for ML workloads Track record of running stable multi-week training jobs and debugging distributed
Employer contacts (email/phone/telegram) are hidden from the public preview — send your CV, and we will connect you directly.
Urgent question? Message @ekaterinovikova