Senior Manager & Principal Research Scientist
Working on Nemotron pre-training, focusing on architecture research, ablation recipes, scaling and distributed training.
I am a Senior Manager & Principal Research Scientist at NVIDIA, working on Nemotron pre-training with a focus on architecture research, ablation recipes, scaling and distributed training.
Previously, I led pre-training at poolside inside the Applied Research organisation, where I was responsible for the pre-training of Laguna M.1, XS.2, and S 2.1 on thousands of GPUs.
Before that, as a Senior Researcher in Natural Language Processing (NLP) and Automatic Speech Recognition (ASR) at AssemblyAI, I was responsible for the pre-training of Universal-1 and contributed to multimodal LLMs. Earlier, as an AI Research Engineer at InstaDeep's and BioNTech's joint AI innovation lab, I led a team of four engineers developing AI models for the treatment of cancer and the prevention and therapy of infectious diseases, including SARS-CoV-2.
I have a Master of Science in Machine Learning from University College London, where I was part of the Deciding, Acting, and Reasoning with Knowledge (DARK) lab, being supervised by Mikayel Samvelyan, Patrick Lewis, and Tim Rocktäschel. Before joining UCL, I completed my Bachelor of Science in Natural Language Processing at the University of Stuttgart under the supervision of Roman Klinger.
Working on Nemotron pre-training, focusing on architecture research, ablation recipes, scaling and distributed training.
Pre-trained large language models for code generation. Joined as a Member of Engineering, promoted to Tech Lead (Apr 2025) and then Team & Tech Lead (Dec 2025).
Set the technical direction for architecture research and distributed training behind the Laguna models, scaling Mixture-of-Experts training runs to 8,192 H200 GPUs and hundreds of billions of parameters. Led a team of 9 researchers and engineers.
Training multimodal (text and speech) large language models in JAX. Joined as Researcher II, promoted to Senior Researcher (Feb 2024).
Led self-supervised pre-training research for Universal-1, a SOTA multilingual speech-to-text model, on 12.5M hours of audio (~4.5T tokens) for models ranging from 660M to 2B parameters, trained on 512 TPU v5e. Implemented and optimised paired JAX and PyTorch codebases for model training and deployment.
Led a team of four research engineers, developing, evaluating and deploying state-of-the-art AI models (Transformers and Graph Neural Networks) based on the latest AI research applied to biology.
Developed a state-of-the-art relation extraction model based on distributional similarity with TensorFlow and transformers. Generated an automatically annotated text corpus with unique named entity identifiers based on the English Wikipedia encyclopedia.
Researching and reporting about current IT topics including artificial intelligence for one of the largest online computer magazines in Germany (>30 million PIs per month).
In arXiv preprint · 2026
In NAACL 2025 · 2025
In arXiv preprint · 2024
In NeurIPS MLSB 2023 · 2023
In arXiv preprint · 2023
In NeurIPS MLSB 2022 · 2022
In NAACL 2019 · 2019
Feel free to reach out about pre-training, distributed training, or research collaborations.