Multimodal Generative AI Researcher
Remote
Full Time
#Research
#Artificial Intelligence
#Machine Learning
#Language
#Language Models
#Representation Learning
#PyTorch
#Distributed Training
#DeepSpeed
#Ray
#3D
We're building the next generation of models that can reason across vision, language, and 3D. Our team works at the intersection of cutting-edge research and scalable engineering, creating systems that move beyond text-only intelligence into truly multimodal understanding. If you're excited by the challenge of training and adapting large vision-language models for real-world tasks, we'd love to hear from you.
What you'll be doing
- Design and fine-tune large-scale vision-language and language models, along with hybrid architectures, to tackle tasks such as visual reasoning, retrieval, 3D understanding, and embodied interaction.
- Build and maintain robust training and evaluation pipelines, including data curation, distributed training setups, mixed-precision techniques, and scalable fine-tuning workflows.
- Analyse model performance through detailed experiments, ablations, bias and robustness checks, and generalisation studies while collaborating with research, engineering, and 3D teams to move models from prototype into production.
What you'll bring
You hold a PhD or possess equivalent experience in Machine Learning, Computer Vision, NLP, Robotics, or Computer Graphics. You have a proven track record of training or fine-tuning large-scale vision-language and language models for downstream applications. You combine deep technical knowledge with a strong engineering mindset, allowing you to design, debug, and scale complete training systems.
You understand multimodal alignment and representation learning, including vision-language fusion, CLIP-style pre-training, and retrieval-augmented generation. You're familiar with current trends such as video-language and long-context models, spatio-temporal grounding, agentic multimodal reasoning, and Mixture-of-Experts fine-tuning. You also have awareness of 3D-aware multimodal approaches that leverage NeRFs, Gaussian splatting, or differentiable renderers.
Hands-on experience with PyTorch, DeepSpeed, Ray, and distributed or mixed-precision training is essential. You communicate clearly and work effectively across teams. English proficiency is required for this remote role.
What you'll get
We offer the opportunity to work remotely from anywhere, giving you the flexibility to contribute from the location that suits you best. You'll join a collaborative environment where research and engineering work closely together to push the boundaries of multimodal AI.
Stability AI
13 views
Markets












