ML Ops Engineer
175k - 275k USD
On-site
Full Time
#Engineering
#AI
#Kubernetes
#Terraform
#Cloud Services
#Distributed Training
#VPC
#Data
#Learning
#DevOps
#Parallel Computing
Radical AI, Inc. is an artificial intelligence company that is accelerating scientific research and development. We are at the forefront of innovation in the field of materials R&D, a critical driver for advancing our most cutting-edge industries and shaping the future. Breaking away from the traditionally slow and costly R&D process, Radical AI leverages artificial intelligence and machine learning to pioneer generative materials science. This innovative field blends AI, engineering, and materials science, revolutionizing how materials are created and discovered. Radical AI's approach speeds up R&D and addresses global challenges, setting new benchmarks in technology and sustainability.
What is this role?
We are hiring a Senior ML Ops Engineer to join our team on a full-time, on-site basis in the United States. In this position you will report to the Vice President of Research and help build our machine learning and data platform from the ground up, supporting the development, training, and deployment of models for materials research applications.
What will you do?
- Deploy and manage advanced machine learning models, with a focus on generative models for materials discovery, using Kubernetes, Terraform, and cloud services such as Lambda to scale models efficiently and adapt them to high-demand scenarios.
- Optimize computing infrastructure by improving GPU utilization, distributed training, bandwidth efficiency between machines, and VPC connections to maximize system performance.
- Work closely with the AI research team and cross-functional engineering groups to ensure effective model deployment and integration into production systems while maintaining documentation, running rigorous testing, and promoting engineering best practices.
What makes you a great fit?
You bring solid experience with DevOps, cloud infrastructure, and deploying machine learning models, along with expertise in network optimization and parallel computing. You are comfortable working with Kubernetes, Terraform, and cloud computing platforms to deliver scalable AI model deployments, and you have experience scaling model training across GPU clusters. You also have basic machine learning knowledge, including training generative models at scale, and you have built data pipelines and managed data infrastructure. Strong written and verbal communication skills help you convey complex technical information clearly, and you thrive in a collaborative team environment. We conduct our work in English.
What's in it for you?
We offer a salary range of $175000 – $275000, plus equity. Base pay may vary depending on job-related knowledge, skills, and experience. In addition to competitive compensation, we provide a comprehensive benefits package that includes medical, dental, and vision insurance for you and your family, a mental health and wellness budget, unlimited vacation along with 14 or more company holidays per year, and a 401k plan. You will work closely with a team at the cutting edge of AI research and contribute to a mission that aims to fundamentally change the way humanity makes progress through materials science discovery.









