
Member of Technical Staff, Data Pipeline
On-site
Full Time
#Engineering
#Machine Learning
#Data Processing
#Python
#PyTorch
#Data Labeling
#Database Management
#Cloud Platforms
#Data Privacy
#Data Collection
Boson AI is an early-stage startup dedicated to making large language tools accessible to everyone. Our team is led by founders Alex Smola and Mu Li, and we are supported by a group of scientists and engineers specializing in deep learning, optimization, natural language processing, AutoML, and statistics. We are currently focused on developing high-quality generative AI models that extend beyond language. We are looking for a senior machine learning engineer to join our team on a full-time basis at our office in Santa Clara. In this role, you will play a vital part in building the infrastructure for data collection, extraction, filtering, and synthetic generation, which are essential for creating more lifelike AI models.
Key outcomes
- Design and build robust data processing pipelines that handle extraction, filtering, and labeling.
- Implement machine learning models to enhance the quality and diversity of our data, including the development of quality classifiers, document layout models, and speech transcription tools.
- Collaborate closely with our research and engineering teams to advance our next generation of large multimodal models.
- Identify and resolve data anomalies and biases to ensure the highest standards of data quality.
Requirements
- A Master's degree or PhD in computer science or a related field.
- Proven experience in machine learning projects involving audio, text, or vision, with a track record of training models to solve specific problems.
- Strong expertise in constructing large-scale data processing pipelines and familiarity with distributed workloads using tools like Ray, Docker, Kubernetes, and multiprocessing.
- High proficiency in Python and the ability to produce clean, maintainable code.
- Hands-on experience with at least one deep learning framework, specifically PyTorch.
- Excellent problem-solving skills and meticulous attention to detail.
- Full professional proficiency in English.
Preferred qualifications
- A history of active contributions to open-source projects on GitHub.
- Direct experience in building large-scale datasets.
- Familiarity with data labeling tools such as LabelStudio, data collection techniques like Selenium, or processing frameworks such as Hadoop and Datasketch.
- Working knowledge of database management.
- Practical experience with cloud platforms including AWS, Azure, or GCP.
- Multilingual abilities that can help enrich the language diversity of our training data.
- Understanding of fairness, toxicity, and data privacy regulations.
How to apply
If you are passionate about building the future of generative AI and possess the technical background we are looking for, we invite you to apply. Please submit your application to be considered for this position at our Santa Clara office.







