Site Reliability Engineer
Remote
Full Time
#Engineering
#Rust
#Containerization
#Helm
#Docker
#Kubernetes
#AWS
#Azure
#Google Cloud
#Network Protocols
#Infrastructure as Code
We are a dedicated engineering team focused on building robust, high-performance systems that serve our users reliably. We believe that great software is defined by its stability and its ability to scale under pressure. We are currently searching for a talented professional to help us maintain the health of our infrastructure and ensure our services remain available and efficient at all times.
About the Role
We are hiring a Senior Site Reliability Engineer for a full time position. In this role, you will be a key contributor to our engineering department, taking ownership of the design, implementation, and ongoing maintenance of our infrastructure. You will work closely with our developers to automate processes and improve the overall reliability of our platform.
Key Responsibilities
- Collaborate with cross-functional teams to ensure that all software is built and deployed with a focus on maximum system reliability.
- Design and implement automation tools that support the stability of our production services.
- Participate in incident response, troubleshoot complex issues, and create detailed root cause analysis reports for our customers.
Requirements
To be successful in this position, you should have a strong background in systems engineering and a passion for operational excellence. We are looking for the following qualifications:
- At least 8 years of professional experience working with production-grade systems.
- Demonstrated expertise in managing large-scale production environments.
- Strong proficiency in the Rust programming language.
- Hands-on experience with containerization tools, including Docker, Kubernetes, and Helm.
- Solid working knowledge of major cloud platforms such as AWS, Azure, or Google Cloud.
- A deep understanding of network protocols, DNS management, and load balancing.
- Experience using Infrastructure as Code to deploy and monitor complex systems.
- Excellent troubleshooting skills and a proactive approach to problem-solving.
Location
This position is fully remote and can be performed from anywhere.
Compensation and Benefits
We are committed to supporting our team members by providing a flexible work environment. This role includes the following benefit:
- Full remote work flexibility.
Grafbase.com
6 views
Markets



