J

Systems Reliability Engineer

Jobgether • United State
Remote
Apply
AI Summary

Design, automate, and operate highly reliable distributed systems in a modern engineering environment. Maintain and improve reliability, performance, and scalability of large-scale production systems. Requires strong software engineering, infrastructure expertise, and automation skills.

Key Highlights
Design, build, and maintain reliable distributed systems
Develop automation tools using Python, Go, or Java
Implement observability solutions with Prometheus, Grafana, OpenTelemetry
Lead incident response and post-mortem reviews
Key Responsibilities
Maintain and improve the reliability, performance, and scalability of large-scale production systems
Design, build, and maintain reliable distributed systems while improving availability and performance
Develop automation tools and solutions using programming languages such as Python, Go, or Java
Operate and troubleshoot Linux-based systems at scale, including networking, performance optimization, and system-level issues
Manage Kubernetes and containerized workloads in production environments
Build, maintain, and improve CI/CD pipelines supporting both applications and infrastructure
Implement and enhance observability solutions using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or similar platforms
Monitor system health, define reliability improvements, and contribute to proactive performance management
Lead incident response activities, conduct post-incident reviews, and implement preventative improvements
Support the development and adoption of reliability practices including automation, service ownership, and operational excellence
Collaborate with engineering and cross-functional teams to improve system design and resilience
Technical Skills Required
Python Go Java Linux Kubernetes Prometheus
Benefits & Perks
Competitive annual salary range of $100,000-$150,000
Fully remote work environment within the United States
Full-time employment opportunity
Nice to Have
SLOs, error budgets, chaos engineering practices, and reliability frameworks
Familiarity with cloud platforms such as AWS, Azure, or GCP
Background in capacity planning, performance engineering, load testing, or service mesh technologies

Job Description


This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Systems Reliability Engineer based in United States.

This role offers the opportunity to design, automate, and operate highly reliable distributed systems in a modern engineering environment.

You will work at the intersection of software development and infrastructure operations to improve platform stability, scalability, and performance.

The position focuses on reducing operational complexity through automation, observability, and engineering best practices.

You will help ensure critical systems remain available and resilient while proactively addressing reliability challenges.

The ideal candidate will bring strong systems expertise, programming skills, and a passion for building dependable technology platforms.

You will collaborate with engineering teams to solve complex production challenges and continuously improve operational excellence.

This is a remote opportunity for a technically driven professional who values ownership, innovation, and measurable impact.

Accountabilities

The Systems Reliability Engineer will be responsible for maintaining and improving the reliability, performance, and scalability of large-scale production systems. The role requires a balance of software engineering, infrastructure expertise, automation, and operational leadership.

  • Design, build, and maintain reliable distributed systems while improving availability and performance.
  • Apply software engineering principles to infrastructure and operations challenges to reduce manual effort and operational complexity.
  • Develop automation tools and solutions using programming languages such as Python, Go, or Java.
  • Operate and troubleshoot Linux-based systems at scale, including networking, performance optimization, and system-level issues.
  • Manage Kubernetes and containerized workloads in production environments.
  • Build, maintain, and improve CI/CD pipelines supporting both applications and infrastructure.
  • Implement and enhance observability solutions using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or similar platforms.
  • Monitor system health, define reliability improvements, and contribute to proactive performance management.
  • Lead incident response activities, conduct post-incident reviews, and implement preventative improvements.
  • Support the development and adoption of reliability practices including automation, service ownership, and operational excellence.
  • Collaborate with engineering and cross-functional teams to improve system design and resilience.

Requirements

The ideal candidate combines strong software engineering capabilities with deep infrastructure knowledge and experience operating complex production environments.

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • 5+ years of experience in Site Reliability Engineering, DevOps, or production engineering roles supporting large-scale distributed systems.
  • Strong programming experience in at least one of the following: Python, Go, or Java.
  • Deep hands-on experience managing Linux environments, including networking, troubleshooting, and performance tuning.
  • Production experience with Kubernetes and container-based architectures.
  • Strong knowledge of observability and monitoring platforms such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or equivalent solutions.
  • Experience designing and managing CI/CD pipelines for applications and infrastructure.
  • Understanding of distributed system concepts including consistency models, partitioning, and failure handling.
  • Proven experience leading incident response and conducting effective post-mortem reviews.
  • Excellent communication, collaboration, and technical documentation skills.
  • Preferred experience with SLOs, error budgets, chaos engineering practices, and reliability frameworks.
  • Preferred familiarity with cloud platforms such as AWS, Azure, or GCP.
  • Preferred background in capacity planning, performance engineering, load testing, or service mesh technologies.

Benefits

  • Competitive annual salary range of $100,000-$150,000.
  • Fully remote work environment within the United States.
  • Full-time employment opportunity with career growth potential.
  • Opportunity to work on large-scale distributed systems and advanced technology projects.
  • Collaborative environment focused on engineering excellence and continuous improvement.
  • Exposure to modern cloud, automation, observability, and reliability practices.
  • Inclusive workplace committed to equal employment opportunities.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

Bright Vision Technologies

United State

Chief Technology Officer

Devops
•
28m ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Executive

leorna

United State

Lead AWS Data Engineer

Devops
•
29m ago
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Iris Software Inc.

United State

Subscribe our newsletter

New Things Will Always Update Regularly