Linux Systems Administrator responsible for maintaining the operational health of core compute services, managing tickets, and driving operational work. Requires 4-8 years of experience with Linux, Python, and CI/CD processes.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Job Description
Linux Systems Administrator (Linux, Python, CI/CD)
100% Remote
Fulltime (Permanent)
Job Description:
Experience:
- 4–8 years of relevant professional experience.
- Strong experience with Linux operating systems and Linux system administration.
- Experience using Linux and Unix commands.
- Experience automating operational tasks using Python, Bash, and JavaScript.
- Experience supporting mission-critical Tier 1 services in an operational environment with pager-duty responsibilities.
- Software development experience focused on services and operational tools.
- Experience with CI/CD processes and software practices.
Interested in remote work opportunities in IT & Network Engineering? Discover IT & Network Engineering Remote Jobs featuring exclusive positions from top companies that offer flexible work arrangements.
Core Skills:
- Operating Systems: Linux, Unix
- Programming and Scripting Languages: Python, Bash, JavaScript
- System Administration: Linux System Administration, Linux/Unix Commands
- Monitoring and Observability: Service Metrics, Dashboards, Service KPIs, Alarming Systems
- Development and Delivery: CI/CD, Services Development, Operational Tools
- Service Operations: Mission-Critical Tier 1 Services, Pager Duty, Incident Response, Root-Cause Analysis
- Operational Excellence: Runbooks, Operational Toil Mitigation and Reduction, Automation
Browse our curated collection of remote jobs across all categories and industries, featuring positions from top companies worldwide.
Key Responsibilities:
- Maintain the operational health of core compute services to support API availability and low latency.
- Manage and triage tickets based on business and technical impact.
- Drive the prioritization and execution of operational work.
- Scale systems sustainably through easy-to-use tooling and automation.
- Collaborate with service developers to improve scalability, reliability, and development velocity.
- Develop runbooks to reduce the mean triage time for incidents.
- Prioritize and automate frequently used runbooks.
- Practice sustainable incident response and drive root-cause analysis.
Similar Jobs
Explore other opportunities that match your interests