Back to Jobs
Site Reliability Engineer (SRE II)
Berkeley, CAOn-SiteTemporaryTechnology$80+/hrPosted 3 days ago
Job Description
Site Reliability Engineer (SRE II) | 100% Onsite | Berkeley, CA | $80/hr | 1-Year Contract
Important Notes:
Summary:
Perfect Timing Personnel is seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley National Laboratory's National Energy Research Scientific Computing Center (NERSC). As part of a 24x7 operations team, you'll help ensure the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery.
This role is ideal for a Linux-focused infrastructure professional with experience in systems operations, monitoring, automation, and incident response who enjoys solving complex technical challenges in a mission-driven environment.
What You'll Do:
Qualifications:
Required
Preferred
Important Notes:
- Permanent overnight schedule of midnight to 8:00 a.m., five days per week.
- 100% onsite in Berkeley, California
- No third-party agencies, Corp-to-Corp (C2C), or subcontracting arrangements
- Candidates must be authorized to work in the United States without current or future sponsorship
Summary:
Perfect Timing Personnel is seeking a Site Reliability Engineer (SRE II) to support Lawrence Berkeley National Laboratory's National Energy Research Scientific Computing Center (NERSC). As part of a 24x7 operations team, you'll help ensure the reliability and performance of critical high-performance computing infrastructure that supports scientific research and discovery.
This role is ideal for a Linux-focused infrastructure professional with experience in systems operations, monitoring, automation, and incident response who enjoys solving complex technical challenges in a mission-driven environment.
What You'll Do:
- Monitor computing, storage, network, and facility systems to ensure reliable operations
- Respond to alerts, troubleshoot issues, and coordinate with on-call teams as needed
- Develop and maintain automation, monitoring, and alerting solutions that improve operational efficiency
- Build tools and integrations that support incident management and infrastructure monitoring
- Support incident tracking and operational workflows through ServiceNow
- Collaborate across technical teams to coordinate maintenance activities and improve reliability
- Perform periodic data center walkthroughs to verify environmental, cooling, and power systems are operating properly
Qualifications:
Required
- Experience supporting large-scale IT infrastructure, data centers, or other highly available production environments
- Strong Linux administration and command-line experience
- Experience developing tools or automation using Python, Perl, Java, C, C++, or similar scripting/programming languages
- Experience troubleshooting production issues and working across technical teams to resolve them
- Strong written and verbal communication skills
- Ability and willingness to work a regular onsite overnight schedule
Preferred
- Experience in Site Reliability Engineering (SRE), DevOps, NOC, Systems Administration, Infrastructure Operations, or Platform Engineering
- ServiceNow experience
- Experience with Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, or similar monitoring platforms
- Familiarity with IT Service Management (ITSM) best practices
- Experience supporting HPC, research computing, scientific computing, or other mission-critical environments
- Experience developing automation tools; experience with AI-driven automation is a plus
Technology