Explore
Production reliability
Site Reliability Engineer
Engineering discipline applied to the question of whether it stays up.
Typical pay
$120k-$210k
How you get in
Software engineering plus systems and cloud depth
Outlook
Strong; every company running services at scale needs it
SREs keep production systems reliable through automation, monitoring and incident response. It is software engineering pointed at operations, and the on-call pager is a defining feature of the role.
A day in the life
- 9:00aReview overnight alerts and error budget burn
- 10:00aWrite automation to eliminate a recurring manual fix
- 12:30pLunch
- 1:30pCapacity planning and load testing
- 3:00pPostmortem review for last week's incident
- 4:30pImprove monitoring so the next one is caught earlier
How to get in
Software engineer to SRE
Duration
2-4 years
Cost
$0-$3,000
Credential
None — production experience
- Work as a backend or infrastructure engineer
- Learn distributed systems, observability and incident response
- Take on-call and own the reliability of your own services
- Move into a dedicated SRE role, internally or externally
Sysadmin or DevOps into SRE
Duration
2-5 years
Cost
$1,000-$5,000
Credential
Cloud certifications plus programming skill
- Start in systems administration or DevOps
- Close the gap on real software engineering — this is the usual blocker
- Learn Go or Python well enough to build tooling, not just scripts
- Move to SRE, which pays notably more than traditional operations
CS degree into SRE
Duration
4-6 years
Cost
$40,000-$200,000
Credential
BS in computer science
- Complete a CS degree with strong systems coursework
- Intern on an infrastructure or platform team
- Start as a software engineer, then specialize toward reliability
- Understand you rarely start directly in SRE out of school
What people love
- · Among the highest-paid engineering roles
- · Deep, systems-level technical problems
- · Remote-friendly
- · Blameless postmortem culture in mature organizations
What wears people down
- · On-call rotations wake you up
- · Incident stress is real and cumulative
- · Not entry level — requires strong engineering first
- · Some companies use the title for pure operations work