Explore

Production reliability

Site Reliability Engineer

Engineering discipline applied to the question of whether it stays up.

Typical pay

$120k-$210k

How you get in

Software engineering plus systems and cloud depth

Outlook

Strong; every company running services at scale needs it

SREs keep production systems reliable through automation, monitoring and incident response. It is software engineering pointed at operations, and the on-call pager is a defining feature of the role.

A day in the life

  1. 9:00aReview overnight alerts and error budget burn
  2. 10:00aWrite automation to eliminate a recurring manual fix
  3. 12:30pLunch
  4. 1:30pCapacity planning and load testing
  5. 3:00pPostmortem review for last week's incident
  6. 4:30pImprove monitoring so the next one is caught earlier

How to get in

Software engineer to SRE

Duration

2-4 years

Cost

$0-$3,000

Credential

None — production experience

  1. Work as a backend or infrastructure engineer
  2. Learn distributed systems, observability and incident response
  3. Take on-call and own the reliability of your own services
  4. Move into a dedicated SRE role, internally or externally

Sysadmin or DevOps into SRE

Duration

2-5 years

Cost

$1,000-$5,000

Credential

Cloud certifications plus programming skill

  1. Start in systems administration or DevOps
  2. Close the gap on real software engineering — this is the usual blocker
  3. Learn Go or Python well enough to build tooling, not just scripts
  4. Move to SRE, which pays notably more than traditional operations

CS degree into SRE

Duration

4-6 years

Cost

$40,000-$200,000

Credential

BS in computer science

  1. Complete a CS degree with strong systems coursework
  2. Intern on an infrastructure or platform team
  3. Start as a software engineer, then specialize toward reliability
  4. Understand you rarely start directly in SRE out of school

What people love

  • · Among the highest-paid engineering roles
  • · Deep, systems-level technical problems
  • · Remote-friendly
  • · Blameless postmortem culture in mature organizations

What wears people down

  • · On-call rotations wake you up
  • · Incident stress is real and cumulative
  • · Not entry level — requires strong engineering first
  • · Some companies use the title for pure operations work

From people doing it

Add yours