Remote Sr. Site Reliability Engineer Job at Reddit, Remote

MHZSOVdxWmtwckJGbXphOXdQY1NMaVVqcUE9PQ==
  • Reddit
  • Remote

Job Description

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities

  • Advise :
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify :
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit.
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate :
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose :
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities.
  • Optimize :
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits

  • Pension Scheme
  • Private Medical and Dental Scheme
  • Life Assurance, Income Protection
  • Workspace benefit for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Jobicy job ID: 116922

Job Tags

Remote job, Full time, Home office, Flexible hours,

Similar Jobs

Boyd's Electrical Service, Inc.

Summer Internship / Work Study Apprentice at Clearwater, NE Job at Boyd's Electrical Service, Inc.

 ...offering an opportunity for students of a technical college, or those interested in...  ...would typically be during the summer months when on summer...  ...time. This is a part time paid position. The applicant must...  ....Job Descripion - Internship/Work Study Apprentice Electrician... 

C&R Management Group LLC

Affordable Property Manager (Bilingual) Job at C&R Management Group LLC

 ...Description Description: Commercial and Residential Management Group (CRMG) is looking for an Affordable Property Manager with amazing attention to detail and...  ...Options: Option 1: $26$28 per hour + Live on-site with a 100% discount on a 3-bedroom apartment... 

D.C. United

Director, Corporate Communications (Washington) Job at D.C. United

 ...new challenges and new opportunities? Join our team! D.C. United is seeking an experienced and dynamic Director of Corporate Communications to lead and execute strategic communication initiatives that enhance the club's reputation, drive brand awareness, and foster... 

Purple Unicorn

Customer Service, Inside Sales Representative Job at Purple Unicorn

 ...Sales Representative Location: Fully Remote Job Overview: Purple Unicorn...  ...and help others protect what matters most. Experience is great but not required. If you bring...  ...% employer match ~ Remote & Flexible: Work from home on your schedule ~ Professional... 

Farm Job Search

Farm Equipment Operator Job at Farm Job Search

 ...Farm Equipment Operator (6151) Location: New York State JobNumber: 6151 Farm Equipment Operator Position available on 1,200 cow dairy and 8,000 acre crop farm in North Central New York. MUST be able to operate and repair large modern John Deere equipment and...