SoftwareCareers
Loading page...
SoftwareCareers
Loading page...
at GitLab • Remote
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and over 50% of the Fortune 100 trust GitLab to ship better, more secure software faster. We embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. We hire Site Reliability Engineers from Intermediate through Senior Staff, evaluating skills holistically and matching candidates to opportunities that best align with their experience and our hiring needs. We seek engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly, and will support your success with GitLab's tools, systems, and ways of working. Our SRE hiring process includes a Recruiter Screen, Core Technical assessment, Peer Technical interview, Hiring Manager Interview, and a Skip-Level Interview. Your level is calibrated during the process; Intermediate roles involve meaningful contributions within a scoped area, diagnosing issues, prioritizing, and documenting work. Senior SREs drive reliability improvements, lead investigations, coordinate incident response, and influence team patterns. Staff SREs shape reliability strategy across teams, introduce prevention strategies, design execution at organizational scale, and connect reliability work to business needs. Senior Staff SREs set technical direction for reliability across a sub-department, tackle ambiguous systems problems, establish standards, and mentor other engineers. What you'll do: Keep user-facing services and production systems reliable, scalable, and efficient. Build automation and tooling that reduces toil, replacing manual work with repeatable, infrastructure-as-code-driven workflows. Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling. Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps. Participate in on-call, triage alerts, improve runbooks, and escalate appropriately. Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early. Take part in incident response and post-incident reviews, turning learnings into automation and process changes. Document runbooks, architecture decisions, and reviews. What you'll bring: Experience keeping production systems reliable, combining an operations mindset with software engineering practice. Experience building net-new infrastructure tooling and automation (e.g., Terraform modules, Kubernetes operators/controllers). Ability to read, debug, and reason about code (Go; some Ruby). Experience with infrastructure as code, Kubernetes, and its ecosystem. Hands-on experience with at least one major cloud provider (GCP or AWS). Familiarity with observability practices, including metrics, logging, alerting, and SLOs/SLIs. Comfort participating in on-call and incident response. Strong written communication and ability to operate as a manager-of-one in an async, distributed environment. A track record of using automation, and increasingly AI, to reduce toil. Alignment with GitLab's values. About the team: Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab's user-facing services, most notably GitLab.com. The department spans Production Engineering and Dedicated, owning everything from the production fleet and networking platform to observability, incident response, and our single-tenant Dedicated offering. We are a globally distributed, all-remote group that works asynchronously, favors automation over toil, and closes the loop with monitoring and metrics to drive accountability.
Application planning
More visa sponsored jobs in United States
No related jobs found.
Receive verified Site Reliability Engineering, Infrastructure Platforms jobs with visa and relocation signals.