Senior Site Reliability Engineer - Cloud

NVIDIA

Jun 11, 2024

Applications are closed

Job
Full-time
Senior Level
Software Engineering
$168K - $264.5K
Santa Clara
Remote

Requirements

What we need to see:
BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience
5+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production
Experience in one or more of the following: Python, Go, Perl or Ruby
In depth knowledge on Linux, Networking and Containers
Ways to stand out from the crowd:
Interest in crafting, analyzing and fixing large-scale distributed systems
Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive
Ability to debug and optimize code and automate routine tasks
Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack and Docker

Responsibilities

Design, implement and support operational and reliability aspects of large scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting
Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinement
Support services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviews
Maintain services once they are live by measuring and monitoring availability, latency and overall system health
Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity
Practice sustainable incident response and blameless postmortems
Be part of an on call rotation to support production systems

FAQs

What are the key responsibilities of a Senior Site Reliability Engineer - Cloud at NVIDIA?

The key responsibilities include designing, building, and maintaining large scale production systems with high efficiency and availability using software and systems engineering practices, ensuring reliability and uptime of GPU cloud services, enabling developers to make system changes through careful planning and preparation, automating manual work, performance tuning, and optimizing production systems.

What skills and knowledge are required for this role?

Skills and knowledge required include expertise in systems, networking, coding, database, capacity management, continuous delivery and deployment, Kubernetes, OpenStack, and other cloud enabling technologies. Additionally, a mindset of problem solving, diversity, intellectual curiosity, and openness is important for success in this role.

What is the culture like at NVIDIA's Site Reliability Engineering organization?

The culture at NVIDIA's Site Reliability Engineering organization is one of diversity, collaboration, intellectual curiosity, problem solving, and openness. The organization encourages collaboration, thinking big, taking risks in a blame-free environment, self-direction on meaningful projects, and providing support and mentorship for learning and growth.

NVIDIA

Manufacturing & Electronics

Industry

10,001+

Employees

1993

Founded Year

Mission & Purpose

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

OpportunitiesView all

Verification Engineer

Job

Bangalore

Food and Beverage Manager

Job

Pune

Government Affairs Program Manager

Job

Berlin