Logo of Huzzle

Senior Site Reliability Engineer - Cloud

image

NVIDIA

2mo ago

Applications are closed

  • Job
    Full-time
    Senior Level
  • Software Engineering
  • $168K - $264.5K
  • Santa Clara
    Remote

Requirements

  • What we need to see:
  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience
  • 5+ years of experience with Infrastructure automation, distributed systems design, experience with design, develop tools for running large scale private or public cloud system in Production
  • Experience in one or more of the following: Python, Go, Perl or Ruby
  • In depth knowledge on Linux, Networking and Containers
  • Ways to stand out from the crowd:
  • Interest in crafting, analyzing and fixing large-scale distributed systems
  • Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive
  • Ability to debug and optimize code and automate routine tasks
  • Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack and Docker

Responsibilities

  • Design, implement and support operational and reliability aspects of large scale Kubernetes clusters with focus on performance at scale, real time monitoring, logging and alerting
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation and refinement
  • Support services before they go live through activities such as system design consulting, developing software tools, platforms and frameworks, capacity management and launch reviews
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health
  • Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity
  • Practice sustainable incident response and blameless postmortems
  • Be part of an on call rotation to support production systems

FAQs

What are the key responsibilities of a Senior Site Reliability Engineer - Cloud at NVIDIA?

The key responsibilities include designing, building, and maintaining large scale production systems with high efficiency and availability using software and systems engineering practices, ensuring reliability and uptime of GPU cloud services, enabling developers to make system changes through careful planning and preparation, automating manual work, performance tuning, and optimizing production systems.

What skills and knowledge are required for this role?

Skills and knowledge required include expertise in systems, networking, coding, database, capacity management, continuous delivery and deployment, Kubernetes, OpenStack, and other cloud enabling technologies. Additionally, a mindset of problem solving, diversity, intellectual curiosity, and openness is important for success in this role.

What is the culture like at NVIDIA's Site Reliability Engineering organization?

The culture at NVIDIA's Site Reliability Engineering organization is one of diversity, collaboration, intellectual curiosity, problem solving, and openness. The organization encourages collaboration, thinking big, taking risks in a blame-free environment, self-direction on meaningful projects, and providing support and mentorship for learning and growth.

Manufacturing & Electronics
Industry
10,001+
Employees
1993
Founded Year

Mission & Purpose

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.