It all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.
Benefit options available through Magnit Global, depending on contract factors and upon meeting requirements.
We are seeking a highly motivated Site Reliability Engineer (SRE) with a strong operational focus to join our growing team. In this role, you will play a vital role in ensuring the smooth operation and performance of our critical infrastructure and services. You'll work cross-functionally to create alignment and deliver results alongside builders who have helped to shape the success of companies such as Google, Okta, AWS, Snowflake.
What you will do in this role:
- Deploy software for Cloud Prem and SAAS customers.
- Respond to and diagnose system incidents in a timely and efficient manner, minimizing downtime
and impact on users. - Collaborate with other engineers to establish root causes and implement effective resolutions.
- Continuously improve incident response processes and documentation for future occurrences.
- Proactively monitor and maintain the health and performance of our infrastructure and services.
- Perform routine administrative tasks such as system configuration, user management, and data
backups. - Identify and implement operational improvements to ensure ongoing system reliability and
efficiency. - Develop and implement scripts and automated solutions to streamline operational tasks and
reduce manual workload. - Participate in the on-call rotation to address critical incidents outside of regular business hours.
- Ensure effective handoff between on-call engineers and document post-incident information for future reference.
- Document processes for support and create, maintain and execute run-books for identified situations
- Provide tier 2/3 technical support to customers experiencing platform issues or requiring
advanced troubleshooting - Work directly with customer technical teams to resolve complex deployment, configuration, and integration challenges
- Conduct technical onboarding sessions and provide guidance on best practices for customer
implementations - Collaborate with customer success teams to ensure smooth customer experiences and rapid
issue resolution - Create and maintain customer-facing technical documentation, troubleshooting guides, and knowledge base articles
- Escalate customer feedback and feature requests to product and engineering teams
- Participate in customer calls and technical discussions to provide expert-level platform guidance
- Track and analyze customer support metrics to identify trends and areas for improvement
What you will need to be successful in this role:
- Education:BS degree in Computer Science or related field
- Experience:3+ years of experience in Site Reliability Engineering
- 2+ years experience working with cloud platform and cloud automation tools especially in AWS
- Strong experience with Kubernetes, Helm, Linux, AWS networking(VPC) and Terraform
- Experience with the GitOps model for deployment
- Familiarity with distributed version control
- Experience with monitoring and alerting tools (e.g., Prometheus, Grafana).
- Bazel and CueLang experience a plus
- Understanding of software configuration best practices
- Ability to wear multiple hats in a fast-paced environment
- Hands-on, “can do” attitude and a bias for action
- Low ego and high intellectual curiosity
- Comfortable working across time zones to support global customer base
- Excellent communication skills with ability to explain technical concepts to both technical and non-technical audiences
- Strong customer service orientation with patience and empathy when working with frustrated customers
Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation.