It all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone.
Benefit options available through Magnit Global, depending on contract factors and upon meeting requirements.
About The Work
We are seeking a Senior Systems Engineer to join our Hyperscaler Performance & Reliability Engineering (HPRE) team. This role is focused on improving the reliability, performance, and operational excellence of infrastructure services running across AWS, Azure, and GCP.
Operating within a Site Reliability Engineering (SRE)-inspired model, you will own infrastructure-level reliability challenges spanning compute, storage, and networking services. You will investigate production issues, develop performance insights from telemetry and observability platforms, validate infrastructure designs through testing, and implement engineering improvements that reduce customer-impacting incidents.
This is a hands-on engineering role for someone who enjoys troubleshooting complex distributed systems, analyzing system behavior at scale, and turning operational learnings into lasting reliability improvements. The ideal candidate combines strong cloud infrastructure experience with a data-driven approach to performance tuning and operational excellence.
Key Responsibilities
- Own reliability and performance investigations for compute, storage, and network-related issues across AWS, Azure, and GCP, driving incidents from detection through root-cause analysis and remediation.
- Design, implement, and maintain observability solutions using hyperscaler-native monitoring platforms and internal telemetry systems to identify bottlenecks, capacity constraints, and performance regressions.
- Analyze customer workload behavior using metrics, logs, traces, and infrastructure telemetry to optimize resource utilization, latency, throughput, and availability.
- Develop and maintain Infrastructure-as-Code, automation, and operational tooling that improve reliability, reduce toil, and enable safe infrastructure changes at scale.
- Execute performance, load, stress, and resilience testing of cloud infrastructure platforms and validate infrastructure changes before production adoption.
- Identify and drive reliability improvements, operational processes, and technical projects within the team while serving as a trusted technical resource for peers and partners.
Required Qualifications
- Strong understanding of cloud compute, storage, and networking fundamentals, including virtual machines, block storage, VPC/VNet design, load balancing, routing, and access management.
- Experience implementing and operating observability platforms, including metrics, logging, tracing, alerting, and performance analysis workflows.
- Proficiency with Infrastructure-as-Code and automation technologies such as Terraform, Git-based workflows, CI/CD pipelines, and scripting with Python or Bash.
- Demonstrated ability to independently troubleshoot complex production issues, perform structured root-cause analysis, and implement preventative solutions.
- Experience working within SRE, platform engineering, infrastructure engineering, or cloud operations environments where reliability, scalability, and operational accountability are core responsibilities.
Preferred Qualifications
- Experience with Kubernetes, GitOps, and declarative infrastructure management models.
- Familiarity with CloudWatch, Azure Monitor, Google Cloud Monitoring, Splunk, OpenTelemetry, or similar observability ecosystems.
- Cloud platform or Kubernetes certifications (AWS, Azure, GCP, CKA, CKAD, or equivalent).
- 3+ years of hands-on experience operating and troubleshooting production infrastructure in at least one major public cloud, with working knowledge of AWS, Azure, and/or GCP services.
Equal Opportunity Employer
Magnit Global is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.