Sr. Observability Engineer
Lehi, Utah, United States Hybrid
Salary not listedJobFig found this opening at its original source and checks that it remains available.
About the role
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand. This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
What you'll bring
- BS in Computer Science or equivalent experience
- 5+ years running production observability, SRE, or DevOps
- 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency
- AI- and workflow-literate.
- You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs
- Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis
- Shared on-call, Incident Commander-capable
- Fintech experience with MX-like architectures
- Google SRE practices: toil elimination, incident management, automation for self-healing
- Cross-functional influence without authority.
- You've improved teams that don't report to you
- Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)