Staff Observability Engineer - Golang, Terraform, AI

Intuit
Intuit

Software Engineering, Data Science

Multiple locations

USD 188,500-274k / year + Equity

Posted on Aug 26, 2026

Staff Observability Engineer - Golang, Terraform, AI

Category Software Engineering Location Charlotte, North Carolina; Oakland, California; San Diego, California Job ID 23517

Company Overview

Intuit is the global financial technology platform that powers prosperity for the people and communities we serve. With approximately 100 million customers worldwide using products such as TurboTax, Credit Karma, QuickBooks, and Mailchimp, we believe that everyone should have the opportunity to prosper. We never stop working to find new, innovative ways to make that possible.

Job Overview

Come join Intuit Credit Karma's Observability team as a Staff Site Reliability Engineer. The team owns the telemetry platform that 700+ engineers rely on to understand, operate, and troubleshoot ~250 microservices in production. The footprint is large: roughly 140 TB/day of metrics, traces, and events plus 80 TB/day of logs (about 80 PB/yr combined), flowing through OpenTelemetry collectors, Telegraf, Pub/Sub-based routing, Splunk, and New Relic, running on GKE across multiple production clusters on GCP.

This is a builder's role, not a dashboard-tuning role. Our platform is code: obs_nrtf, a Terraform monorepo with GitOps workflows, automated NRQL linting, daily state reconciliation, and CI/CD that lets several hundred engineers manage their own alerting and dashboards without touching Terraform internals. Alongside it sit our canary and regional-failover tooling and o11y, our internal service platform and Slackbot.

Our tracing stack is already OTel-native. The next two years are defined by the harder half: re-architecting the telemetry data plane onto Pub/Sub, moving our custom metrics pipeline off legacy InfluxDB line protocol and pod-level gauges onto OTel Metrics, and building the streaming aggregation and post-processing layer that will decide whether our observability spend scales with the business or with our traffic. You will lead that work end to end: the design, the code, and the multi-quarter program management that gets several hundred service owners across the finish line.


Responsibilities

  • Own the platform as software. Extend and operate obs_nrtf and our canary and failover tooling: Terraform modules, Go services and CLIs, CircleCI pipelines, linters, and validation that catch bad config before it reaches production. You will be writing code most weeks.
  • Lead large-scale migrations end to end. The Pub/Sub telemetry data plane, custom metrics onto OTel Metrics, service-level metric re-aggregation, and New Relic estate consolidation. Several of these run concurrently and touch every service in the company. Sequencing them, sizing the blast radius, running dual-write periods, publishing timelines, and driving hundreds of service owners to adopt is as much of the job as the engineering.
  • Solve real-time telemetry at scale, cost-effectively. Design the streaming ingest, routing, enrichment, and windowed-aggregation layer that carries our full telemetry volume. Metrics dominate our ingest bill, and the levers (resolution reduction, aggregation, cardinality control, consumption governance) all live in this layer. This is the single biggest technical problem on the team.
  • Apply AI where it earns its place. Two fronts: ML on telemetry, including anomaly detection, sensitivity tuning, and forecasting, so that alerting gets sharper rather than noisier; and AI-assisted engineering, using coding agents and MCP-based tooling to accelerate migration mechanics, config refactors, and documentation across hundreds of repos.
  • Set technical direction and raise the bar. Author TDDs, drive semantic conventions and SLO/alerting standards, review designs across the team, and mentor engineers. You will be the deepest platform voice in the room, working directly with App Platform, PaaS, Security, and our GCP and New Relic partners.
  • Participate in oncall for the observability platform.

Qualifications

What you'll bring

  • 8+ years in software engineering, SRE, infrastructure, or platform engineering, with a meaningful stretch of it building platforms other engineers depend on.
  • Strong Go. You have shipped and operated production services, pipelines, or developer tooling in Go, not just scripts.
  • Deep Terraform. Module design, state management, provider behavior, drift reconciliation, and the operational reality of a large shared IaC monorepo. Experience migrating state and refactoring modules without orphaning resources.
  • Real-time data experience. Streaming pipelines at high volume (Pub/Sub, Kafka, Kinesis, Dataflow, Flink, or equivalent) with hands-on work on aggregation, windowing, backpressure, and the cost and correctness tradeoffs that come with them.
  • A track record leading enterprise-scale migrations. You have taken a platform migration from proposal to done across many teams. You can talk concretely about how you sequenced phases, handled dual-write or dual-read periods, drove adoption from teams who did not ask for the work, and knew when to cut scope.
  • Program management instincts. Comfortable running multiple concurrent multi-quarter efforts: tracking adoption, communicating status upward and outward, unblocking other teams, and holding a timeline without a PM doing it for you.
  • GCP fluency (GKE, Pub/Sub, IAM, BigQuery, billing and cost surfaces) and CI/CD depth (CircleCI or similar) for pipelines that gate infrastructure changes.
  • Observability fundamentals. Metrics, logs, and traces as a system: cardinality, sampling, retention, resolution, and the cost model behind each.
  • Strong written communication. Much of the influence in this role happens through design docs, migration guides, and runbooks.

What we'd like to see

  • Hands-on OpenTelemetry: collector configuration and operation, pipeline design, semantic conventions, and OTel Metrics. Experience with tail-based sampling on a stateful collector tier is a strong plus; it's a design problem we have ahead of us.
  • New Relic at scale (NRQL, alerting, entity synthesis, multi-account governance) or comparable depth in Datadog, Grafana Cloud, or Chronosphere.
  • Splunk administration: index and ingestion management, search performance, RBAC, retention.
  • Streaming aggregation applied directly to vendor cost reduction. If you have built pipelines that materially cut a telemetry bill, we want to talk to you.
  • Practical use of LLMs and coding agents in an engineering workflow, especially applied to large-scale codebase or config migration.
  • ML applied to operational telemetry: anomaly detection, forecasting, alert-noise reduction.
  • Kubernetes at production scale, Helm, and GitOps patterns.
  • Scala or TypeScript familiarity, useful for working inside our service frameworks.
  • Experience in a regulated or high-sensitivity data environment (PII handling, DLP in telemetry pipelines).

Footer

Intuit provides a competitive compensation package with a strong pay for performance rewards approach. This position may be eligible for a cash bonus, equity rewards and benefits, in accordance with our applicable plans and programs (see more about our compensation and benefits at IntuitĀ®: Careers | Benefits). Pay offered is based on factors such as job-related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender.

The expected base pay range for this position is:
Oakland, CA $202,500- $274,000
San Diego, CA $188,500- $255,000