Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Operations, Data Science
Seattle, WA, USA
Description
We are looking for an seasoned Applied Scientist to design, build, and deploy the ML/AI models that power our decision intelligence platform. You will work at the intersection of causal inference, time-series forecasting, anomaly detection, and LLM-based reasoning — all applied to real operational problems with measurable business impact.
Key Job Responsibilities
Decision Intelligence Models
- **Causal inference & root cause analysis:** Build models that decompose fleet-wide metric movements into root causes, distinguishing correlation from causation across operational dimensions (site, service, failure mode, time)
- **Dose-response modeling:** Develop models that learn the quantitative relationship between intervention intensity and outcome magnitude
- **Forecasting & projection:** Build time-series models that project metric trajectories under different intervention scenarios, enabling "if we do X, expect Y by date Z" recommendations
- **Anomaly detection & trend identification:** Develop multi-variate anomaly detection that distinguishes signal from noise in noisy operational data, and identifies emerging patterns before they become crises
- **Confidence calibration:** Build and maintain calibrated confidence scores for recommendations, ensuring the system knows what it knows and what it doesn't
- **Outcome attribution:** Design experiments and causal methods to measure the true impact of interventions
LLM Integration & Reasoning
- **Structured reasoning:** Design LLM prompting architectures that reliably transform operational data into executive-quality narrative summaries, decision framings, and recommendation rationales
- **LLM evaluation:** Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time
- **RAG systems:** Design retrieval-augmented generation systems that ground LLM outputs in operational data, historical playbooks, and institutional knowledge
- **Progressive autonomy:** Design the trust-calibration system where AI gradually earns expanded authority based on demonstrated accuracy over time
Research & Production
- **End-to-end ownership:** Take models from research through production deployment — you ship, you monitor, you iterate
- **Experimentation:** Design A/B tests and quasi-experiments to validate model improvements and measure business impact
- **Stakeholder communication:** Translate complex scientific results into actionable insights for non-technical senior leaders