AI/ML Engineer, Amazon Global Data Center Ops Central Insight and Analytics Team

Amazon
Amazon

Software Engineering, Operations, Data Science

Seattle, WA, USA

Posted on Aug 25, 2026

Description

We are looking for an **AI/ML Engineer** to build, deploy, and operate the ML/AI systems that power the agentic decision intelligence workflow we are building. You are the person who takes a model from a notebook to production, builds the LLM integration layer, implements RAG pipelines, creates evaluation frameworks, and ensures our AI systems are reliable, observable, and continuously improving.

This is a hands-on engineering role with deep ML/AI focus — you write production code that runs AI systems, not research papers. If you love the intersection of ML infrastructure, LLM applications, and production engineering, this role is for you.

Key job responsibilities
- Build and maintain LLM-powered components: structured reasoning chains, narrative generation, recommendation rationale
- Implement and optimize prompt engineering pipelines with version control, A/B testing, and regression detection
- Build RAG (Retrieval-Augmented Generation) systems that ground LLM outputs in operational data, historical playbooks, and domain knowledge
- Build guardrails, validation layers, and output parsing for LLM responses. Optimize latency, cost, and quality trade-offs across LLM providers
- Deploy ML models to production. Implement model monitoring: drift detection, performance degradation alerts, automated retraining triggers
- Build A/B testing infrastructure for model experiments. Manage model versioning, rollback, and canary deployment. Ensure SLA compliance for inference latency and availability
- Own the operational health of AI/ML services: monitoring, alarming, on-call, incident response, observability across the AI stack (prompt traces, latency histograms, token usage, error rates)
- Write comprehensive tests (unit, integration, end-to-end) for ML pipelines