Data Engineer II, AWS Analytics Engineering
Data Science
Seattle, WA, USA
Description
The AWS Analytics Engineering (AAE) organization is the analytics backbone of AWS — we build and operate the data platform that powers business decisions across more than 150 AWS services. Every insight surfaced to AWS product leadership, from service adoption trends to revenue drivers, flows through systems our team designs, builds, and maintains. We operate at massive scale — processing petabytes of data daily through thousands of jobs consisting of transformations, reporting queries, ingestions, and infrastructure management scripts. As AWS launches new services and features at an accelerating pace, the need for a unified, self-service Data Ingestion Platform has become critical.
We are seeking a Data Engineer to join our Data Ingestion Platform team. In this role, you will design, build, and operate the next-generation ingestion service that automatically onboards data from any new AWS service or feature at launch — without manual intervention. You will build scalable ingestion frameworks capable of handling petabytes of data, implement big data procurement pipelines, and establish the governance, compliance, and audit controls that ensure all data flowing through the platform meets AWS’s security and data integrity standards.
The platform you build must scale to N services — as AWS’s service portfolio grows, your ingestion infrastructure grows with it. You will work on event-driven architectures, self-registration mechanisms, and configuration-driven onboarding so that new data sources are ingested automatically, reliably, and with full lineage and audit trails.
The ideal candidate is a strong engineer who is passionate about building platforms, thinks in terms of frameworks and abstractions rather than one-off pipelines, and is excited about building the foundational layer that enables all of AAE’s analytics capabilities.
Key job responsibilities
Design and build reusable, configuration-driven ingestion frameworks and abstractions — not one-off pipelines — that automatically onboard new AWS services and data sources at launch without manual intervention. Architect for scale-to-N, ensuring the platform grows with AWS's expanding service portfolio.
Implement event-driven architectures and self-registration mechanisms that enable service teams to onboard data sources through standardized, self-service interfaces.
Build and operate big data procurement pipelines at petabyte scale — ensuring reliability, fault-tolerance, cost-effectiveness, and performance across thousands of concurrent ingestion jobs.
Establish governance, compliance, and audit controls across the ingestion platform — enforcing full data lineage, provenance tracking, and security/data integrity standards for every source flowing through the system.
Identify and resolve data quality issues proactively. Implement automated guardrails, data certification standards, and SLAs for ingestion completeness, freshness, and accuracy.
Produce high-quality platform code that is secure, maintainable, and understandable by engineers unfamiliar with the system. Make sound architectural trade-offs balancing pragmatic short-term delivery with sustainable long-term design.
Contribute to infrastructure decisions within the team's data architecture. Efficiently manage compute, storage, and AWS infrastructure resources — building solutions that are stable, performant, and cost-optimized.
Own operational health for ingestion systems — monitoring, alarming, runbooks, on-call rotation, and incident resolution. Drive automation to reduce operational toil.
Drive engineering best practices through code reviews, design discussions, and operational reviews. Establish standards for dependency management and continuous improvement.
Collaborate cross-functionally with upstream service teams and downstream analytics consumers. Resolve differing technical perspectives and build consensus on platform direction.
Mentor peers, participate in hiring, and contribute to team knowledge-sharing.