Software Development Engineer - Data
Software Engineering
Cupertino, CA, USA
USD 150,400-225,300 / year + Equity
Posted on Jul 28, 2026
At Apple, the content that reaches hundreds of millions of people: music, films, books, podcasts, and apps, is underpinned by sophisticated data systems that run quietly and reliably at scale. The ASE Content Data Engineering team builds and maintains the reporting infrastructure, data pipelines, and data models that keep Apple's online media platform informed, compliant, and moving forward. We are looking for a Software Development Engineer - Data to join our Reporting & Data Models team. In this role, you will modernize and automate legacy data pipelines, build new reporting solutions, maintain a large and complex reporting estate across Oracle and cloud-native platforms, and help deliver compliance extracts that meet real-world legal deadlines across multiple jurisdictions. If you enjoy working on systems that matter, where your engineering choices directly impact operational accuracy, regulatory compliance, and business intelligence for one of the world's largest media platforms, we would love to hear from you.
The ASE Content Data Engineering team is responsible for the internal business intelligence infrastructure that powers Apple's online media store. The data models and reports developed by this team supply information to engineering, business operations, and external content providers regarding the status of online media, including workflow status, availability, and incoming and outgoing volume. This engineer will own and evolve a large portfolio of data pipelines and reporting systems, driving the ongoing migration from legacy infrastructure to modern pipeline orchestration and cloud-based storage, while also delivering net-new reports and data models in response to business, engineering, and regulatory needs. The role demands a blend of engineering rigor, data fluency, and a strong reliability mindset: the team's reports directly support legal and compliance obligations with hard, externally-imposed deadlines. The ideal candidate brings a bias toward automation and self-healing pipelines over manual intervention, and is comfortable managing a broad portfolio of concurrent work, balancing urgent pipeline fixes against longer-horizon migration and enhancement projects. They can independently troubleshoot and resolve ambiguous, production-impacting issues under time pressure, and communicate technical findings clearly to non-engineering stakeholders across legal, business operations, and content provider teams. Knowledge of or genuine curiosity about online media is a plus.
- Design, build, and maintain automated data pipelines and reporting systems across a hybrid stack spanning Oracle, Spark, Hadoop, and cloud object storage
- Drive the migration of legacy report jobs to modern pipeline orchestration and cloud infrastructure, ensuring reliability, observability, and maintainability
- Investigate and resolve pipeline failures, performance bottlenecks, and data quality issues across a large, interdependent reporting estate covering all content domains
- Develop and maintain legal and regulatory-compliance reports where accuracy, auditability, and on-time delivery are non-negotiable
- Onboard and publish curated datasets from operational data stores to the Apple Data Lakehouse, ensuring proper data governance, lineage documentation, and PII tagging
- Extend and enhance internal content data platform models across all content domains: propagating schema changes, adding new columns and descriptors, and maintaining the accuracy of the platform's content-intelligence surface
- Execute controlled, legally-gated production data modifications and archival/deletion programs under change control and with required legal approval
- Contribute to the evolution of AI-assisted data tooling (e.g., Text-to-SQL), automating and accelerating how stakeholders access and query content data
- Minimum 1 year of experience as a Software Engineer, Data Engineer, or equivalent role
- Strong experience with Spark and Spark SQL, including performance tuning and resource optimization
- Strong experience with Hadoop, HDFS, and/or S3-compatible object storage
- Hands-on experience with Python scripting for data processing and automation
- Experience writing SQL and PL/SQL on Oracle, including query optimization
- Experience building and maintaining automated data pipelines and job scheduling systems (e.g. Airflow)
- Experience with CI/CD workflows and version-controlled deployment of data pipeline code
- Hands-on experience with Kubernetes for container-based workload execution
- Bachelor's degree in Computer Science, Software Engineering, Data Science, or a related technical field OR Hands-on engineering experience and a demonstrable track record of building and maintaining production data systems are weighted equally alongside formal academic credentials.
- Familiarity with Iceberg / Parquet-based table formats
- Proficiency in macOS and Linux server environments
- Experience using GenAI tools (e.g., Claude, OpenAI tools) to accelerate engineering tasks
- Familiarity with cloud object storage migration patterns (HDFS to S3)
- Experience with EKS or other cloud-based compute platforms
- Familiarity with Trino or other distributed SQL query engines
- Experience with Flink or Flink SQL for stream-based data processing
- Familiarity with GenAI-assisted data access patterns (e.g., MCP servers, Text-to-SQL, LLM-assisted data documentation)
- Familiarity with Tableau or similar data visualization and dashboarding tools
- Experience with shell scripting for automation and pipeline support
- Experience with Java development