Software Dev Engineer, Software Controller, CLOS Fabric Engineering, AWS
Accounting & Finance
Santa Clara, CA, USA
Description
Are you a developer passionate about building incredible software? Committed to quality, agility, and consistency? Curious how software and AI/ML can operate and maintain one of the world's largest networks?
If so, come join us — automate the future of global networks at unprecedented scale, and thrive in a team where your voice matters.
AWS Infrastructure Services is looking for a Software Development Engineer to join the Data Center Network organization. We're the people who keep the cloud running — owning the design, planning, delivery, and operation of all AWS global infrastructure. In this role, you'll work with customers, leadership, and peers to automate and invent new ways of operating the AWS network. You'll build best practices, improve operational procedures, and deliver iterative impact with a proactive mindset. You'll collaborate across AWS to deliver the highest standards of safety and security while providing seemingly infinite capacity at the lowest possible cost. You'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion. You'll be responsible for building software services that deploy and scale the Amazon networks supporting AWS, customers, and other business units across multiple global datacenters.
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.
You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
Key job responsibilities
- Shape the automation future in networking
- Build and maintain production-quality distributed systems at startup pace
- Develop tools and processes that collect and rationalize data from multiple sources to reduce workloads
- Use data to measure success; take ownership of service quality and proactively prevent customer-facing faults
- Partner with Network Engineering teams for fast, smooth software roll-outs
- Identify and troubleshoot recurring platform issues; escalate effectively to senior engineering teams
- Design and build cloud-computing system software for a diverse set of customers
- Mentor junior engineers and bring clarity to ambiguous situations
- Identify and implement optimizations for performance, scalability, and efficiency
A day in the life
You will work with customers to gather requirements and generate technical designs, and you will carry the project through all the software lifecycle stages. You’ll develop products that enable builders to develop and operate robust, high-quality software and safely, securely, and reliably deploy it. You will use your technical expertise and communication skills to mentor other engineers and provide training and support for our technologies. You will have access to leadership and engineering staff.
About the team
The AWS DCFC (DCF Controllers) team owns the software controller within the CLOS Fabric Engineering (CFE) organization. The team is responsible for building software that monitors the network/fabric state, recovers out-of-service capacity, enables network scaling, manages IP allocation, and plans and executes decommissioning.
The broader CFE organization focuses on designing cost-effective CLOS topologies, maximizing port availability, and managing the full platform lifecycle—from hardware qualification through decommissioning. CFE also ensures configuration and OS compliance across the fleet, and leads OS promotion, testing, and deployment.
Meet Matt, VP, Core Networking --- https://youtu.be/DqTStjRtjX4