Applied Scientist, Generative AI and Synthetic Data
Software Engineering, Data Science · Full-time
Sunnyvale, CA, USA
Description
Amazon Lab126 builds the synthetic data that trains the AI in Amazon's devices, spanning wearables, smart home, and streaming. The team is looking for Applied Scientists to advance the science of making synthetic data maximally useful to the models it trains, across video, audio, and 3D. This is generation, but it is also the full loop around it: deciding what data to generate through data-centric methods like importance sampling and curriculum design, combining it with real data, and measuring its effect on the models that ship to customers.
A large part of the opportunity lies in the physical environment and how it responds to activity around it, including fast-moving, rule-governed events where signals across modalities must stay aligned in space and time. You will develop generative and simulation methods, from diffusion and multimodal models to physics-based digital twins and sim-to-real transfer, grounded in precise 3D spatial understanding, and deliver science artifacts that feed the models shipping in Amazon's devices. This is a hands-on research role with technical lead scope, reporting to a manager within a cross-discipline organization of Applied Scientists and Software Development Engineers.
Key job responsibilities
- Develop and fine-tune generative models and 3D simulation pipelines across video, audio, and sensor modalities, advancing temporal coherence, identity preservation, controllability, and causality so that generation is action-conditioned and physically plausible.
- Develop training recipes that combine synthetic and real data effectively, including mixing ratios, curriculum ordering, and sampling strategies, and characterize when synthetic data helps, when it hurts, and why.
- Build evaluation methods that predict synthetic data's impact on downstream models without full training runs, using proxy signals and embedding-based diagnostics that shorten the feedback loop from days to minutes.
- Model structured, rule-governed activity so that generated multimodal data stays coherent across modalities, 3D space, and time, and build 3D-consistent scene and event representations that keep generated data spatially faithful.
- Validate all work by measured improvement on production models, ship science artifacts to products, and mentor junior scientists and interns.
A day in the life
You might start your morning reviewing experiment results from an overnight training run, then join a design discussion with engineers to plan how your model will be integrated into a production pipeline. After lunch, you could be prototyping a new approach to a problem your team recently identified, writing clean and well-documented code as you iterate. Later, you might pair with a teammate to review their research methodology or prepare findings for an upcoming science review.
About the team
Our team owns the synthetic data strategy that accelerates time to market across Amazon's devices business. Demand for this data is growing, and we are actively expanding into new product categories. Our goal is to own this strategy across multiple flagship product lines for the entire devices organization, dramatically reducing how long it takes to get new capabilities into customers' hands.
We are a cross-discipline group of Applied Scientists and Software Development Engineers who partner closely to move ideas from research to production. If you want your science to directly shape the AI powering the next generation of Amazon devices, this is an exciting place to build your career.