doordashusa
Senior/Staff Deep Reinforcement Learning Engineer - DoorDash Dot
<div class="content-intro"><p><img style="display: none; max-width: 100%;" src="https://click.appcast.io/greenhouse-te8/a31.png?ent=34&e=22630&t=1701374353806" width="1px"> <img style="display: none; max-width: 100%;" src="https://track.jobadx.com/v1/i.gif?utm_pixel=224e990b-8ff4-4287-8d5d-2ff09647f181&utm_ptz=EST&utm_rqt=track" alt="" width="1"></p></div><h2>About the Team</h2> <p>Our DD Labs team builds real-time autonomous delivery systems. The Planning & Decision-Making group is investing heavily in deep reinforcement learning to move beyond classical planning, learning policies that generalize across novel driving scenarios, handle long-tail edge cases, and improve continuously from large-scale fleet data. Our models jointly handle prediction and planning in a single unified architecture. Our stack is pure JAX end-to-end: the same code you train with is the code that runs on the robot. No C++ rewrites, no TensorRT export. A new policy goes from training to on-vehicle deployment in minutes.</p> <h2>About the Role</h2> <p>As a Senior/Staff Deep RL Engineer, you will design, train, and deploy deep reinforcement learning policies that make real-time driving decisions for our autonomous vehicles. You will own the full lifecycle, from problem formulation and reward design through large-scale distributed training to on-vehicle inference. You'll help define how learned components compose with the rest of the autonomy stack to produce robust, shippable behavior.</p> <h2><strong>You’re excited about this opportunity because you will…</strong></h2> <ul> <li>Formulate complex driving tasks as RL problems with well-shaped reward functions and expressive state/action representations.</li> <li>Design and train model-based deep RL agents using GPU-accelerated simulation at massive scale, including improving the simulator itself.</li> <li>Build and maintain distributed training infrastructure in JAX across large compute clusters.</li> <li>Build agentic optimization systems that automatically improve code, run experiments, analyze metrics, and iterate on RL policies with minimal human intervention.</li> </ul> <h2><strong>We’re excited about you because…</strong></h2> <ul> <li>BS/MS/PhD in CS, EE, Robotics, or a related field, with a strong foundation in reinforcement learning and deep learning.</li> <li>You have proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software</li> <li>Hands-on experience training RL agents at scale, ideally in robotics, autonomous driving, or other real-time decision-making domains.</li> <li>Proficiency in JAX or a similar functional ML framework; comfort with JIT compilation, vectorized environments, and GPU-accelerated simulation.</li> <li>Deep grasp of core RL concepts: policy gradients, value functions, exploration-exploitation, model-based RL, reward shaping, and sim-to-real transfer.</li> <li>Data-driven mindset: comfortable building experiment pipelines, analyzing training runs, and letting metrics guide architectural decisions.</li> </ul> <h3>Nice to Have</h3> <ul> <li>Publications at top venues (NeurIPS, ICML, ICLR, CoRL, RSS, ICRA) on RL or learned planning.</li> <li>Experience building or working with GPU-accelerated simulators for RL training.</li> <li>Track record of shipping a learned component in a production robotics or autonomous vehicle stack.</li> </ul> <p> </p> <p><strong>Notice Regarding Use of AI and Automated Tools: </strong>To streamline our hiring process, DoorDash utilizes an automated recruitment tool called Gem.</p> <p><strong>How it works: </strong>Gem assists our recruiting team by evaluating job related qualifications and characteristics in connection with hiring. The tool is designed and used to support - rather than replace - human decision-making; trained personnel make final decisions with meaningful human review
前往雇主官网申请
联系前请核实雇主身份、工作地点、薪资和用工条件;不要向陌生人提供银行卡密码或验证码。