waypointjobs

talentpluto

RL Environment Software Engineer

Remote — United States (see country and timezone requirements)

Check who can apply and the requirements below before continuing.

Job description

Location: Remote (United States)

Work Model: Remote

Industry: Applied AI / AI research data

Compensation: $180K-$220K base, ~$400K+ OTE (uncapped profit share)

About the Company

Our partner is a fast-growing applied AI research lab that builds high-quality reinforcement-learning environments and agents sold to the world's leading AI labs. In under two years they have scaled to a nine-figure revenue run rate and grown their team severalfold in a matter of months, backed by leading venture investors. Quality is their core differentiator, and they are rapidly expanding into new domains.

The Opportunity

As an RL Environment Software Engineer, you will sit at the intersection of research engineering and traditional software engineering, building the environments that simulate real-world workflows and the agents that automate them. This is forward-looking work, you will help research and predict what high-quality environments the frontier will need next, then build them from the ground up.

You will join a brand-new RL team being assembled with exceptional talent, with a clear path to grow alongside it as the function scales into industry pods.

Responsibilities

Design and build high-quality RL environments that simulate real working environments end to end.

Develop agents for the tasks within those environments and iterate until they are efficient and production-ready.

Partner with the research team to scope which environments to build and why, staying ahead of future demand rather than only meeting present needs.

Own the backend and infrastructure layers that make environments reliable and scalable.

Help set engineering standards for a zero-to-one team as the RL function grows.

Requirements

Strong machine-learning engineers who code heavily and build systems from scratch, with strong intuition for reinforcement learning.

Proficiency across a modern stack, Node.js and Python on the backend and React/TypeScript on the frontend, with strong Kubernetes and Docker skills.

Comfort operating in a fast-paced startup environment with high ownership and long hours.

A track record of meaningful tenure and impact at previous companies.

Reinforcement-learning experience or an RL research background is a strong plus, though not required.

Bachelor's degree in computer science or a related technical field, or equivalent practical experience.

Originally posted on Himalayas

Who can apply

Eligible countries: United States. Accepted UTC offsets: UTC-10, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.

Ready for your next step?Apply on the official website
Apply on Himalayas ↗

Explore related searches

Current related jobs

infisical

Jobicy

Senior Full Stack Engineer

Remote — Brazil, Canada, Europe, USA

Salary not specifiedRemote

Infisical is the open source security infrastructure platform that engineers use for secrets management, certificates, and privileged access mana…

Listing review due 2026-10-07View job

Spotify

Jobicy

Data Scientist - Music Mission

Remote — USA

Salary not specifiedRemote

The Music Mission enables music creators to grow, engage, and monetize their fan bases on Spotify. Central to the Music Mission's vision is the d…

Listing review due 2026-10-07View job

Spotify

Jobicy

Data Scientist - Music Promotion

Remote — USA

Salary not specifiedRemote

The Music Mission enables Music creators to grow, engage & monetize their fan bases on Spotify. Central to the Music Mission's vision is the deve…

Listing review due 2026-10-07View job

infisical

Jobicy

Strategic Finance

Remote — Canada, USA

Salary not specifiedRemote

Infisical is the open source security infrastructure platform that engineers use for secrets management, certificates, and privileged access mana…

Listing review due 2026-10-07View job