About this opportunity
Meta lists this Principal Software Engineer, Systems opportunity in cheyenne, Wyoming. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Summary
Meta is seeking an experienced Software Engineer to build the next generation of AI systems for production reliability. This role sits at the intersection of large-scale distributed systems, incident response, and applied AI. You will lead the development of an AI agent that can autonomously investigate production incidents, identify likely root causes, create safe mitigation plans, and execute them while humans supervise and can intervene. You will improve the agent's accuracy, autonomy, and real-world impact, while identifying other opportunities to apply agentic systems across incident prevention, detection, mitigation, observability, and infrastructure operations. This may include improving the agent's architecture, context, and tooling, or adapting foundation models through fine-tuning, reinforcement learning, distillation, pruning, and inference optimization.This is an applied engineering role, not a research position. The ideal candidate combines deep infrastructure expertise with sound judgment about when and how to apply AI to production problems, a long-term technical vision, and the ability to move quickly from an ambitious idea to a production system operating safely at Meta scale.This is also a deeply hands-on senior IC role. You will personally prototype, build, evaluate, and ship AI systems, working directly in the code and with production data. A central challenge is creating systems that improve from experience: learning from outcomes, identifying their own failure modes, testing changes, and safely increasing their effectiveness over time.
Required Skills
Principal Software Engineer, Systems Responsibilities:
Define the technical vision and architecture for agentic reliability systems across Meta
Personally design, code, and ship production agentic systems for complex infrastructure problems
Lead the development of an AI agent for production incident investigation and mitigation, advancing it toward accurate, trusted, and safely supervised autonomous action
Develop major improvements in agent reasoning, context, tool use, planning, evaluation, learning, and safe execution
Explore and apply techniques including automated hill climbing, fine-tuning, reinforcement learning, model routing, distillation, pruning, and inference optimization, selecting the simplest approach that produces measurable gains
Identify new high-value applications of AI across incident prevention, detection, mitigation, observability, and infrastructure operations
Move rapidly from ambiguous problems to prototypes, validate them against real production workloads, and develop successful approaches into reliable systems at scale
Build evaluation and experimentation systems that connect agent quality to outcomes such as investigation accuracy, successful mitigation, incident duration, and reduced operational work
Build closed-loop improvement systems that turn production outcomes into evaluations, experiments, and better agent behavior
Establish architectures and guardrails for production actions, including authorization, independent validation, auditability, rollback, and human oversight
Partner with Infrastructure, AI, Product, and Reliability leaders to integrate agentic capabilities into Meta’s production ecosystem
Influence technical strategy across organizations and mentor other engineers working on distributed systems and applied AI
Minimum Qualifications
Minimum Qualifications:
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
12+ years of software engineering experience, including experience building and operating large-scale distributed or infrastructure systems
Experience setting technical direction and leading complex, multi-year engineering efforts across organizational boundaries
Experience applying AI or machine learning systems to production problems
Demonstrated experience moving from technical concept to production deployment and measurable impact
Experience diagnosing complex production systems using telemetry, code, configuration, and dependency information
Experience coding in languages such as C++, Java, Python, Rust, or equivalent
Recent hands‑on experience building and shipping complex production systems, with the ability to move directly between architecture, experimentation, debugging, and implementation
Experience influencing senior engineers and leaders without direct organizational authority
Preferred Qualifications
Preferred Qualifications:
Experience with autonomous or semi-autonomous production actions and their safety, authorization, and rollback mechanisms
Experience building systems that operate at hyperscale under strict reliability and latency requirements
Experience improving agent quality through evaluation, context engineering, fine‑tuning, reinforcement learning, distillation, or model‑serving optimization
Track record of identifying unconventional opportunities, rapidly prototyping solutions, and changing the technical direction of a large organization
Deep expertise in observability, incident response, change safety, distributed systems, or production infrastructure
Experience building self‑improving or self‑evolving systems, including automated experimentation, feedback loops, hill climbing, reinforcement learning, or recursive self‑improvement
Experience building production AI agents that reason across telemetry, code, configuration, deployments, and operational knowledge
Public Compensation
$271,000/year to $347,000/year + bonus + equity + benefits
Industry
Internet
Equal Opportunity
Meta is proud to be an Equal Employment Opportunity and We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at
#J-18808-Ljbffr
Worksite address
cheyenne, WY, 82007, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.