waypointjobs

Oracle

Principal Software Engineer, Core Infrastructure (Large Infrastructure)

nashville, TN

Check who can apply and the requirements below before continuing.

About this opportunity

Oracle lists this Principal Software Engineer, Core Infrastructure (Large Infrastructure) opportunity in nashville, Tennessee. Review the employer’s description below for duties, qualifications and application requirements.

Job description

hackajob is collaborating with Oracle to connect them with exceptional professionals for this role.

About the team

Scaled Manufacturing (SMF) builds and operates container-based services for validating infrastructure across factory environments. SMF enables repeatable production of liquid-cooled GPU platforms - including GB200, GB300, VR, MI355, and MI455 - by embedding quality, yield, throughput, and readiness controls into manufacturing.

SMF delivers factory infrastructure; a Control Plane for test lifecycle management, operator workflows, quality gates, fleet health, and reporting; and a Data Plane that automates tests validation. Manufacturing Repair and Triage identifies root causes, guides recovery and uses data to reduce repeat failures. Working with Supply Chain Operations, Oracle Hardware Development, manufacturing partners, data-center teams, NVIDIA, and AMD, SMF detects hardware, firmware, software, and configuration issues before racks ship. This improves yield, reduces rework and downstream failures, and accelerates reliable AI infrastructure deployment at hyperscale.

Description

* Leads development and begins architecting scalable, container-based services that build, validate, and monitor liquid-cooled GPU infrastructure across factory environments.

* Develops secure infrastructure, control-plane workflows, and data-plane test capabilities that orchestrate manufacturing test lifecycles, operator workflows, quality gates, fleet status, and business reporting for platforms including GB200, GB300, VR, MI355, and MI455.

* Automates and maintains GPU test validation; manages consistent firmware, software, and hardware configuration; and develops repair and triage capabilities that identify root causes and guide recovery.

* Partners with Supply Chain Operations, Hardware Development, external manufacturing partners, data-center operations, NVIDIA, and AMD to resolve issues before racks ship.

* Establishes manufacturing yield, throughput, quality, and deployment-readiness metrics that reduce rework and downstream failures, improve data-center ingestion, and accelerate reliable hyperscale AI infrastructure delivery.

Responsibilities

Responsibilities

* Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.

* Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.

* Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.

* Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.

* Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.

* Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.

* Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.

* Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.

* Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.

* Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.

Qualifications

* Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.

* 7+ years of professional software-engineering experience, with demonstrated impact on large-scale distributed systems or cloud infrastructure.

* Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.

* Proficiency in one or more object-oriented or systems programming languages, such as Java, C++, C#, or Go.

* Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices.

* Experience with system-level test automation, performance/load testing, reliability engineering, and production incident response.

* Demonstrated experience leading or influencing technical architecture and mentoring engineers.

* Strong problem-solving, communication, and cross-functional collaboration skills.

Preferred Qualifications

* Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms.

* Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.

* Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.

* Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure.

Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

BNY

WhatJobs

Vice President, Dynamics 365 CE Developer

lake mary, FL

Salary not specified

hackajob is collaborating with BNY to connect them with exceptional professionals for this role. Vice President, Dynamics 365 CE Developer At BNY…

Last received from source 2026-10-10View job

J.P. Morgan

WhatJobs

Lead Software Engineer - AI & Machine Learning

columbus, OH

Salary not specified

hackajob is collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION You will influence outco…

Last received from source 2026-10-10View job