waypointjobs

The Mutual Group

Lead AI Infrastructure Operations Engineer

des moines, IA

Check who can apply and the requirements below before continuing.

Availability awaiting confirmation

We are waiting for a fresh update from the source. This page preserves the last received job details; current availability is not confirmed.

Job description

Job Description

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions.

Department

Information Technology

Job Description

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions. As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs. The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG’s managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations. This is a hands‑on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed‑services partners.

Work Arrangement

Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in‑office schedule. The initial training period may require additional in‑office days.

Accountabilities

Enable AI Infrastructure and Environments

Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases.

Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services.

In Collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG’s AWS managed-services provider and other technology partners.

Support operational readiness for AI solutions transitioning into production.

Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle.

Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases.

Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use.

Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.

Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure.

Operate and Improve Production AI Solutions

Work with managed services to Implement and maintain observability, logging, tracing, monitoring, dashboards, and alerting for production AI applications.

Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption, and business outcomes.

Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations, and supporting infrastructure.

Support root‑cause analysis and coordinate resolution with AI engineers, application teams, platform teams, and managed‑services providers.

Support AI‑related incident management, problem management, operational reviews, release validation, and production‑readiness activities.

Create and maintain service‑health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths.

Define and track service health indicators, SLIs, SLOs, and operational KPIs for production AI solutions.

Use telemetry, evaluations, and production data to validate fixes, releases, configuration changes, and system improvements.

Support AI Risk and Operational Governance

Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements.

Operationalize evaluation processes and support AI Governance team to monitor AI performance, regressions, drift, grounding, retrieval quality, and overall effectiveness.

Support the collection and retention of operational evidence, including model and prompt versions, evaluation results, incidents, exceptions, and corrective actions.

Identify material changes in AI behavior and help ensure they are evaluated, documented, and appropriately addressed.

Analyze trends and proactively identify degradation, drift, capacity constraints, reliability risks, and quality issues before they become production incidents.

Key Outcomes

AI teams receive timely and consistent infrastructure and operational support.

AI use cases move efficiently from experimentation to reliable production operation.

Reusable infrastructure and operational patterns are applied across AI initiatives.

End‑to‑end visibility exists across AI applications and supporting services.

AI quality, reliability, cost, usage, and business impact are consistently measured.

Production issues are detected, diagnosed, and resolved more quickly.

Releases result in fewer regressions and operational disruptions.

Secure, compliant, and audit‑ready AI platforms meeting enterprise governance and risk requirements.

Qualifications

Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent practical experience.

8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering.

Experience supporting cloud‑hosted, distributed, data‑intensive, or AI‑enabled production applications.

Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting.

Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines.

Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed‑services providers.

Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations.

Experience with automation, CI/CD pipelines, infrastructure‑as‑code, configuration management, and release‑management practices.

Experience operating or supporting LLM, generative AI, machine‑learning, RAG, or agent‑based applications in production is strongly preferred.

Familiarity with AI observability, evaluation, model monitoring, hallucination detection, grounding, retrieval quality, model or data drift, and prompt‑related risks.

Familiarity with vector databases, model APIs, AI gateways, prompt‑management platforms, or agent orchestration frameworks.

Strong analytical, documentation, communication, and cross‑functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives.

Experience in insurance, financial services, or another regulated industry, along with relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications, is preferred.

Pay Range

Anticipated Hiring Range:

$130,000 - $150,000 annual base salary depending on experience, qualifications, and geographic location

Benefits

Competitive base salary plus incentive plans for eligible team members

401(K) retirement plan that includes a company match of up to 6% of your eligible salary

Free basic life and AD&D, long‑term disability and short‑term disability insurance

Medical, dental and vision plans to meet your unique healthcare needs

Wellness incentives

Generous time off program that includes personal, holiday and volunteer paid time off

Flexible work schedules and hybrid/remote options for eligible positions

Educational assistance

Equal Opportunity Employer

The Mutual Group is an Equal Opportunity Employer. It is our policy to recruit, hire, train and promote individuals in all job classifications without regard to race, color, religion, sex, national origin, age, veteran status, disability, sexual orientation, gender identity or any other characteristic protected by law.

Know Your Rights: Workplace Discrimination is Illegal

Your Rights Under USERRA

Applicants requiring a reasonable accommodation due to a disability at any stage of the employment application process should contact .

The Mutual Group participates in the E-Verify program and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. You are protected from employment discrimination based on your citizenship status and national origin.

All offers of employment are contingent upon the successful completion of a background check.

#TMG

#J-18808-Ljbffr

Worksite address

des moines, IA, 50319, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Explore related searches

Current related jobs

Amazon Data Services, Inc.

WhatJobs

Data Center Chief Engineer

sparks, NV

Salary not specified

Join our dynamic Data Center Engineering Operations Team and become a critical architect of the infrastructure that powers global cloud computing…

Listing review due 2026-10-08View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive business …

Listing review due 2026-10-08View job

Kohler

WhatJobs

Engineer, New Product Integration

kohler, WI

Salary not specified

Engineer, New Product Integration Work Mode: Onsite Location: Onsite, four days per week - Kohler, WI Opportunity This is mo…

Listing review due 2026-10-08View job

GE Vernova

WhatJobs

Principal Engineer - AI Engineering

niskayuna, NY

See pay details in description

Job Description Summary GE Vernova is embracing cutting-edge technologies to streamline operations, improve customer experiences, and drive gr…

Listing review due 2026-10-08View job

GE Vernova

WhatJobs

Lead Product Safety and Compliance Engineer

rochester, NY

See pay details in description

Job Description Summary The Product Safety & Compliance Engineer works directly within the engineering development team to ensure our products…

Listing review due 2026-10-08View job

Hobbs Brook Real Estate

WhatJobs

Commercial Facilities Engineer

waltham, MA

$30.88 to $38.61 per hour

Job Description: Hobbs Brook Real Estate LLC is an innovative commercial real estate leader with a portfolio of forward-thinking, sustainable pr…

Listing review due 2026-10-08View job