waypointjobs

Critical Solutions

Mid - Systems Engineer – MECM / Site Reliability Engineering (SRE) (w/ active Secret)

norfolk, VA

Check who can apply and the requirements below before continuing.

About this opportunity

Critical Solutions lists this Mid - Systems Engineer – MECM / Site Reliability Engineering (SRE) (w/ active Secret) opportunity in norfolk, Virginia. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Mid - Systems Engineer – MECM / Site Reliability Engineering (SRE) (w/ active Secret)

Location: Norfolk, VA and Pearl Harbor, HI

Clearance: Secret

Type: Full-time, On-site

Salary Range: $85,000 - $110,000

PRIMARY ROLES AND RESPONSIBILITIES:

Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform. In addition, as you discover and document system bugs, you have the motivation to go off and fix them yourself.

Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines.

Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools

Identify performance bottlenecks and optimize the performance of cloud infrastructure

Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code.

Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies.

Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.

Create, script, and run performance tests to measure system behavior under varying levels of load and traffic. Identify bottlenecks, performance degradation, and areas for optimization.

Design, implement, and maintain automated test suites for infrastructure and application components. Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release.

Build automated systems for continuous performance testing, stress testing, and load testing.

Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals.

Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.

Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.

Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks.

Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management

BASIC QUALIFICATIONS:

Requires BS degree and Master's degree.

Currently possess and ability to maintain an active DoD Secret security clearance

Minimum of DoD IAT Level II Certification required prior to onboarding and must maintain certification while supporting the SMIT Contract

Must be able to support program execution in classified environments and access SIPRNet from an NMCI location on short notice (local travel).

Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred

Experience designing, deploying, and maintaining Microsoft Endpoint Configuration Manager (MECM) environments at enterprise scale.

Proven ability to optimize MECM infrastructure for performance, reliability, and high availability.

Hands‑on experience creating and managing MECM device collections, task sequences, applications, packages, and operating system deployment images.

Expertise in automating MECM administrative tasks using PowerShell, APIs, and infrastructure‑as‑code tooling.

Strong background in Windows endpoint management, including patching strategies, compliance baselines, configuration items, and reporting.

Demonstrated ability to improve endpoint reliability through monitoring, observability tooling, and proactive remediation.

Experience integrating MECM with cloud‑based services such as Intune, Azure AD, and Windows Update for Business.

Ability to troubleshoot complex MECM client and server issues using logs, diagnostics, and platform health data.

Experience supporting enterprise‑wide operating system deployments and in‑place upgrades using MECM automation pipelines.

Knowledge of SQL Server administration relevant to maintaining MECM databases and enabling performance tuning.

Experience contributing to system reliability through continuous improvement, infrastructure monitoring, and configuration drift reduction.

Strong understanding of networking concepts supporting MECM components such as DP distribution, boundary groups, and PXE services

Ability to work in a highly collaborative, forward thinking, and innovation‑driven environment

Knowledge of Agile and DevSecOps/SRE concepts and best practices, with a desire to grow that knowledge

Hand‑on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.).

Experience creating JIRA and/or Azure DevOps workflows, projects, custom configurations

Experience administrating/maintaining SRE platform via Ansible playbooks (e.g. upgrading Jenkins)

Experience in automating tasks with scripting languages like PowerShell, or Python

Integrating/maintaining with various 3rd party CI/CD tools like Jenkins and Gitlab.

Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.

Experience with automated provisioning and configuration tools like Terraform, Cloud Formation, Chef, Puppet, Ansible, or similar technologies.

Working knowledge of the Risk Management Framework (RMF), DISA STIGs

PREFERRED QUALIFICATIONS:

List additional skills and experience that is "nice to have" but not required.

Previous work experience providing support to the NGEN-NMCI program

Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation for automating test environments.

ITILv4, Scrum Master, or Agile SAFe certification(s) or applicable experience

KEY METRICS OF SUCCESS FOR THE TEAM:

Improved system reliability, as measured by adherence to Service Level Objectives (SLOs) and reduced Mean Time to Recovery (MTTR).

Comprehensive and regularly updated automated test coverage for all critical systems and infrastructure components.

Timely identification and resolution of performance bottlenecks and failure points.

Integration of automated testing into the CI/CD pipeline, ensuring continuous reliability validation.

Increased scalability and performance of systems under high load due to effective performance testing.

LOCATION:

Onsite

Norfolk, VA and Pearl Harbor, HI

ADDITIONAL INFORMATION:

CLEARANCE REQUIREMENT:

Must possess an active DoD Secret clearance. In addition, selected candidate must undergo background investigation (BI) and finger printing by the federal agency and successfully pass the preceding to qualify for the position. US CITIZENSHIP IS REQUIRED.

CRITICAL SOLUTIONS PAY AND BENEFITS:

Salary range $85,000 - $110,000. The salary range for this position represent the typical salary range for this job level and this does not guarantee a specific salary. Compensation is based upon multiple factors such as responsibilities of the job, education, experience, knowledge, skills, certifications, and other requirements.

BENEFIT SNAPSHOT: 100% premium coverage for Medical, Dental, Vision, and Life Insurance, Supplemental Insurance, 401K matching, Flexible Time Off (PTO/Holidays), Higher Education/Training Reimbursement, and more.

#J-18808-Ljbffr

Worksite address

norfolk, VA, 23500, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Annapurna Labs (u.s.) Inc.

WhatJobs

PD Engineer, Annapurna Labs

cupertino, CA

Salary not specified

As a member of the Cloud-Scale Machine Learning Acceleration team you'll be responsible for the design and optimization of Hardware in our data c…

Last received from source 2026-10-07View job

Amazon.com Services Llc - A57

WhatJobs

Senior Automation Engineer

suffolk, VA

Salary not specified

Operations is at the heart of Amazon's business. We are known for our speed, accuracy, and exceptional service. Our buildings deliver tens of tho…

Last received from source 2026-10-07View job

Annapurna Labs (u.s.) Inc.

WhatJobs

DFT Design Engineer, Machine Learning Acceleration

austin, TX

Salary not specified

Custom SoCs (System on Chip) are at the heart of AWS Machine Learning servers. As a member of the Cloud-Scale Machine Learning Acceleration team,…

Last received from source 2026-10-07View job