About this opportunity
Alexander Chapman Ltd lists this Machine Learning Infrastructure Engineer opportunity in new york, New York. Review the employer’s description below for duties, qualifications and application requirements.
Job description
We're looking for a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists and ML engineers rely on to train, deploy, and monitor models. You'll focus less on the models themselves and more on making the systems around them fast, reliable, and easy to use.
What You'll Do
Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling)
Build internal tooling for experiment tracking, feature stores, and model versioning
Optimize model serving for latency, throughput, and cost at scale
Set up and maintain CI/CD pipelines for ML workflows
Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm)
Implement monitoring and alerting for model performance, data drift, and system health
Partner with ML engineers and data scientists to understand their workflows and remove friction
Contribute to platform architecture decisions as the ML org scales
What We're Looking For
2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads
Strong proficiency in Python and/or Go
Experience with containerization and orchestration (Docker, Kubernetes)
Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar
Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi)
Understanding of distributed systems and data pipeline designComfort working with GPUs and understanding of training/serving performance trade-offs
Strong communication skills and ability to work cross-functionally with ML practitioners
Nice to Have
Experience with high-performance model serving (Triton, TorchServe, vLLM)
Familiarity with data versioning tools (DVC, LakeFS) or feature stores (Feast, Tecton)
Experience scaling distributed training (Horovod, DeepSpeed, PyTorch DDP)
Background in SRE or DevOps practices applied to ML systems
What We Offer
Competitive salary and equity
Health, dental, and vision insurance
#J-18808-Ljbffr
Worksite address
new york, NY, 10261, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.