About this opportunity
ZealHire Inc. lists this Principal Machine Learning Engineer – Agentic AI Platforms opportunity in northern, Kentucky. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Principal Machine Learning Engineer – Agentic AI Platforms | JPC-8589
USC/GC/H4-EAD- W2Need LinkedIn, VISA DL100% remote
Position Summary
We are seeking a Principal Machine Learning Engineer with 10+ years of experience building large-scale distributed machine learning systems and modern AI platforms. This role will lead the architecture and implementation of Agentic AI ecosystems, LLM infrastructure, RAG platforms, model-serving systems, and AI engineering frameworks supporting enterprise-wide AI adoption.
This is a deeply hands‑on role requiring expertise in platform architecture, ML systems design, distributed computing, MLOps, LLMOps, and production AI deployment.
Key Responsibilities
Agentic AI Platform Engineering
Design and implement enterprise-grade agent frameworks supporting: Multi-agent collaboration
Planner‑executor architectures
Tool‑use workflows
Memory systems
Human‑review workflows
Build orchestration services and runtime environments for AI agents.
Implement MCP and agent interoperability standards.
Design state-sharing and context management systems.
LLM Platform Development
Build scalable LLM platforms supporting: OpenAI
Claude
Gemini
Llama
Mistral
Fine‑tuned models
Implement model routing and workload balancing across models.
Design low‑latency AI serving architectures.
RAG & Retrieval Engineering
Build retrieval systems using: Hybrid search
Dense retrieval
BM25
Re‑ranking pipelines
Graph‑based retrieval
Optimize vector stores and embedding pipelines.
Design retrieval infrastructure capable of supporting billions of documents.
MLOps & LLMOps
Build automated pipelines for: Training
Evaluation
Deployment
Monitoring
Retraining
Implement CI/CD pipelines for AI systems.
Build observability frameworks including: OpenTelemetry
LangSmith
LangFuse
Custom tracing
Reliability & Evaluation
Implement release gates and quality controls.
Develop evaluation pipelines for: Hallucination detection
Retrieval quality
Agent performance
Cost optimization
Latency management
Design monitoring frameworks for drift, degradation, and reliability.
Technical Leadership
Lead architecture reviews for AI systems.
Establish engineering standards and best practices.
Mentor engineering teams.
Partner with Data Scientists, Product Managers, and Business Stakeholders.
Required Qualifications
Bachelor's, Master's, or PhD in Computer Science, Engineering, Data Science, or related field.
10+ years of software engineering or machine learning engineering experience.
5+ years building production ML systems.
3+ years building production GenAI platforms.
Strong expertise in: Python
SQL
Spark
Distributed computing
Kubernetes
Docker
Experience with
LangGraph
LangChain
LlamaIndex
Vector databases
FastAPI
REST APIs
Cloud platforms (AWS, Azure, GCP)
Preferred Qualifications
Experience building multi‑agent platforms at enterprise scale.
Experience implementing MCP and A2A architectures.
Knowledge graph and GraphRAG experience.
Fine‑tuning experience (LoRA, QLoRA, PEFT).
Experience deploying AI systems in healthcare, finance, legal, insurance, or cybersecurity domains.
Ideal Candidate Profile
The strongest candidates for either role should be able to confidently discuss: Agent architecture tradeoffs
Retrieval evaluation methodologies
Hallucination measurement
LLM evaluation frameworks
Fine‑tuning vs RAG decisions
Production AI failures and learnings
Cost optimization strategies
MCP/A2A interoperability
AI governance and model risk management
Scaling AI systems from prototype to enterprise production
Significant experience operating AI systems supporting thousands to millions of users.
$000,000 – $000,000 Salary range is not shown on the candidate portal
#J-18808-Ljbffr
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.