About this opportunity
Tala lists this Manager of Machine Learning Engineering opportunity in new york, New York. Review the employer’s description below for duties, qualifications and application requirements.
Job description
We’re looking for a Manager, Machine Learning Engineering to lead Tala’s ML Platform team. This person will manage a team of Machine Learning Engineers responsible for building the platforms, frameworks, and infrastructure that enable our Data Science teams to securely train, deploy, monitor, and operate machine learning models at scale
This is a player-coach management role. You’ll be responsible for developing and growing the team while also providing enough technical leadership to guide architecture, engineering practices, reliability, and production systems. The role has a particular focus on real-time machine learning inference and streaming data systems, as well as the platforms that support batch model development and deployment
Manage and develop a team of 4–6 Machine Learning Engineers across mid-to-senior levels
Hire, source, interview, and close strong MLE talent
Establish clear expectations, provide regular feedback, and create development plans for direct reports
Coach engineers toward growth and promotion while addressing performance gaps directly and thoughtfully
Create opportunities for engineers to take on challenging projects and grow their technical leadership
Set quarterly goals and ensure the team consistently delivers against them
Own prioritization across product roadmap work, run-the-business activities, and operational excellence
Balance team capacity across new development, maintenance, technical debt, and production support
Improve team productivity by reducing context switching and delegating effectively
Partner with engineers and technical leads to estimate and scope complex work
Guide the development of platforms and frameworks that allow Data Scientists and Analysts to explore data, develop features, and train, test, deploy, and monitor ML models
Provide technical leadership across model infrastructure, real-time inference, streaming feature extraction, batch processing, and production ML systems
Drive strong engineering practices around testing, automation, observability, fault tolerance, infrastructure-as-code, and deployment
Own and improve SLOs, on-call health, capacity planning, reliability, and incident response
Review technical designs and help drive architectural standards and technical debt reduction
Work closely with Data Science, Data Engineering, Data Platform, Product, Credit, and Business Development teams
Translate business and technical needs into scalable ML platform solutions
Coordinate dependencies and delivery across multiple engineering and data teams
Help create structure and clarity in an environment where priorities and requirements can evolve
Requirements
Experience owning team goals, prioritization, estimation, and delivery
2+ years of directly managing engineers, including hiring, performance management, coaching, and career development
Experience managing a team through at least one full performance cycle
Demonstrated ability to coach engineers toward promotion and address underperformance effectively
Willingness to be actively involved in sourcing, interviewing, and closing engineering talent
Experience with production on-call, incident response, and capacity planning
Strong understanding of software quality, security, reliability, testing, and production operations
Experience building and operating machine learning or causal inference systems in production
Earlier-career experience personally building and deploying ML models or ML infrastructure
At least 3 years of hands-on Python experience
Ability to participate in technical architecture and system-design discussions and provide technical direction without needing to be the primary coder
6+ years of backend software engineering experience in consumer-scale applications
Technical Skills
Cloud & Infrastructure: AWS, GCP, Azure, Kubernetes, Docker
Streaming: Kafka, Kinesis, Beam, Flink, Spark Streaming
Languages: Python, SQL
ML/Analytics: Machine learning, causal inference, scalable algorithms
APIs: REST, GraphQL, gRPC, Protocol Buffers
Production Engineering: DevOps, SLOs, monitoring/observability, on-call, capacity planning, root-cause analysis
Batch Processing: Airflow, Metaflow
Machine Learning: Jupyter, Pandas, Scikit-Learn, XGBoost, TensorFlow, PyTorch, Hugging Face
Databases: MySQL, PostgreSQL, Cassandra, Snowflake, Druid, and/or similar technologies
We’re particularly interested in candidates with experience across:
#J-18808-Ljbffr
Worksite address
new york, NY, 10261, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.