waypointjobs

Maxinsights

Machine Learning Engineer (Video Understanding & Segmentation)

santa clara, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Maxinsights lists this Machine Learning Engineer (Video Understanding & Segmentation) opportunity in santa clara, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Job Description: We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Responsibilities

Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.

Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.

Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.

Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.

Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.

Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.

Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.

Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.

Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.

Minimum Qualifications

MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.

3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.

Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.

Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.

Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).

Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).

Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).

Preferred Qualifications

PhD with a research focus in video understanding, multi-modal learning, or vision-language models.

Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.

Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.

Track record of deploying production-scale video search or retrieval systems.

Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.

Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).

Experience with humanoid robotics or embodied AI data pipelines is a plus.

Default Benefits

Health insurance

Vision care

Dental coverage

401(k)

Paid holidays

PTO (Paid Time Off)

Sick leave

#J-18808-Ljbffr

Worksite address

santa clara, CA, 95053, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Space Systems Mechanical Engineer

laurel, MD

See pay details in description

Description Do you want to design and build unique space structures and spacecraft for NASA missions that enable groundbreaking scientific disco…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Security Engineer

laurel, MD

See pay details in description

Description Are you looking for an opportunity to utilize your technical skills to solve complex, real-world problems? If so, we're looking …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Advanced Reentry Mission Engineer

laurel, MD

See pay details in description

Description Are you interested in hypersonic and reentry system design and prototyping? Do you want to make contributions to next generation …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Thermal and EO/IR Modeling and Simulation Engineer

laurel, MD

See pay details in description

Description Are you looking for a unique opportunity to impact significant advances to the nation's groundbreaking integrated air and missile de…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Network Effects Engineer

laurel, MD

See pay details in description

Description Do you want to perform advanced research, development, and test & evaluation of communications systems and network technologies that…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Realization and Resilience Engineer

laurel, MD

See pay details in description

Description Are you passionate about applying system engineering principles to influence the development and resilience of future strategic weap…

Listing review due 2026-10-06View job