About this opportunity
Moe lists this Founding ML Engineer, Computer Vision (Item Identification) opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Moe's entire pitch to users rests on one claim: we can identify any item in the world from a photo and price it as accurately as a human expert, instantly. You'll build the model that makes that claim true. This isn't a research exercise — every category you get right becomes a category users trust Moe with real money, and every one you get wrong becomes a support ticket and a churn risk. You'll set the technical direction for identification from day one, with real ownership over architecture, data strategy, and the accuracy bar we hold ourselves to.
What you'll do
Design and own the computer vision architecture for fine-grained item identification — brand, model, edition, variant — starting from foundation vision models and fine-tuning toward Moe's specific catalog
Define what "accurate enough" means per category, and build calibrated confidence scoring so the product can say "we're not sure" instead of guessing
Build the feedback loop between model errors and what gets labeled next, in partnership with the labeling lead
Decide where to invest: broader category coverage vs. deeper accuracy on today's categories
Own the model serving path from research to production — latency, cost, and reliability at scale
Represent the identification model's capabilities and limits to the rest of the company, including in investor and customer conversations when needed
What we're looking for
5+ years in applied computer vision, with at least one system shipped to production at meaningful scale
Hands-on experience with fine-grained/instance-level classification, not just general object detection — you've worked on a problem where "close" isn't good enough (e.g., telling two similar sneaker colorways or watch references apart)
Strong fluency in PyTorch or TensorFlow, and experience fine-tuning and deploying vision transformers or CNNs in production
Experience designing and running evaluation frameworks for vision models — you know how to measure whether a model is actually getting better, not just achieving a lower loss
Comfortable being the most senior technical voice on a hard, open-ended problem with no existing internal playbook
Strong written and verbal communication — you'll need to explain technical tradeoffs to non-technical stakeholders, including investors
Nice to have
Prior work at a resale/marketplace company (StockX, GOAT, Vinted, Rebag, The RealReal) or a visual search company (Pinterest Lens, Google Lens, Syte)
Experience with active learning or human-in-the-loop labeling pipelines
Familiarity with deploying models behind low-latency APIs at scale
Founding or early-stage startup experience, ideally as the first ML hire
What success looks like
30 days: fully ramped on the current state of the identification problem; has picked the first category to ship (e.g., sneakers or watches) and defined the accuracy bar for it
60 days: v1 identification model is live for that category, with a measured accuracy baseline against a held-out test set
90 days: confidence scoring is in place, and there's a clear, prioritized plan for expanding into the next 2–3 categories
Process
Comp: $200K–$260K base + 0.75%–1.5% equity (negotiable for the right candidate)
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.