Stord

Senior Computer Vision Engineer (Egocentric), Data Foundry

Stord Atlanta, Georgia, United States

Transportation, Logistics, Supply Chain and Storage · 1,001-5,000 employees

Yesterday
Principal (10+ yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the end-to-end egocentric video stack, including data collection, vision models, and hardware rigs. You will also lead the warehouse capture operation and coordinate across teams to deliver high-quality datasets on schedule.

What they look for

Computer Vision Egocentric Perception Python C++ Machine Learning Robotics Camera Calibration 3D Reconstruction Pose Estimation CNN Vision Transformers Multimodal Models Data Pipelines Hardware Integration Geometric Computer Vision

Requirements

Candidates must have 8+ years of experience building production computer vision systems or an advanced degree with 6+ years of hands-on experience. Strong expertise in geometric computer vision, Python, and C++ is required, along with a proven track record of scaling perception stacks.

Full description

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord is rapidly growing and is on track to double our revenue in the next 18 months. To meet and exceed this target, Stord is strategically scaling teams across the entire company, and seeking energetic experts to help us achieve our mission.

By combining comprehensive commerce-enablement technology with high-volume fulfillment services, Stord provides brands a platform to compete with retail giants. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator-built software suite including OMS, Pre- and Post-Purchase, and WMS platforms. Stord is leveling the playing field for all brands to deliver the best consumer experience at scale.

With Stord, brands can increase cart conversion, improve unit economics, and drive sustained customer loyalty. Stord’s end-to-end commerce solutions combine best-in-class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order.

Hundreds of leading DTC and B2B companies like AG1, True Classic, Native, Seed Health, quip, goodr, Sundays for Dogs, and more trust Stord to deliver industry-leading consumer experiences on every order. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top-tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

About the role

Stord operates the largest independent e-commerce fulfillment network in the US — 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are building a new business line that turns this operational infrastructure into some of the most valuable training data assets in physical AI.

We are looking for an experienced computer vision engineer and technologist to build and scale this business from the ground up.

What You Will Own

You will own the early egocentric video stack -- data collection, vision models and pipelines, and rigs. You'll partner closely with a small team to operationalize. This is a builder-operator role. You will:

  • Define and deliver the product. You will own the data product across quality tiers — from RGB egocentric video through depth-enhanced and full multimodal capture with hand pose and annotations. You will decide what gets built, in what order, based on what buyers will actually pay for. You will hold the line on quality.
  • Run the capture and delivery program. You will stand up the warehouse capture operation: camera and rig hardware selection, enrollment, edge processing, and the processing pipelines that package datasets for delivery. You will coordinate across warehouse operations, engineering, and customers to ship datasets on spec and on schedule.
  • Build the perception stack. Detection, tracking, and segmentation, plus depth/3D reconstruction and 6DoF, multi-view 3D hand/body pose estimation from egocentric and fixed-camera capture.
  • Stand up VLM-assisted and automated labeling with human-in-the-loop QA to drive down cost per annotated hour, and integrate the annotation tooling.
  • Own the hardware<>vision intersection. Camera calibration, epipolar/multi-view geometry, and frame-accurate time-sync across multi-camera and egocentric rigs; derive 3D pose by triangulation where no direct sensor exists.
  • Train and ship models. Design, fine-tune, and optimize CV/multimodal models on large unstructured video datasets, and get them reproducible and production-ready, not stuck in a notebook.

What You Bring

  • Experiencing standing up and scaling an egocentric perception stack. You have built and run a similar product end to end at a robotics or AI data company. You have driven the full lifecycle: hardware setup, embedded perception, data pipelines, ensuring quality, and delivering it to production teams who depend on it.
  • 8+ years building and shipping production computer-vision/perception systems (or an MS/PhD in CV, ML, or robotics plus 6+ years hands-on), including systems that ran on messy real-world data, not just benchmarks.
  • Deep expertise in computer vision and tooling — track record of leveraging existing tooling and designing, training, and debugging CNNs and vision transformers from scratch.
  • Strong command of geometric computer vision: camera calibration, depth estimation, and 2D/3D pose estimation
  • End-to-end ownership of a major perception problem: from data and model design through evaluation, optimization, and deployment, with measurable accuracy and reliability outcomes.
  • Track record of setting technical direction for a team or large workstream and raising the bar for other engineers.
  • Proven ability to take ambiguous, 0→1 problems with no established playbook and drive them to a working system with limited resources.
  • Experience with large unstructured datasets (video/multimodal) and the eval discipline to instrument accuracy rather than eyeball it.
  • Expert Python and strong software-engineering fundamentals; C++ where performance demands it.

Why This Role

 This is a rare opportunity to build a high-growth business from the ground-up with infrastructure and resources to support. You will have:

  • A structural moat that no startup can replicate
  • Direct access to the fastest-growing buyer market in AI
  • CTO/Co-Founder as your direct partner.