M

Founding AI Engineer (Multimodal Vision-Language Systems) - San Francisco

meeboss United State
Visa Sponsorship
Apply
AI Summary

Join meeboss as a Founding AI Engineer to design, build, and deploy production-grade agentic vision-language models (VLMs) for industrial smart glasses. You’ll own end-to-end pipelines, optimize edge inference, and drive real-time multimodal AI systems for field service workflows. Requires 3-5 years of applied AI/ML experience with a focus on shipping multimodal systems in fast-paced environments.

Key Highlights
Build and ship production agentic-VLM pipelines for industrial smart glasses with real-time visual reasoning and tool integration
Optimize edge inference for latency, power, and connectivity constraints with graceful fallbacks
Own model orchestration, evaluation harnesses, and data flywheels for measurable model improvements
Key Responsibilities
Design and build production agentic-VLM pipelines for industrial smart glasses, enabling multi-step visual reasoning and tool integration
Optimize model orchestration and runtime for edge inference, balancing latency, power, and connectivity with graceful fallbacks
Develop evaluation harnesses and data flywheels to capture failure modes, fine-tune models with customer data, and drive measurable quality improvements
Build real-time voice-video AI interfaces tailored to diverse end-user profiles (video-heavy, conversational speech, proactive alerts)
Construct RAG pipelines for enterprise knowledge bases using field operator data and integrate with customer-specific workflows
Deploy and fine-tune open-source models (SFT, RLHF, quantization) for on-premise industrial deployments
Technical Skills Required
Python Vision-Language Models (VLM) Model Orchestration & Edge AI
Benefits & Perks
$180K - $240K annual salary
Visa sponsorship for H-1B transfers and TN visas (no new H-1B sponsorship)
On-site work policy (5 days/week in San Francisco)
Nice to Have
Wearable AI or autonomous driving experience (e.g., Meta Reality Labs, Snap, Apple Vision)
Industrial domain exposure (data centers, energy grid, aerospace, manufacturing)
Experience with wearable AI, AR, or real-time video/streaming systems
Founder/CTO or early-stage startup experience in AI product development

Job Description


We’re seeking Founding AI Engineer to join Full-time onsite in San Francisco, CA 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems. Location : San Francisco, CA On-site work policy: On-site 5 days per week in San Francisco. Full-time position Salary: $180K - $240K Hiring Count: Looking to hire 2 candidate for this role Tech stack: Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, LangChain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker Visa sponsorship available: For H-1B transfers and TN visas supported. No new H-1B sponsorship. ⚠️ Candidates must currently reside in the USA or Canada. About This Role We're looking for an AI Engineer with 1-5 years of experience in applied AI/ML who has shipped agentic multimodal systems to real users — not just RAG wrappers or single-turn chatbots. You should be comfortable building production VLM pipelines with real hardware constraints (latency, power, connectivity) and have a track record of owning AI systems end-to-end in fast-paced environments. Bonus points if you have wearable AI, autonomous driving, or industrial domain experience. What you'll do:

  • Build and ship the production agentic-VLM pipeline running on industrial smart glasses — multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service)
  • Own model orchestration and runtime optimization for edge inference, trading off model quality against latency with graceful fallback across connectivity conditions
  • Design and build the eval harness and data flywheel from scratch — failure-mode capture, customer-data fine-tune loops, and measurable model quality improvements each release
  • Ship real-time voice-video AI interfaces adapted to different end-user profiles: video-heavy, conversational speech, and proactive alerts
  • Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data
  • Drive multimodal model training when needed for on-premise deployments: open-source model SFT, RL post-training, and quantization

Role requirements Seniority · 3 - 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems. Work experience · Shipped multimodal and computer vision systems in the VLM era. (production, not demos, not pure research). Owned the model layer end to end · Experience at a startup OR AI team building production AI products, preferably as a founder/CTO/founding engineer. · Big- company tenure (Meta Reality Labs, Snap, Apple Vision, Google, big streaming infra) is a BONUS only when on a directly- relevant team (AR/smart- glasses, real- time video/streaming, on- device/edge ML) OR paired with a builder signal (founder / early- startup / side projects / OSS). A long single big- tech tenure with no builder signal and no relevant- team work is a NEGATIVE. · Production AR / wearable AI experience (Meta Reality Labs, Snap, Apple Vision) or autonomous driving Computer Vision. · Industrial domain exposure (data centers, energy grid, aerospace, manufacturing). Education · Strong CS/ML/Eng background OR demonstrated equivalent shipping record. · Master's with vision or multimodal research component. Hard skills Applied VLM / multimodal engineering: makes VLMs reliably do visual reasoning in production (image/video understanding, detection/segmentation as needed) in the modern VLM/VLA era. Depth is in shipping, hardening, and applied fine- tuning, not pretraining from scratch — this means VISION- language / video- language multimodality (images, video, VLM/VLA), NOT sensor- fusion, materials, audio- only, or time- series 'multimodal'. Applied agentic AI / model orchestration experience (vs. pure research) Evals discipline: rigorous eval harnesses (ground- truth, trajectory and tool- call accuracy, regression) to compare models/orchestrations and drive iteration On- prem / self- hosted model deployment: serving and optimizing open- weight ML models on customer hardware (for high- IP environments). Hands- on fine- tuning and deployment of a vLLM is a proxy In- context grounding / RAG against a knowledge base, with tool- and- KB wiring (a real plus; FDEs own the customer- specific integration) Soft skills Balances core AI depth with applied product mindset. Genuinely mission- aligned with CV/wearables/industrial AI. Miscellaneous · Ideally willing to join a hacker house (live on site). At least willing to work on- site 5 days/week in SF. · We are not looking Pure researcher focused on pretraining / training from scratch with no shipped product · Classical CV vocabulary only (Faster R- CNN, IOU segmentation) without VLM and agentic awareness. · We are not looking Pure Big- tech- only profile without ownership of a shipped AI product. · 'Multimodal' that means sensor/materials/signal fusion or audio- only/time- series rather than vision- language. 💰 $5,000 Referral Bonus for successful placements! If you have the required skills and experience, please send your resume (with details) to: 📧 skumar@cognistack.co · Join WhatsApp Group: StartupStack WhatsApp Group


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Associate

agilegrid solutions

United State

People Operations Manager

Programming
1h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

fluidstack

United State

Head of Talent

Programming
2h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Wispr Flow

United State

Subscribe our newsletter

New Things Will Always Update Regularly