Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Member of Technical Staff - VLM

Requirements

  • You've pretrained or significantly advanced a VLM (not just SFT'd or LoRA'd one) that was deployed in a production system or released publicly
  • Strong publication record or unambiguous production track record showing you push the frontier on multimodal architectures
  • Deep understanding of how vision and language representations interact: tokenization, alignment, grounding, cross-modal attention, and the failure modes of each
  • Experience with distributed training at multi-node scale
  • Comfortable at the research/production boundary — you care whether the work ships and generalizes, not just whether it reads well
  • Experience with diffusion or flow-based generative models is a strong plus — especially if you've thought about how autoregressive and diffusion paradigms can compose

What You'll Be Doing

  • Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack — innovating on architectures, not just applying existing ones
  • Design fine-tuning strategies that adapt VLMs to specialized creative use cases (captioning, editing instructions, prompt enhancement) that general-purpose models can't handle
  • Research integrations between VLM/LLM capabilities and our diffusion and flow pipelines — finding creative ways to improve generation quality and controllability without computational bottlenecks
  • Evaluate emerging multimodal architectures, translating the best of recent research into practical improvements

Nice to Haves

  • Experience with diffusion or flow-based generative models

Perks and Benefits

  • Base Annual Salary: EU €130,000-€340,000 + Equity
AI Summary ✨

Similar jobs

  • 3 hours agoNew
  • 3 days agoRemote
  • 4 days agoRemote EMEA
  • 6 days ago
  • See all jobs in Germany →