Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Staff Research Engineer - Multimodal Generative Modelling

Requirements

  • The ability to bring novel ideas and designs that advance the field of interactive multimodal systems.
  • Strong understanding of generative modelling, ideally applied to sequential or multimodal data.
  • Hands-on experience with large language models or similar transformer-based architectures.
  • High proficiency in PyTorch, including distributed training and model optimization.
  • A solid grasp of time-series modeling and tokenization, preferably in the context of audio, speech, or video.
  • A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently.
  • Proven experience training deep learning models end-to-end, from data preparation through evaluation.
  • Strong general software engineering skills, enabling contributions to a large, shared research infrastructure.

Nice to Haves

  • Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it.
  • Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints, not afterthoughts.
  • Working on LLMs with large scale trainings leading to models with decent reasoning capabilities.
  • Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment.
  • Collaborating across modalities or teams (e.g. audio and video, or research and product) to ship a unified system.

What you'll be doing

  • Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons.
  • Propose novel multi-modal system architectures (especially text and voice).
  • Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
  • Design solutions that reinforce emotional expressiveness and natural interaction.
  • Implement and bring designs to life, from pretraining through post-training.
  • Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness.
  • Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements.
  • Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs.
  • Curate new datasets to complement existing data.
  • Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality.
  • Ship models to production with optimized runtime to serve customers, and address their feedback thereafter.

Perks and Benefits

  • Experience with real-time or streaming architectures.
  • Familiarity with state-of-the-art architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders.
  • Excellence in one or more of the following modalities: voice, text, video.
  • Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).
AI Summary ✨
Synthesia logo

Synthesia

Remote EMEA

Remote EMEA
Experience: Staff
Posted: September 16, 2026
Last seen: an hour ago
machinelearning

Why we track Synthesia

Synthesia is the London-based AI video platform (avatars, dubbing, generative video) and one of the UK's biggest AI scale-ups. Engineering and research sit in London, with several roles open to anyone in Europe or remote in the UK. Senior engineers in London earn around £184K on levels.fyi (median £137K), and remote-Europe L5 roles come in at about €116K.

Similar jobs

  • 5 days ago
    Remote
  • 5 days ago
    Remote
  • 5 days ago
    Remote
  • See all jobs in undefined