Machine Learning Engineer - Voice Conversion
Research/ML engineer building large-scale speech models for voice conversion and TTS at Cantina. You own the model-data-eval flywheel end to end, from training to production inference.

Cantina is a San Francisco-based AI startup specializing in advanced speech and video models, including text-to-speech, voice conversion, and joint audio-video modeling.
Research/ML engineer building large-scale speech models for voice conversion and TTS at Cantina. You own the model-data-eval flywheel end to end, from training to production inference.
This role involves building and scaling data pipelines for large video generation models. You will manage the full lifecycle of training data, from ingestion to model-ready samples.
Cantina is a San Francisco-based AI startup specializing in advanced speech and video models, including text-to-speech, voice conversion, and joint audio-video modeling. The company operates at the forefront of generative AI, building proprietary models and infrastructure for real-time voice and video applications. As a young, research-driven company, Cantina focuses on pushing the boundaries of what's possible in multimodal AI, with a lean team of engineers and researchers.
Cantina is expanding its engineering presence in Spain, hiring machine learning engineers for roles in TTS, voice conversion, and data & ML infrastructure. The Spanish team will work closely with the San Francisco headquarters, contributing to core model development and evaluation. For international professionals, Cantina offers the opportunity to work on cutting-edge AI research in a fast-paced startup environment, with remote-friendly culture and a focus on technical excellence.
Cantina is a San Francisco-based startup focused on building AI models for speech and video, with a strong emphasis on voice conversion and text-to-speech technologies. The company is relatively new and operates in the competitive AI research and development space, attracting top engineering talent.